AI Engineer, Evaluation

Expired

This listing is older than 60 days and may no longer be accepting applications.

Distyl AI

This listing is older than 60 days and has likely been filled. Here are open roles like it:

Get new agentic engineering jobs in your inbox every Monday.

One curated email a week. No spam, unsubscribe anytime.

$150K - $250K/yr

Tech Stack

About the Role

ABOUT DISTYL AI

Distyl is an applied AI technology company partnering with the world's most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.

We research and deploy technologies that power AI-native operations, both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations, from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys.

Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s. The results reflect this approach: a 100% production deployment success rate for our customers and one of the few enterprise AI companies to run a profitable business.

WHAT WE ARE LOOKING FOR

At Distyl, we build AI systems using Evaluation-Driven Development, an approach where evaluation is not an afterthought, but the primary mechanism for iterating, improving, and trusting AI behavior in production.

AI Evaluation Engineers focus on designing and implementing the evaluation systems that drive this process. They are hands-on engineers who write production Python code, build evaluation pipelines, and use structured signals to guide system design, prompt iteration, and deployment decisions for real customer-facing AI systems.

This role is for engineers who believe that AI systems only improve when measurement is tightly coupled to development, and who want to apply that philosophy directly to systems that matter.

KEY RESPONSIBILITIES

  • Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments
  • Define how system quality is measured in each domain, ensuring that evaluation signals reflect real user needs, domain constraints, and business objectives
  • Build and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases. These test suites are treated as first-class system components that evolve alongside the AI system itself
  • Develop and maintain evaluation pipelines, offline and online, that integrate directly into system iteration loops. Evaluation results inform prompt design, agent logic, model selection, and release readiness
  • Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments. They investigate where evaluation signals diverge from real-world outcomes and refine grading approaches
  • Work closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks meaningfully guide system development and deployment in production

WHAT WE REQUIRE

  • 2+ years of software engineering experience
  • Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments. You treat evaluation code with the same rigor as application code
  • Experience with Evaluation-Driven or Experiment-Driven Development: Experience using structured evaluation or experimentation frameworks to drive system iteration
  • Ability to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders that scale
  • Systems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment
  • AI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration
  • Travel: Ability to travel 25-50%

WHAT WE OFFER

  • The base salary range for this role is $150K to $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package
  • 100% covered medical, dental, and vision for employees and dependents
  • 401(k) with additional perks (e.g., commuter benefits, in-office lunch)
  • Access to state-of-the-art models, generous usage of modern AI tools, and real-world business problems
  • Ownership of high-impact projects across top enterprises
  • A mission-driven, fast-moving culture that prizes curiosity, pragmatism, and excellence

Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday-Thursday) in-office.

See similar open roles

More jobs like this

InstaLILY logo
Software Engineer, Full Stack

InstaLILY

$150K - $190K/yr
🇺🇸
2026-07-20
Observe.AI logo
AI Agent Engineer, Client Facing

Observe.AI

$108K - $170K/yrLangChainLlamaIndex+4 more
🇺🇸
2026-07-20
Mixpanel logo
Software Engineer, AI Product Insights

Mixpanel

$188K - $254K/yrAgno
🇺🇸
2026-07-20
Zoox logo
Software Engineer - Tools & Automation

Zoox

$184K - $231K/yrAgno
🇺🇸
2026-06-28
Bunkerhill Health logo
AI Product Engineer

Bunkerhill Health

$160K - $260K/yr
🇺🇸
2026-06-28
Sema4.ai logo
Forward Deployed Engineer - AI Agent Engineer

Sema4.ai

CrewAILangChain+1 more
🇺🇸
2026-07-20
Guild.ai logo
AI Engineer, Production Agents

Guild.ai

🇺🇸
2026-06-28
NextGen Federal Systems logo
Application/Agentic AI Engineer

NextGen Federal Systems

AutoGenCrewAI+3 more
🇺🇸
2026-06-13
Docker logo
Staff Software Engineer, Agentic Platform

Docker

$170K - $276K/yrMCP
🇺🇸
2026-07-20
Everlaw logo
Senior Software Engineer, AI Platform

Everlaw

$173K - $251K/yr
🇺🇸
2026-07-20
Klaviyo logo
Senior Software Engineer, Agent Platform

Klaviyo

$148K - $222K/yrLangChainArize+6 more
🇺🇸
2026-07-20
Mixpanel logo
Senior Software Engineer, AI Product Insights

Mixpanel

$226K - $306K/yrAgno
🇺🇸
2026-07-20
Everlaw logo
Staff/Principal AI Engineer

Everlaw

$228K - $340K/yrAnthropicMCP+1 more
🇺🇸
2026-07-20
Writer logo
Software engineer, generative AI

Writer

$112K - $304K/yrPineconeWeaviate+1 more
🇺🇸
2026-07-20
Uncountable logo
Platform Engineer - Generative AI

Uncountable

$120K - $160K/yr
🇬🇧
2026-07-07
FloQast logo
Senior Software Engineer, Applications

FloQast

$144K - $216K/yrMCP
🇺🇸
2026-06-28
Allstate logo
AI Engineer Lead

Allstate

$152K - $222K/yrMCP
🇺🇸
2026-06-28
Metropolitan Commercial Bank logo
AI Engineer

Metropolitan Commercial Bank

$130K - $200K/yrLangChainLlamaIndex+2 more
🇺🇸
2026-06-28
Zoox logo
Senior/Staff Software Engineer - ML/AI Applications

Zoox

$225K - $249K/yrMCP
🇺🇸
2026-06-28
Zep AI logo
Senior AI Engineer

Zep AI

$180K - $250K/yrGoogle ADKLangGraph+1 more
🇺🇸
2026-06-24

Explore related roles

Get jobs like this weekly

Join 82 subscribers