remote

Salary
$258,000–$348,000from the description
Posted
Jul 15, 2026
Location
Last confirmed open
Jul 21, 2026

What this job asks for AI summary

A director-level role focused on building and owning the AI evaluation function for a product organization. The work centers on designing evaluation frameworks, rubrics, benchmark datasets, and automated pipelines to measure the quality of AI and LLM-powered features, then translating findings into guidance for senior leadership and product decisions. Suited to a seasoned research professional with hands-on AI evaluation experience who has also managed people and built new practices from scratch.

Senior level · 10+ years · Remote · Full-time

Advertised as Principal, but the requirements read as Senior.

Must have (4)
LLMsAI evaluationbenchmark creationinter-rater reliability
Nice to have (1)
Braintrust, Langsmith or Deepeval

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

What the occupation pays Median $178,991 (middle half $141,096–$225,584). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (7)

SOC classification is a genuine judgment call: the role carries explicit people-management duties (team lead, 2+ years management required) pointing to 11-3021, but the day-to-day work is deeply hands-on applied AI research and evaluation methodology, which could support 15-2051. 11-3021 was chosen because the Director title, direct reports, and organizational strategy ownership are the primary framing.

The advertised title is 'Director' — mapped to Principal advertised seniority — but the actual scope (small team, building a new function, individual contributor research depth) is consistent with a Senior band rather than a true Principal/org-wide authority role.

AI evaluation tooling (Braintrust, LangSmith, DeepEval) is explicitly called 'a strong asset' and listed separately from the Requirements block, so it is treated as preferred.

Several desirable backgrounds (product design, product management, data science, engineering, front-end development) are described as 'advantageous' — these are not emitted as skills because they name disciplines rather than concrete technologies.

The posting is listed via Jobgether on behalf of an unnamed partner company; the actual employer is not disclosed.

Ignored 3 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): rubric design, LLM-as-a-judge, people management.

Caller marked this a fully-remote role — scored against the national candidate pool.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗