Director, Research - AI Evals at Figma
New York, NY
$258,000–$348,000
Jul 13, 2026
New York, NY
Jul 21, 2026
What this job asks for AI summary
A director-level research role focused on building and owning the evaluation infrastructure for Figma's AI-powered product features. Day-to-day work involves defining quality standards, constructing rubrics and benchmark datasets, combining human and automated assessment methods, and producing decision-ready signal for product and engineering teams. Suited to a senior practitioner with hands-on AI/LLM evaluation experience who can also manage a small team and influence cross-functional stakeholders.
Senior level · 10+ years · Remote · Full-time
Pay in the description: $258,000–$348,000
Advertised as Principal, but the requirements read as Senior.
“or” means any one of them counts — you don't need all of them.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
What the occupation pays Median $122,874 (middle half $87,544–$162,374). This posting is about at that midpoint.
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.
Why we read it this way (9)
This role sits at the intersection of AI/ML evaluation research and people management (2+ years of management required, managing a small team). The primary day-to-day work is designing and running AI evaluation frameworks — closer to applied data science / research than pure engineering management — hence 15-2051 as primary SOC, with 11-3021 as a genuine runner-up given the team-management scope.
The 'Director' title implies a senior leadership level, but the actual scope is owning a single function (AI evals) with a small team, which maps to Senior rather than Staff or Principal on the IC/scope ladder. The advertised title is Director, which is treated as Principal for the level advertised in the title.
The salary range ($258K–$348K) is stated for SF and NY hub offices; remote roles are localized at 80–100% of that range per the posting.
Several skills listed under the required qualifications ('human evaluation programs', 'rubric/benchmark construction', 'inter-rater reliability', 'LLM-as-judge') are methodologies rather than named tools, but they are concrete, named evaluation techniques that function as hard gates for this specialized role and are listed explicitly in the requirements section.
Braintrust, LangSmith, and DeepEval appear under the 'nice to have' / 'added plus' section and are listed as interchangeable examples of eval tooling — emitted as a single preferred skill with alternatives.
'Evaluation pipelines' and 'regression testing' appear under the preferred/plus section as something built 'in partnership with engineering' — marked preferred accordingly.
Familiarity with Figma's own products is listed as a plus, not a requirement.
Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): data analysis.
This posting reads as a fully-remote role, so it was scored against the national candidate pool rather than a single metro.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.