Beverly Hills, CA

Salary
$180,000–$280,000from the description
Posted
Jul 27, 2026
Location
Beverly Hills, CA
Last confirmed open
Jul 29, 2026

What this job asks for AI summary

This is a founding individual-contributor role responsible for building the evaluation and measurement infrastructure that governs model training and product decisions across an AI-powered student learning platform and a workforce learning product. The person hired will design LLM evaluation suites, construct internal multi-turn tutoring benchmarks, translate product telemetry into training signal, and own analytics and experimentation standards for the entire team. It suits a rigorous quantitative researcher with hands-on experience shipping LLM evaluations to production.

Senior level · 5+ years · Full-time

Must have (18)
PythonPandasNumPySciPyJupyterSQLLLMsexperimental designcausal inferenceBayesian statisticshypothesis testingRAGembeddingsfine-tuningMongoDBPostgreSQLvector databasesGCP
Nice to have (3)
agent workflowspsychometricsitem response theory

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

What gives you an edge
experimental design4%vector databases6%RAG8%embeddings10%SciPy12%GCP13%

Rare in this occupation — lead with these, and say what you built with them.

What won't set you apart
Python88%Pandas72%SQL72%Jupyter65%NumPy62%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $122,874 (middle half $87,544–$162,374). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 29, 2026. It is a model, not a headcount.

Why we read it this way (6)

The posting does not name a city or state; it says only 'in-person role.' The CBSA is unknown and has been left blank.

The role accepts either a PhD plus 5+ years of applied experience OR a non-PhD candidate with a strong track record owning AI evaluation in production — so no degree is hard-required.

The 5-year figure comes from the PhD pathway ('5+ years applying it to real products'); the non-PhD pathway has no stated minimum, so total experience minimum is treated as 5 years.

SOC classification is Medium confidence: the role is primarily about statistical evaluation, measurement science, and LLM evals rather than classic predictive modeling, making it a close call between Data Scientists (15-2051) and Operations Research Analysts (15-2031). The LLM fine-tuning and ML evaluation emphasis tips it toward 15-2051.

The stack section is framed as 'you don't need every item below, but you should be deep in most' — items are listed as a single block without a hard-required vs. preferred split. Skills have been marked required based on the depth of emphasis ('expert-level SQL', 'deep in most') and the overall framing; learning science / psychometrics / item response theory are explicitly called 'a real plus' and are marked preferred.

Equity is offered alongside base salary but no equity range is stated.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗