Scale AI · San Francisco, CA

Salary
$216,000–$270,000
Posted
Jul 13, 2026
Location
San Francisco, CA
Last confirmed open
Jul 21, 2026

What this job asks for AI summary

A senior ML engineering role focused on keeping production AI agents reliable over time. The work centers on building observability tooling, designing and automating evaluation frameworks, and creating systems that detect drift or misalignment in live agent behavior — then closing the loop with experiments that validate improvements before they ship. Suited to engineers with hands-on experience building production ML or LLM-powered systems end to end.

Senior level · 5+ years · San Francisco-Oakland-Berkeley, CA · Full-time

Must have (4)
LLMsagent architecturesML evaluation frameworksML observability/monitoring
Nice to have (1)
RLHF, Sft or Reward Modeling

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

What the occupation pays Median $173,851 (middle half $133,298–$217,653).

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (6)

The role sits at the intersection of applied ML engineering and systems engineering — it could reasonably map to Software Developers (15-1252), but the primary day-to-day work (evaluation frameworks, monitoring for drift/anomalies, experimentation, reward modeling) is closer to Data Scientists (15-2051).

No location is explicitly stated in the posting; Scale AI is headquartered in San Francisco, CA, so that metro is assumed.

Remote eligibility is not stated; defaulting to false.

The requirements section lists broad capability areas ('strong grounding in at least two of the following') rather than specific named tools or frameworks, making it difficult to extract discrete hard-gated technologies beyond LLMs and agent architectures. The skills above reflect the clearest named technical requirements.

RLHF, SFT, reward modeling, fine-tuning, and inference optimization all appear under the explicit 'Nice to have' section.

Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): inference optimization.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗