Recruiting From Scratch · San Francisco, CAremote

Salary
$150,000–$300,000from the description
Posted
Jun 30, 2026
Location
San Francisco, CA
Last confirmed open
Jul 21, 2026

What this job asks for AI summary

This role centers on building and owning evaluation infrastructure for a production LLM-based document intelligence platform. Day-to-day work involves designing benchmarks, automated evaluation pipelines, and statistical methodologies to measure and improve model quality across large volumes of unstructured documents such as PDFs and OCR outputs. It suits engineers with hands-on experience in ML evaluation, Python, and production LLM systems, ideally from a startup or high-bar technical environment.

Mid level · 1+ years · Remote · Full-time

Must have (4)
PythonLLMsML evaluation pipelinesstatistical evaluation
Nice to have (5)
Flask or TypeScriptAWS S3TinybirdOLAPVision-Language Models

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

Roughly 71,000 people nationally plausibly meet what this posting asks for (software developers). range 20,900–106,500

Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.

What won't set you apart
Python51%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $138,970 (middle half $107,524–$175,762). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (9)

The job posting specifies onsite in San Francisco 5 days per week, but per caller instruction this is being treated as a fully remote, national-pool role.

The title carries no seniority level word; the experience sweet spot of 2–4 years (range 1–5) maps to Mid-level despite the wide $150K–$300K comp band, which likely reflects equity variance at a Series B startup rather than a senior scope.

SOC classification is a genuine judgment call: the role builds evaluation software and tooling (→ 15-1252 Software Developers) but is deeply ML-methodology-focused (→ 15-2051 Data Scientists); 15-1252 was chosen because the primary deliverable is production evaluation infrastructure and internal tooling, not modeling or statistical research.

Degree is listed as 'preferred' throughout the education section, so the degree requirement is set to None.

Flask and TypeScript are listed together as interchangeable lightweight web-app options ('Flask, TypeScript, or similar frameworks') under the technical requirements section and are captured as a single skill with alternatives.

AWS S3, OLAP systems, Tinybird, and Vision-Language Models are explicitly marked 'preferred' in the JD and are captured as preferred skills.

The skill 'unstructured data / document processing' consolidates the JD's repeated references to PDFs, OCR outputs, spreadsheets, and document extraction as a single hard-gate capability.

Ignored 2 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): LLM-as-a-Judge, unstructured data / document processing.

Caller marked this a fully-remote role — scored against the national candidate pool.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗