Computational Linguist, AI Evaluation
TwelveLabs · San Francisco, CA
$140,000–$160,000
Jul 14, 2026
San Francisco, CA
Jul 21, 2026
What this job asks for AI summary
This role sits within an ML Data Operations team focused on video-language AI, where the core work involves designing model evaluation protocols, building data pipelines that surface model performance gaps, and managing large-scale labeling projects with external vendors. It suits someone with 5+ years in AI data operations who has hands-on experience with benchmark design, annotation workflows, and Python-based automation, and who is comfortable working across research and engineering teams.
Senior level · 5+ years · Remote
Posted 2 times — it's one opening, so apply once.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
Roughly 3,000 people nationally plausibly meet what this posting asks for (operations research analysts). range 890–4,550
Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.
Most people in this occupation already list these. Still required — just not what gets you shortlisted.
What the occupation pays Median $90,896 (middle half $69,863–$128,761).
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.
Why we read it this way (8)
The posting states the role is hybrid in San Francisco (onsite Tuesdays & Thursdays), but per caller instruction it is being treated as fully remote with no metro attached.
SOC classification is a judgement call: the role sits between data/operations analysis (15-2031) and data science (15-2051). The primary day-to-day work is building evaluation pipelines, labeling operations, and surfacing model-quality insights — closer to structured analysis and operations research than to ML modeling itself, so 15-2031 was chosen with 15-2051 as the runner-up.
The 'model evaluation pipelines' skill is a functional capability requirement (benchmark design, human eval frameworks, automated scoring) stated in the main qualifications block, not a named tool — it is retained because it is a concrete, gateable competency central to the role.
AWS (S3, EC2), FFMPEG, Encord, Notion, and Linear appear only in the 'Tech Stack' narrative section, not in a requirements or qualifications block, so they are treated as preferred/stack context rather than hard gates.
LLMs/VLMs are listed as a required 'foundational understanding' in the main qualifications block; 'VLMs' is captured under the canonical 'LLMs' skill with multimodal context implied.
No compensation figures are provided in the posting.
Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): data pipeline development.
Caller marked this a fully-remote role — scored against the national candidate pool.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.