Senior Software Engineer, AI Benchmarking at Jobgether
—
Jul 20, 2026
—
Jul 22, 2026
What this job asks for AI summary
A senior engineering role focused on building and operating infrastructure for evaluating frontier AI models, with a particular emphasis on biosecurity and dual-use risk. Day-to-day work spans developing model adapters, scaling evaluation pipelines, contributing to cloud infrastructure, and analyzing results to inform safety research. Best suited to an experienced Python engineer with hands-on background in LLM-based applications or AI agent systems who has a genuine interest in AI safety.
Senior level · 5+ years · Remote · Full-time
“or” means any one of them counts — you don't need all of them.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
Roughly 20,400 people nationally plausibly meet what this posting asks for (software developers). range 12,200–36,600
Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.
Most people in this occupation already list these. Still required — just not what gets you shortlisted.
What the occupation pays Median $138,970 (middle half $107,524–$175,762).
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.
Why we read it this way (6)
The role is posted via Jobgether on behalf of an unnamed partner company; the actual employer is not disclosed, though office locations in Cambridge, MA and Berkeley, CA are mentioned as optional access points for this fully remote role.
No compensation figures are provided — only a 'competitive compensation package based on experience and qualifications' is mentioned.
The alt SOC 15-1299 (Computer Occupations, All Other) reflects the specialized AI evaluation/benchmarking focus, which sits at the boundary of software engineering and applied AI research; however, the primary day-to-day work is clearly software development.
AWS ECS is specifically called out in the requirements section as an example AWS service; captured under the broader AWS skill rather than as a separate entry since ECS is a sub-service rather than a separately marketed product in the same way as, e.g., AWS CDK.
Terraform, OpenTofu, and AWS CDK are listed together under preferred qualifications as interchangeable infrastructure-as-code options; emitted as one preferred skill with alternatives.
Inspect AI, OpenAI-compatible APIs, Anthropic APIs, Together APIs, Codex, and Claude Code all appear under preferred qualifications; grouped into two preferred skills reflecting their functional categories (evaluation framework vs. LLM API/coding tools).
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.