Somerville, MAremote

Salary
—
Posted
Aug 4, 2026
Location
Somerville, MA
Last confirmed open
Sep 24, 2026

What this job asks for AI summary

A Senior Software Engineer role at an early-stage biotech AI startup, focused on designing and scaling a data platform that transforms raw mass spectrometry outputs into enriched, production-ready datasets for ML research and product use. The work spans pipeline architecture, data lake infrastructure, quality gating, and developer-facing SDKs — all at petabyte scale. Well-suited to engineers with deep experience in data systems who are comfortable operating autonomously in a scientific domain.

Senior level · Remote · Full-time

Quick apply — this platform usually takes a CV and a few fields.

Must have (11)
PythonSQLKubernetesArgo Workflows, Airflow, Dagster or MetaflowEKSAWS GlueApache IcebergParquetDuckDBTerraformLLMs
Nice to have (2)
RDKit or ProteowizardmzML

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

Roughly 1,250 people nationally plausibly meet what this posting asks for (database architects). range 370–2,300

Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.

What won't set you apart
SQL90%Python55%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $142,568 (middle half $111,775–$173,013).

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Aug 6, 2026. It is a model, not a headcount.

Why we read it this way (6)

No explicit years-of-experience number is stated; the posting says the company 'levels on scope and judgment rather than years.'

The role is described as flexible hybrid — fully remote is explicitly accommodated, though the office is in Somerville, MA.

Argo Workflows and Metaflow are listed together in the requirements block as the primary orchestration tools; Airflow and Dagster are called out as transferable alternatives. All four are captured under one skill entry with alternatives.

LLM/agent production experience (including eval loops and cost-per-record) is stated in the requirements block and treated as a hard gate, even though it is a relatively novel requirement for a data engineering role.

RDKit and ProteoWizard (along with mzML) appear in a 'comfort with messy scientific formats' line that reads as contextual/preferred rather than a strict gate — listed as preferred accordingly.

SOC classification is a genuine toss-up: the role builds data pipelines and lake architecture (15-1243 Database Architects) but also ships SDKs and production software systems (15-1252 Software Developers). 15-1243 was chosen as primary given the dominant focus on data contracts, enrichment pipelines, and lake-scale storage design.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗