Somerville, MAremote

Salary
—
Posted
Aug 4, 2026
Location
Somerville, MA
Last confirmed open
Sep 24, 2026

What this job asks for AI summary

A data engineering role at an early-stage biotech AI startup focused on mass spectrometry and molecular data. The engineer will design, build, and operate production data pipelines that transform raw scientific data into training corpora for ML models and queryable datasets for a product layer. The role suits someone comfortable owning pipelines end-to-end — including quality checks, monitoring, and messy vendor formats — in a small, cross-disciplinary team.

Mid level · 2+ years · Remote · Full-time

Quick apply — this platform usually takes a CV and a few fields.

Must have (2)
PythonSQL
Nice to have (7)
Argo Workflows, Metaflow, Airflow or PrefectKubernetesAWS Glue, GCP or AzureApache IcebergParquetDuckDBTerraform

“or” means any one of them counts — you don't need all of them.

Posted 2 times — it's one opening, so apply once.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

Roughly 7,800 people nationally plausibly meet what this posting asks for (database architects). range 4,700–11,700

Applicant volume Heavy — This req sits in a large pool with little in its requirements to thin it, and auto-apply tools fire at everything in the occupation. Applying early and leading with the rare skills below is what gets read.

What won't set you apart
SQL90%Python55%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $142,568 (middle half $111,775–$173,013).

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Aug 6, 2026. It is a model, not a headcount.

Why we read it this way (5)

The posting names a specific stack (Argo Workflows, Metaflow, EKS, Glue, Athena, Apache Iceberg, Parquet, DuckDB, Terraform) but explicitly states 'depth in any comparable stack transfers fine' — these are listed as context for the team's environment, not hard gates. All are marked preferred accordingly.

Cloud infrastructure knowledge is required ('working knowledge of cloud data infrastructure'), but no specific cloud provider is mandated; AWS is inferred from the named stack (Glue, Athena, EKS) and listed as preferred.

The role is classified as a data pipeline/architecture role (15-1243) rather than general software development (15-1252) given the primary focus on ETL, data quality, ingestion, and dataset ownership — though the boundary is close given the engineering depth expected.

No compensation figures are disclosed; the posting references a 'competitive base salary' without quoting numbers.

The role is described as hybrid-flexible, explicitly accommodating fully remote team members, so remote is set to true.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗