Data Engineer at Matterworks, Inc.
Somerville, MA
—
Aug 4, 2026
Somerville, MA
Sep 24, 2026
What this job asks for AI summary
A data engineering role at an early-stage biotech AI startup focused on mass spectrometry and molecular data. The engineer will design, build, and operate production data pipelines that transform raw scientific data into training corpora for ML models and queryable datasets for a product layer. The role suits someone comfortable owning pipelines end-to-end — including quality checks, monitoring, and messy vendor formats — in a small, cross-disciplinary team.
Mid level · 2+ years · Remote · Full-time
Quick apply — this platform usually takes a CV and a few fields.
“or” means any one of them counts — you don't need all of them.
Posted 2 times — it's one opening, so apply once.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
Roughly 7,800 people nationally plausibly meet what this posting asks for (database architects). range 4,700–11,700
Applicant volume Heavy — This req sits in a large pool with little in its requirements to thin it, and auto-apply tools fire at everything in the occupation. Applying early and leading with the rare skills below is what gets read.
Most people in this occupation already list these. Still required — just not what gets you shortlisted.
What the occupation pays Median $142,568 (middle half $111,775–$173,013).
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Aug 6, 2026. It is a model, not a headcount.
Why we read it this way (5)
The posting names a specific stack (Argo Workflows, Metaflow, EKS, Glue, Athena, Apache Iceberg, Parquet, DuckDB, Terraform) but explicitly states 'depth in any comparable stack transfers fine' — these are listed as context for the team's environment, not hard gates. All are marked preferred accordingly.
Cloud infrastructure knowledge is required ('working knowledge of cloud data infrastructure'), but no specific cloud provider is mandated; AWS is inferred from the named stack (Glue, Athena, EKS) and listed as preferred.
The role is classified as a data pipeline/architecture role (15-1243) rather than general software development (15-1252) given the primary focus on ETL, data quality, ingestion, and dataset ownership — though the boundary is close given the engineering depth expected.
No compensation figures are disclosed; the posting references a 'competitive base salary' without quoting numbers.
The role is described as hybrid-flexible, explicitly accommodating fully remote team members, so remote is set to true.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.