Senior Data Engineer - Databricks at Steampunk
McLean, VA
$125,000–$165,000from the description
Jul 13, 2026
McLean, VA
Jul 21, 2026
What this job asks for AI summary
A senior-level data engineering role focused on designing and migrating enterprise data platforms, primarily on Databricks, using Python and SQL. Day-to-day work involves building batch and streaming pipelines within lakehouse architectures, supporting GenAI and ML workflows, and advising both technical and non-technical stakeholders on data solutions. The role suits an experienced engineer comfortable operating in regulated, government-adjacent cloud environments and Agile delivery teams.
Senior level · 8+ years · Remote · Bachelor's required · Full-time
Posted 2 times — it's one opening, so apply once.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
Roughly 320 people nationally plausibly meet what this posting asks for (database architects). range 190–480
Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.
Rare in this occupation — lead with these, and say what you built with them.
Most people in this occupation already list these. Still required — just not what gets you shortlisted.
What the occupation pays Median $142,568 (middle half $111,775–$173,013). This posting is about at that midpoint.
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.
Why we read it this way (8)
The role sits at the boundary between 15-1243 (Database Architects) and 15-1252 (Software Developers): the JD emphasises architecting lakehouse/data-warehouse systems and data pipelines, which leans toward 15-1243, but the volume of pipeline-building and distributed-processing work is also consistent with 15-1252.
Clearance: the JD requires the ability to hold a US government public-trust position, which is a suitability determination, not a formal security clearance — the clearance requirement is therefore set to None.
Degree requirement: the JD states '12+ years with a Bachelor's or 8+ years with a Master's', making a Bachelor's the minimum formal degree gate (the Master's path only reduces the experience threshold).
Total years minimum is set to 8, reflecting the lower bound of the stated range ('8-12 years industry experience in data engineering'); the 12-year figure applies to candidates without a graduate degree.
MLflow is the only explicitly named ML-lifecycle tool; GenAI/ML workflow experience (feature engineering, vector storage, RAG pipelines) is listed as a required qualification, with MLflow cited as the representative tool.
Search/retrieval patterns (indexing, semantic search, vector-based retrieval) are called out as required but name no specific tool, so no additional skill entry is emitted for that capability.
Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): Apache Spark Structured Streaming.
Caller marked this a fully-remote role — scored against the national candidate pool.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.