Data Engineer
Innodata Inc.
$100,000–$120,000from the description
Jul 22, 2026
—
Jul 23, 2026
What this job asks for AI summary
A data engineering role focused on building and maintaining data warehouses, data lakes, and ETL pipelines to support supply chain and real estate operations. The work spans GCP services such as BigQuery and Dataflow, SQL and Python scripting, and extending pipelines to support AI/ML use cases including retrieval-augmented generation and agentic workflows. Suits engineers with solid ETL and cloud data infrastructure experience.
Mid level · Remote · Full-time
“or” means any one of them counts — you don't need all of them.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
Roughly 1,500 people nationally plausibly meet what this posting asks for (database architects). range 300–2,200
Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.
Most people in this occupation already list these. Still required — just not what gets you shortlisted.
What the occupation pays Median $142,568 (middle half $111,775–$173,013). This posting is about at that midpoint.
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.
Why we read it this way (8)
No work location or metro is specified in the posting; the role appears to be remote or location-flexible within the US.
The SOC classification is a close call: the role centers on designing data warehouses, lakes, and pipelines (15-1243 Database Architects), but the heavy Python scripting, ETL code authorship, and AI/ML pipeline development also support 15-1252 Software Developers.
GCP services (BigQuery, Dataflow, Pub/Sub, Cloud Storage, Looker) are listed under 'You'll Thrive in This Role If You Have' with the qualifier 'or similar, preferred', but the role's core responsibilities explicitly require building on GCP — they are treated as hard gates on the GCP capability, with specific services marked required accordingly.
Looker is listed with the 'or similar, preferred' qualifier in the requirements section and is also framed as a Bonus item; it is marked preferred to reflect that ambiguity.
RAG, vector databases, copilots, and agentic AI pipelines appear under 'Exposure to data pipelines for AI/ML' — framed as familiarity rather than deep expertise — and are marked preferred.
Tableau and Power BI appear only in the Bonus section and are marked preferred.
Supply chain or data center domain knowledge is explicitly called 'a strong plus' and is a domain concept rather than a named tool, so it is omitted from the skills list per extraction rules.
No minimum total years of experience is stated in the posting.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.