Innodata Inc.remote

Salary
$100,000–$120,000from the description
Posted
Jul 22, 2026
Location
Last confirmed open
Jul 23, 2026

What this job asks for AI summary

A data engineering role focused on building and maintaining data warehouses, data lakes, and ETL pipelines to support supply chain and real estate operations. The work spans GCP services such as BigQuery and Dataflow, SQL and Python scripting, and extending pipelines to support AI/ML use cases including retrieval-augmented generation and agentic workflows. Suits engineers with solid ETL and cloud data infrastructure experience.

Mid level · Remote · Full-time

Must have (8)
SQLPythonETL/ELT pipelinesGCP or AWSBigQueryDataflowPub/SubCloud Storage
Nice to have (6)
LookerRAGvector databasesTableauPower BIMLOps

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

Roughly 1,500 people nationally plausibly meet what this posting asks for (database architects). range 300–2,200

Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.

What won't set you apart
SQL90%Python55%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $142,568 (middle half $111,775–$173,013). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (8)

No work location or metro is specified in the posting; the role appears to be remote or location-flexible within the US.

The SOC classification is a close call: the role centers on designing data warehouses, lakes, and pipelines (15-1243 Database Architects), but the heavy Python scripting, ETL code authorship, and AI/ML pipeline development also support 15-1252 Software Developers.

GCP services (BigQuery, Dataflow, Pub/Sub, Cloud Storage, Looker) are listed under 'You'll Thrive in This Role If You Have' with the qualifier 'or similar, preferred', but the role's core responsibilities explicitly require building on GCP — they are treated as hard gates on the GCP capability, with specific services marked required accordingly.

Looker is listed with the 'or similar, preferred' qualifier in the requirements section and is also framed as a Bonus item; it is marked preferred to reflect that ambiguity.

RAG, vector databases, copilots, and agentic AI pipelines appear under 'Exposure to data pipelines for AI/ML' — framed as familiarity rather than deep expertise — and are marked preferred.

Tableau and Power BI appear only in the Bonus section and are marked preferred.

Supply chain or data center domain knowledge is explicitly called 'a strong plus' and is a domain concept rather than a named tool, so it is omitted from the skills list per extraction rules.

No minimum total years of experience is stated in the posting.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗