Applied Data Scientist, Finance AI Evaluation & Datasets at Innodata Inc.
$150,000–$175,000from the description
Jul 20, 2026
—
Jul 21, 2026
What this job asks for AI summary
This role centers on designing, measuring, and validating datasets used to train, fine-tune, and evaluate large language models and multimodal AI systems built for financial services workflows — think earnings analysis, fraud detection, credit underwriting, and compliance. The work spans annotation guidelines, evaluation methodology, statistical quality checks, and model risk documentation, with particular emphasis on unstructured and multimodal financial data such as PDFs, scanned documents, tables, and call transcripts. It suits a data scientist with hands-on financial-domain experience who is comfortable bridging rigorous ML dataset practices with regulated-industry governance requirements.
Senior level · 5+ years · Remote · Full-time
“or” means any one of them counts — you don't need all of them.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
Roughly 1,700 people nationally plausibly meet what this posting asks for (data scientists). range 1,000–2,500
Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.
Rare in this occupation — lead with these, and say what you built with them.
Most people in this occupation already list these. Still required — just not what gets you shortlisted.
What the occupation pays Median $122,874 (middle half $87,544–$162,374). This posting is about at that midpoint.
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.
Why we read it this way (9)
The JD requires 5+ years of data science experience overall, with at least 2+ years specifically in financial services, fintech, banking, or a comparable regulated environment — the 5-year figure is used for the overall years minimum as the broader gate.
pandas and scikit-learn are listed together as 'pandas, scikit-learn, or equivalent' in the requirements block; they are treated as interchangeable alternatives on a single skill rather than two separate hard gates.
Hugging Face and PyTorch appear in the same requirements-block clause ('working familiarity with Hugging Face, PyTorch, or model APIs'); 'working familiarity' is softer than the firm language used for Python/SQL, so this is marked preferred.
XBRL, ISO 20022, and GAAP/IFRS are explicitly flagged as 'strongly preferred' in the posting, not required — marked preferred accordingly.
Weights & Biases and LangFuse appear under an 'Experience with … tools such as' clause in what reads as a secondary/preferred qualifications block — marked preferred.
OCR/document AI experience appears in a secondary preferred-qualifications block ('Experience with document AI, OCR/post-OCR quality…') — marked preferred.
The title 'Applied Data Scientist' carries no explicit seniority level word, so the title states no level; the 5+ year requirement and end-to-end ownership scope support a Senior classification.
Degree requirement is None: the JD accepts 'equivalent demonstrated experience' in lieu of a formal degree.
Caller marked this a fully-remote role — scored against the national candidate pool.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.