New York City, NYremote

Salary
$140,000–$200,000from the description
Posted
Jul 21, 2026
Location
New York City, NY
Last confirmed open
Jul 22, 2026

What this job asks for AI summary

This role sits within a small team focused on converting acquired company assets — production codebases, databases, and workspaces — into AI training products such as reinforcement-learning environments, agentic task suites, evaluations, and fine-tuning datasets. The engineer will design and build the pipelines that transform raw assets into derivative works, and collaborate with lab and RL-environment buyers to align output with their training needs. A background in actually building RL environments or model-evaluation products is treated as a hard requirement.

Senior level · 4+ years · New York-Newark-Jersey City, NY-NJ-PA · Full-time

Must have (5)
RL environmentsPythonDockerCI/CDLLMs
Nice to have (3)
SWE-bench, Hud or VerifiersPlaywrightMCP

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

Roughly 1,750 people in the New York-Newark-Jersey City, NY-NJ-PA area plausibly meet what this posting asks for (software developers). range 910–2,250

Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.

What won't set you apart
Docker53%Python51%CI/CD45%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $170,499 (middle half $133,523–$210,499). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (7)

The role sits at the intersection of software engineering and ML/AI-training-data work. 15-1252 (Software Developers) was chosen because the primary day-to-day output is building pipelines, sandboxed environments, and tooling — shipping code — rather than modeling or statistical analysis. 15-2051 (Data Scientists) is a credible runner-up given the RL/eval/training-data focus.

RL-environments background is called out explicitly as 'the core requirement, not a nice-to-have'; it is captured as a required skill. Because it is a domain/background rather than a single named tool, it is labeled generically as 'RL environments'.

LLM evaluation and agent harnesses (SWE-bench-style setups, Verifiers, HUD) are listed under the required qualifications section ('Evals & verification') but framed with 'familiarity with' — a mild softener. They are captured as preferred to reflect that language, with SWE-bench as the primary name and HUD/Verifiers as alternatives.

MCP and Playwright appear only under the 'Nice to have' section and are marked preferred accordingly.

The 4–8 year range is stated as an experience band; the overall years minimum is set to 4 (the lower bound).

The degree requirement explicitly accepts 'equivalent practical experience,' so the degree requirement is set to None.

The role is hybrid (2 days/week in Midtown NYC); candidates must be in the NYC metro area, so remote is false.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗