24-MAG · New York, NYremote

Salary
$55–$85/hrfrom the description
Posted
Jul 25, 2026
Location
New York, NY
Last confirmed open
Jul 26, 2026

What this job asks for AI summary

A full-time remote contract role focused on building and evaluating benchmarks for frontier AI systems. The work involves converting ML research ideas into structured multi-step tasks, implementing and running training experiments in Python, and critically assessing how AI models handle complex machine learning problems — including reinforcement learning scenarios. It suits practitioners with solid hands-on experience across the full ML experimentation lifecycle.

Mid level · 1+ years · Remote · Contract

Must have (5)
PythonGitML model trainingJupyter NotebooksLLMs
Nice to have (3)
reinforcement learningagentic AIablation studies

We read this from the posting text with AI. Skim the description below before ruling yourself out.

Why we read it this way (9)

The role is posted as a W-2 contingent (contract) engagement through staffing firm 24-MAG LLC, not direct employment — classified as Contract accordingly.

The minimum experience threshold is stated as 'at least 1 year' under the Ideal Profile section, which reads as aspirational/preferred framing ('strong candidates may have') rather than a hard gate; however it is the only experience figure given and is treated as the effective floor.

A master's or PhD is described as 'highly relevant' and equivalent practical experience is explicitly accepted — no hard degree requirement exists.

The role sits at the boundary between Data Science (ML research, model training, experimentation) and Software Development (implementing reference solutions, running experiments in code); 15-2051 is chosen because the primary work is ML research evaluation and analysis, with 15-1252 as a close runner-up.

ML model training and notebook environments are listed under the required profile ('working proficiency in Python and Git', 'hands-on experience training and evaluating ML models', 'comfort using both scripting and notebook-based environments') and are treated as hard gates. LLMs are similarly gated ('familiarity with large language model capabilities, limitations, and evaluation techniques').

Reinforcement learning, benchmark development, agentic AI, and ablation studies all appear under the 'Nice to Have' section and are marked preferred.

No specific geographic metro is given; the role is fully remote within the United States.

Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): benchmark development.

Posting is for a contract engagement — the market benchmarks below price full-time roles, so read the comp comparison with that in mind.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗