24-MAG · New York, NYremote

Salary
$55–$85/hrfrom the description
Posted
Jul 25, 2026
Location
New York, NY
Last confirmed open
Jul 26, 2026

What this job asks for AI summary

A full-time remote contract role focused on maintaining the quality and integrity of AI evaluation benchmarks. Day-to-day work involves reviewing multi-step benchmark tasks and reference solutions, designing test cases, debugging Python-based environments, detecting unintended shortcuts in AI agent runs, and building repeatable quality-assurance processes. Suited to test or software engineers with solid Python and debugging experience who are comfortable working independently on open-ended technical problems.

Mid level · 1+ years · Remote · Contract

Must have (2)
PythonGit
Nice to have (5)
AI model evaluationCI/CDcontainersautomated test suitesmachine learning

We read this from the posting text with AI. Skim the description below before ruling yourself out.

Why we read it this way (7)

The posting is structured as a consulting engagement through 24-MAG LLC but is explicitly described as a 'full-time W-2 contingent employment opportunity' — classified as Contract accordingly.

The minimum experience threshold is stated as 'at least 1 year' across QA, test engineering, software engineering, or a related technical role — a low bar that anchors the role at Mid rather than Senior, despite the technically demanding subject matter.

A master's degree or PhD in a STEM field is described as 'highly relevant' and equivalent practical experience 'may also be considered' — no degree is hard-gated, so degree requirement is set to None.

Python and Git are listed under the 'Ideal Profile' section with firm proficiency language ('working proficiency in Python and Git') and are the only two named technologies treated as hard requirements; all other named technologies appear under 'Nice to Have'.

The 'Nice to Have' section names: AI training/model evaluation, agentic systems/multi-step AI benchmarks, ML/research/data-processing workflow testing, automated test suites or validation scripts, CI/CD systems/test harnesses/containers/reproducible environments, and adversarial testing/failure-mode analysis/benchmark design — all captured as preferred.

No specific location or metro is stated; the role is fully remote within the United States.

Posting is for a contract engagement — the market benchmarks below price full-time roles, so read the comp comparison with that in mind.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗