Remote | MCP & AI Connector Evaluation Specialist — $45–$185/hour
24-MAG · New York, NY
$45–$185/hrfrom the description
Jul 25, 2026
New York, NY
Jul 26, 2026
What this job asks for AI summary
A part-time, fully remote contract role for experienced AI power users to support an LLM evaluation research initiative. The work involves designing realistic multi-step personal-assistant tasks, executing those workflows using MCP tools and connectors (such as Google Drive and Notion) with screen recording, and then assessing model outputs by writing structured scoring rubrics. The role suits candidates with deep, frequent hands-on experience using connected LLM tools for complex personal workflows.
Senior level · Remote · Contract
We read this from the posting text with AI. Skim the description below before ruling yourself out.
Why we read it this way (10)
This role is difficult to classify under a standard SOC code — it is an AI evaluation / human-data-labeling contractor role, not a conventional software or data occupation. 15-1299 (Computer Occupations, All Other) is the closest fit; 15-2031 (Operations Research Analysts) is a plausible alternative given the structured evaluation and rubric-development work.
No overall years-of-experience figure is stated. The '6 months or more of regular LLM usage history' is a usage-history signal, not a professional experience gate, so it is not reflected in the overall years minimum.
The degree requirement is explicitly secondary to hands-on experience ('Formal technical credentials are secondary to extensive hands-on experience'), so the degree requirement is set to None.
The wide hourly rate range ($45–$185/hr) reflects stated variation by 'expertise, task complexity, and project scope' — the actual rate for any individual engagement is unspecified at application time.
Google Drive and Notion appear in both the required workflow-testing section and the Nice-to-Have section; they are treated as required given their prominent placement in the core responsibilities.
AI evaluation, rubric design, user research, and quality assurance appear under 'Nice to Have' or as experience that 'may strengthen an application' — listed as preferred accordingly.
Screen recording is explicitly required ('Screen recording is required during task execution') and a desktop/laptop is a hard equipment gate; Chromebooks are explicitly excluded.
The posting is from a staffing/consulting intermediary (24-MAG LLC) and does not name the end client or the specific AI platform being evaluated.
Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): rubric design.
Posting is for a contract engagement — the market benchmarks below price full-time roles, so read the comp comparison with that in mind.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.