remote

Salary
$100,000–$150,000from the description
Posted
Jul 16, 2026
Location
Last confirmed open
Jul 21, 2026

What this job asks for AI summary

A full-stack reinforcement learning role covering the entire lifecycle of RL systems — from designing algorithms and reward functions to building distributed training infrastructure and deploying production-ready models. The work spans simulation environments, policy optimization, safety mechanisms, and evaluation frameworks. Suited to someone with substantial RL research and engineering experience who can bridge theoretical depth with practical, large-scale delivery.

Senior level · 6+ years · Remote · Full-time

Must have (6)
Pythonreinforcement learningdeep learning frameworks, PyTorch, TensorFlow or Jaxsimulation environmentsdistributed trainingGPU clusters
Nice to have (5)
RLHF or Dpohierarchical RLrobotics or Autonomous Systemsoffline RLimitation learning

“or” means any one of them counts — you don't need all of them.

Posted 2 times — it's one opening, so apply once.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

Roughly 12,300 people nationally plausibly meet what this posting asks for (data scientists). range 2,550–18,400

Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.

What won't set you apart
Python88%deep learning frameworks40%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $122,874 (middle half $87,544–$162,374). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (7)

The role sits at the intersection of ML research and software engineering; 15-2051 (Data Scientists) was chosen as the primary SOC because the core work is RL algorithm research, reward modeling, and policy optimization — statistical/modeling work — but 15-1252 (Software Developers) is a genuine runner-up given the strong emphasis on production infrastructure, distributed training pipelines, and deployment.

The JD requires a Master's or PhD 'or equivalent practical experience', so no formal degree is hard-gated; degree requirement is set to None.

Deep learning frameworks are required but no specific framework is named; PyTorch, TensorFlow, and JAX are listed as the well-known interchangeable alternatives.

RLHF/DPO and multi-agent/hierarchical RL, robotics, and offline RL/imitation learning all appear under explicitly preferred or 'is a plus' language in the requirements section and are marked preferred accordingly.

The title carries no seniority level word; the title states no level. The 6+ years requirement and full lifecycle ownership scope support a Senior classification.

Ignored 2 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): reward function design, multi-agent reinforcement learning.

Caller marked this a fully-remote role — scored against the national candidate pool.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗