Member of Technical Staff - ML Training Infrastructure
Rethink recruit · San Francisco, CA
—
Jul 12, 2026
San Francisco, CA
Jul 21, 2026
What this job asks for AI summary
A first-of-its-kind infrastructure role focused on building and owning the full training compute stack for a robotics foundation model currently in pretraining. Day-to-day work spans GPU cluster management, distributed training performance, data pipeline construction, fault tolerance, and evaluation tooling across hundreds of GPUs. Best suited to engineers with hands-on experience running and debugging large-scale multi-node training jobs in production environments.
Senior level · National
“or” means any one of them counts — you don't need all of them.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
What the occupation pays Median $138,970 (middle half $107,524–$175,762).
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.
Why we read it this way (5)
No work location is stated in the posting. The company (Pantheon) appears to be a US-based robotics startup, so the US-based flag of true is assumed, but the specific metro is unknown.
This role sits at the boundary between software development (building training infrastructure, data loaders, evaluation harnesses, fault-tolerance systems) and infrastructure/systems operations (operating GPU clusters, storage, networking). The primary day-to-day work is building and shipping software systems, which tips it to Software Developers (15-1252), but 15-1244 is a credible runner-up given the heavy cluster-operations emphasis.
The required skills section names no specific frameworks or tools by name — it gates on experience with multi-node distributed training at 100+ GPUs, debugging NCCL, and operating production systems. Concrete tool names (Slurm, Kubernetes, Ray, WebDataset) appear only under 'Nice to Have'.
No compensation, employment type, or degree requirement is stated in the posting.
Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): GPU cluster operations.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.