Machine Learning Infrastructure Tech Lead at Reducto
San Francisco, CA
$200,000–$300,000
Jul 15, 2026
San Francisco, CA
Jul 21, 2026
What this job asks for AI summary
A hands-on technical lead role focused on building and optimizing the infrastructure that powers model training and inference at scale. The bulk of the work involves low-level performance tuning — kernels, GPU utilization, batching, scheduling, and distributed multi-node systems — alongside setting the architectural roadmap and tooling that helps ML engineers move from experimentation to production. Suited to someone with deep ML systems experience who is equally comfortable leading technical direction and writing the hardest code themselves.
Senior level · 5+ years · San Francisco-Oakland-Berkeley, CA · Full-time
“or” means any one of them counts — you don't need all of them.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
Roughly 1,550 people in the San Francisco-Oakland-Berkeley, CA area plausibly meet what this posting asks for (software developers). range 820–2,050
Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.
Most people in this occupation already list these. Still required — just not what gets you shortlisted.
What the occupation pays Median $190,744 (middle half $167,095–$224,501).
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.
Why we read it this way (6)
The title 'ML Infrastructure Tech Lead' carries no standard seniority level word, so the title states no level; the 5+ years requirement and scope (owning a roadmap, leading complex projects end-to-end) support a Senior classification.
SOC confidence is Medium: the role is primarily about building and shipping ML infrastructure software (model-serving kernels, training stacks, tooling), which maps to Software Developers (15-1252), but the heavy infrastructure/operations flavor (GPU utilization, distributed systems, Kubernetes) makes 15-1244 a plausible runner-up.
GPU training/inference optimization is listed as a required skill via firm language ('Understand the performance characteristics of modern GPU training or inference workloads') in the requirements section; it is not a named tool but is the closest concrete technical gate the JD states — retained as a skill because it describes a specific, measurable technical capability rather than a generic soft skill.
CUDA, Triton, vLLM, SGLang, PyTorch, TensorRT-LLM, Ray, and observability/scheduling/capacity-management systems all appear exclusively under the 'Bonus Points' section and are therefore preferred, not required.
No compensation figures are disclosed in the posting.
Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): GPU training/inference optimization.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.