Palo Alto, CA

Salary
$140,000–$240,000from the description
Posted
Jul 16, 2026
Location
Palo Alto, CA
Last confirmed open
Jul 21, 2026

What this job asks for AI summary

An ML infrastructure role focused on keeping AI-driven weather forecasting systems running reliably around the clock. The work spans building and maintaining data ingestion pipelines from diverse sources, ensuring uptime for real-time model inference, stabilizing distributed training runs, and managing compute resources across on-premises and cloud environments. Best suited to someone with hands-on production ML systems experience who can bring structure to a fast-moving research team.

Senior level · San Francisco-Oakland-Berkeley, CA · Full-time

Must have (7)
ML systems (production)PyTorchDockerdistributed trainingdata pipelineshealth monitoring / loggingcloud infrastructure
Nice to have (3)
weather datageospatial pipelinesjob schedulers

Posted 2 times — it's one opening, so apply once.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

Roughly 650 people in the San Francisco-Oakland-Berkeley, CA area plausibly meet what this posting asks for (software developers). range 190–980

Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.

What gives you an edge
PyTorch9%ML systems (production)12%

Rare in this occupation — lead with these, and say what you built with them.

What won't set you apart
Docker53%cloud infrastructure50%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $190,744 (middle half $167,095–$224,501). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (7)

The role sits at the intersection of ML infrastructure engineering and systems/ops work — it primarily involves building and shipping production pipelines, monitoring, and training infrastructure (pointing to Software Developers), but has a meaningful operational dimension (on-prem cluster management, incident response). The alt SOC 15-1244 reflects that operational thread.

No explicit years-of-experience requirement is stated; seniority is inferred from the scope (end-to-end ownership of production ML systems, compute strategy, multi-source data pipelines) and the $140k–$240k salary band.

The job title is not given in the posting; 'Unspecified' is used for advertised seniority accordingly.

Location is Redwood City, CA (1600 Bridge Pkwy); the nearest major CBSA is San Francisco-Oakland-Berkeley, CA (41860). The role is hybrid or in-person — not fully remote.

'Cursed memory management' and 'debugging network saturation' are idiomatic descriptions of systems debugging skills tied to PyTorch/distributed training contexts, not distinct named tools; they are captured under the PyTorch and distributed training skills rather than emitted as separate entries.

Weather data, geospatial pipelines, scientific computing, very large (petabyte-scale) datasets, GPU cluster management, and job schedulers all appear under the explicit 'Nice to haves' section.

Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): GPU cluster management.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗