San Francisco, CA

Salary
$150,000–$250,000
Posted
Jul 23, 2026
Location
San Francisco, CA
Last confirmed open
Jul 25, 2026

What this job asks for AI summary

A systems engineering role focused on building and optimizing a high-throughput video dataloader that feeds multi-petabyte multimodal datasets directly to GPUs at line rate. Day-to-day work involves profiling and tuning the full data path from object storage through NVMe, memory, and CPU to the GPU, eliminating bottlenecks and supporting complex sampling patterns without sacrificing training efficiency. Best suited to someone with deep low-level systems expertise in Rust, C++, or C and a strong grasp of OS internals, I/O, and memory hierarchies.

Senior level · San Francisco-Oakland-Berkeley, CA · Full-time

Must have (5)
Rust, C++ or Coperating systems internalsNVMeio_uringNUMA
Nice to have (7)
SLURM or KubernetesCUDAPyAV, Decord or NvdecPyTorchGPU-Direct StoragehugepagesDALI

“or” means any one of them counts — you don't need all of them.

Posted 2 times — it's one opening, so apply once.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

What the occupation pays Median $190,744 (middle half $167,095–$224,501).

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (8)

The role sits squarely in systems software — building a high-performance dataloader in Rust/C/C++ — so 15-1252 Software Developers is the primary classification, though the deep OS/hardware operations focus gives 15-1244 a plausible secondary claim.

Rust, C++, and C are presented as interchangeable options for the same hard gate ('Live and breathe Rust, C++, or C'); Rust is used as the primary name with the others in alternatives.

io_uring is called out with firm gating language ('Strong opinions on io_uring — love it or hate it, you've earned the opinion'), treating it as a required depth signal rather than a nice-to-have.

Operating systems internals (page cache, scheduling, syscalls, NUMA, memory hierarchies) is a required conceptual gate; NUMA is also captured as a discrete named technology because the JD references it specifically in both required and nice-to-have contexts.

GPU experience is explicitly stated as NOT required on day one; CUDA, SLURM/Kubernetes, GPU-Direct Storage, hugepages, video decode libraries (PyAV/decord/NVDEC), and PyTorch DataLoader internals all appear under the 'Nice to have' section.

PyAV, decord, and NVDEC are listed as interchangeable video decode pipeline options under Nice to have; PyAV is used as the primary name.

No compensation figures are disclosed; the posting references 'competitive comp and meaningful startup equity' only.

The role requires 4 days/week in-person at the SF Mission district office — it is not fully remote.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗