fal · San Francisco, CA

Salary
$180,000–$250,000from the description
Posted
Jul 11, 2026
Location
San Francisco, CA
Last confirmed open
Jul 20, 2026

What this job asks for AI summary

Senior level · San Francisco-Oakland-Berkeley, CA · Full-time

Must have (7)
PyTorchTensorRTTransformerEngineNsightCUTLASSTritonmodel quantization
Nice to have (5)
Ring AttentionFlashAttention 3FusedMLPtensor parallelismsequence parallelism

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

What gives you an edge
CUTLASS1%Triton2%PyTorch9%

Rare in this occupation — lead with these, and say what you built with them.

What the occupation pays Median $190,744 (middle half $167,095–$224,501). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (6)

This role sits at the intersection of ML infrastructure engineering and low-level GPU/accelerator programming; 15-1252 (Software Developers) was chosen as the primary code because the role's core output is building and shipping inference engine software and tooling. 15-1299 (Computer Occupations, All Other) is a reasonable runner-up given the highly specialized ML systems focus.

Triton is listed as 'Proficient in Triton or willingness to learn with comparable experience in lower-level accelerator programming' — the JD is in the Requirements section, so it is treated as a hard gate on the underlying GPU programming capability, with Triton as the named tool.

Model quantization and model compilation are called out explicitly in the Requirements section as part of the ML infrastructure stack; they are retained as named technical requirements rather than generic concepts because the JD names them as specific competency areas alongside concrete tools.

Ring Attention, FA3 (FlashAttention 3), FusedMLP, tensor parallelism, and sequence parallelism appear under the 'New frontier' and 'Familiar with internals of' framing within the Requirements block, but the language ('Familiar with', 'New frontier') reads as aspirational/nice-to-have depth rather than a hard rejection gate — treated as preferred.

No overall years-of-experience requirement is stated in the posting.

The role is on-site in downtown San Francisco; relocation assistance is offered but there is no remote option mentioned.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗