remote

Salary
$100,000–$150,000from the description
Posted
Jul 20, 2026
Location
Last confirmed open
Jul 22, 2026

What this job asks for AI summary

A performance-focused engineering role centered on profiling and optimizing large-scale AI training and inference systems. Day-to-day work spans GPU-level tuning, distributed training strategies, model compression, compiler optimizations, and LLM serving improvements, alongside building benchmarking frameworks and cost-reduction initiatives. Suited to someone with deep machine learning systems or high-performance computing experience who can operate across the full AI infrastructure stack.

Senior level · 6+ years · Remote · Full-time

Must have (6)
PythonC++distributed trainingprofiling toolsquantizationFSDP, Zero, Tensor Parallelism or Pipeline Parallelism
Nice to have (4)
Triton, Xla, Torchinductor or TvmvLLM, Tensorrt Llm or DeepspeedCUTLASSFinOps

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

Roughly 1,850 people nationally plausibly meet what this posting asks for (software developers). range 380–3,350

Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.

What gives you an edge
FSDP2%

Rare in this occupation — lead with these, and say what you built with them.

What won't set you apart
Python51%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $138,970 (middle half $107,524–$175,762). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (8)

The role is posted via Jobgether on behalf of an unnamed partner company; the actual employer is not disclosed.

The posting is fully remote within the Continental United States; no specific metro or state is given.

Degree requirement states 'Bachelor's or Master's degree' but does not accept equivalent experience as a substitute in the same sentence — however, the phrasing is aspirational ('ideal profile') rather than a hard gate, so degree is treated as preferred rather than required.

FSDP, ZeRO-style sharding, tensor parallelism, and pipeline parallelism are all listed together as distributed training approaches; FSDP is used as the primary skill name with the others as alternatives.

Compiler-level optimization tools (Triton, XLA, TorchInductor, TVM) appear in the accountabilities/responsibilities section rather than a formal requirements block; they are marked preferred accordingly.

vLLM, TensorRT-LLM, and DeepSpeed appear under Preferred Qualifications and are marked preferred.

SOC classification is Medium confidence: the role is primarily performance/systems engineering on ML infrastructure (shipping optimization tooling and benchmarking frameworks), which maps best to Software Developers (15-1252); Data Scientists (15-2051) is the runner-up given the heavy ML modeling context.

Ignored 2 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): GPU optimization, KV cache optimization.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗