Staff/Principal DevOps Engineer, AI Inference at Lila Sciences
Cambridge, MA
$192,000–$272,000
Jul 29, 2026
Cambridge, MA
Jul 29, 2026
What this job asks for AI summary
A Staff/Principal-level infrastructure engineering role at an early-stage AI company, focused on designing and operating GPU/accelerator infrastructure for serving large-scale machine learning models in production. The work spans Kubernetes-based GPU scheduling, LLM inference serving platforms, autoscaling, CI/CD for model artifacts, and AWS cloud infrastructure — all oriented toward low-latency, high-throughput inference. Suits engineers with deep hands-on experience at the intersection of platform/SRE and ML infrastructure.
Staff level · Full-time
Pay in the description: $192,000–$272,000
“or” means any one of them counts — you don't need all of them.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
Rare in this occupation — lead with these, and say what you built with them.
Most people in this occupation already list these. Still required — just not what gets you shortlisted.
What the occupation pays Median $138,970 (middle half $107,524–$175,762). This posting is about at that midpoint.
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 29, 2026. It is a model, not a headcount.
Why we read it this way (8)
The job title presents this as 'Staff/Principal' — both bands are explicitly named. The requirements describe org-wide, cross-team technical leadership on a novel AI inference platform, which genuinely supports Staff-level classification; Principal was not chosen because the posting hedges with 'Staff/Principal' and the scope, while broad, is not clearly org-wide authority at a large organization.
Work location is not stated in the posting. The salary range is explicitly USD and U.S.-only, and benefits language distinguishes U.S. from international employees, so the role is treated as U.S.-based. No specific city or metro could be determined.
The SOC classification is Medium confidence: this role writes and ships infrastructure code (Terraform, Helm, CI/CD pipelines, Python tooling) — consistent with Software Developers (15-1252) — but has a strong operational/SRE flavor. 15-1244 is noted as a runner-up.
Model serving frameworks (vLLM, Triton Inference Server, TGI) are listed under 'What You'll Be Building' as stack context rather than a hard-gated requirements section; however, the requirements section does gate on 'model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks' as a capability, so vLLM/Triton/TGI are captured as required with alternatives reflecting the 'or' framing in the posting.
GPU infrastructure and NCCL are captured as required skills reflecting the hard-gated requirements language ('Deep experience with Kubernetes for ML workloads: GPU scheduling…' and 'Strong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL…'). These name specific technologies within required sections.
Quantization methods (GPTQ, AWQ, FP8), multi-accelerator families, Rust/Go, chaos engineering, and model registries all appear under 'Bonus Points For' and are marked preferred accordingly.
No minimum total years of experience is stated in the posting.
Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): chaos engineering.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.