remote

Salary
$125,000–$165,000from the description
Posted
Jul 26, 2026
Location
Last confirmed open
Jul 28, 2026

What this job asks for AI summary

An MLOps Engineer role at an AI startup focused on building and scaling inference infrastructure for generative audio models — TTS, voice conversion, and ASR. The work centers on designing high-performance, low-latency model serving systems for both streaming and batch workloads, managing Kubernetes-based deployments with autoscaling, and building CI/CD pipelines that bridge research and production. Suits an engineer with strong MLOps and GPU inference experience.

Senior level · Full-time

Quick apply — this platform usually takes a CV and a few fields.

Must have (4)
KubernetesCI/CDPython or GoGPU-accelerated inference
Nice to have (1)
Triton Inference Server or Vllm

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

What won't set you apart
Python51%CI/CD45%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $138,970 (middle half $107,524–$175,762). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (4)

The job title is 'MLOps Engineer' but the role's primary day-to-day work is writing and shipping infrastructure code (CI/CD pipelines, inference engines, autoscaling systems in Python/Go), which maps more closely to Software Developers (15-1252) than to systems administration (15-1244). The runner-up is 15-1244 given the heavy infrastructure operations component.

The compensation range is quoted in both USD ($125,000–$165,000) and EUR (€110,000–€145,000), suggesting the role may be open to US and European candidates. The USD figure is used here as the primary range given the US benefits section. Location is not explicitly stated in the posting.

Python and Go are presented as alternatives ('Python or Go') for infrastructure tooling and backend services — captured as a single skill with Go as the alternative.

Triton Inference Server and vLLM-Omni appear under a 'is a plus' qualifier alongside other high-performance inference engines, so they are marked preferred. vLLM-Omni is captured under the canonical 'vLLM' name.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗