Austin, TX

Salary
—
Posted
Aug 19, 2026
Location
Austin, TX
Last confirmed open
Sep 25, 2026

What this job asks for AI summary

This Staff-level role owns the workload scheduling and orchestration layer of an AI-native cloud platform, with a focus on eliminating GPU stranding and maximizing utilization across large GPU fleets. The engineer will design and build advanced batch scheduling systems, topology-aware pod placement, GPU sharing mechanisms, and admission control pipelines on top of Kubernetes. It suits a distributed-systems specialist with deep Kubernetes internals knowledge and hands-on experience running HPC or large-scale AI training environments.

Staff level · 6+ years · Bachelor's required · Full-time

Quick apply — this platform usually takes a CV and a few fields.

Must have (9)
KubernetesVolcano or YunikornKueuePyTorch Distributed, Ray or MpiNVLinkInfiniBandGPU MIGTerraformGo

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

What gives you an edge
Terraform11%Go14%

Rare in this occupation — lead with these, and say what you built with them.

What the occupation pays Median $138,970 (middle half $107,524–$175,762).

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Aug 20, 2026. It is a model, not a headcount.

Why we read it this way (8)

The posting is headquartered in Singapore and lists data centers across multiple countries; no specific city or US-based work location is stated, so this role is treated as non-US and no CBSA is assigned.

The degree requirement lists 'Bachelor's or Master's' — Bachelor's is captured as the minimum hard gate.

Volcano and YuniKorn appear together as interchangeable batch scheduling framework options ('like Volcano or YuniKorn'); they are combined into one skill with YuniKorn as an alternative.

PyTorch Distributed, Ray, and MPI are listed as interchangeable distributed training framework examples ('e.g., PyTorch Distributed, Ray, MPI'); Ray and MPI are captured as alternatives.

Terraform and Go-based Operators are listed together as infrastructure-as-code examples; Terraform is named as the primary skill and Go is extracted separately as it is a distinct language used for writing Kubernetes Operators.

GPU MIG and time-slicing are listed together as GPU sharing technologies; MIG is the named product and is captured as the primary skill.

The alt SOC (15-1244) is noted because the role has a significant infrastructure operations dimension, but the primary work — writing custom scheduler plugins, operators, and orchestration code — firmly points to Software Developers (15-1252).

Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): Kubernetes Dynamic Resource Allocation.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗