Austin, TX

Salary
—
Posted
Aug 19, 2026
Location
Austin, TX
Last confirmed open
Sep 25, 2026

What this job asks for AI summary

This Staff-level role is the single technical owner of Slurm cluster architecture, multi-tenant scheduling policy, and GPU cluster reliability across bare-metal and VM-based infrastructure at a Bitcoin mining and AI compute company. The engineer leads adoption of the Slinky operator stack to co-schedule Slurm and Kubernetes workloads on shared GPU pools, builds health-check and observability systems, and serves as the escalation point and customer-facing expert for enterprise HPC onboarding.

Staff level · 8+ years · Full-time

Quick apply — this platform usually takes a CV and a few fields.

Must have (13)
Slurm · 4+ yrsKubernetesInfiniBand or Rocev2NVIDIA DCGMNCCLTerraformAnsiblePythonBashPrometheusMIGGPUDirect RDMASlinky slurm-operator, Coreweave Sunk or Nebius Soperator
Nice to have (11)
PyxisEnrootLustre, Gpfs, Weka, Vast or NfsKVM or QemuGocert-managerHelmMPIPMIxSpackLBNL NHC

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

What won't set you apart
Bash70%Python51%Ansible45%Prometheus45%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $101,310 (middle half $79,725–$129,425).

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Aug 20, 2026. It is a model, not a headcount.

Why we read it this way (8)

The SOC classification is a genuine judgment call. The role is primarily about operating, architecting, and owning production HPC/Slurm cluster infrastructure (15-1244), but it also involves writing automation code in Python/Bash/Go and building tooling (a signal toward 15-1252). 15-1244 was chosen because the dominant day-to-day work is running and owning infrastructure rather than shipping a software product.

No work location is specified beyond the company's Singapore headquarters. The posting lists US data centers but does not state the role is US-based; US-based flag is set to false accordingly. Candidates should confirm the primary work location before applying.

The 4-year Slurm experience requirement is explicitly scoped to production clusters at 100+ GPU-node scale; the 8-year overall figure covers HPC/systems/cloud infrastructure broadly.

Slinky slurm-bridge is listed as an evaluation/pilot responsibility rather than a hard gate, so it is captured under the Slinky slurm-operator skill entry rather than as a separate required skill.

Pyxis, Enroot, LBNL NHC, cert-manager, Helm, MPI/PMIx, and Spack appear in the responsibilities narrative rather than under a formal requirements heading and are marked preferred.

Go is explicitly called 'a plus' in the qualifications and is marked preferred.

Parallel/shared storage (Lustre, GPFS/Spectrum Scale, WEKA, VAST, NFS) is described as 'working knowledge' — a softer gate — and is marked preferred with alternatives.

No compensation range is stated in the posting.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗