NVIDIA Corporation · Santa Clara, CAremote

Salary
$184,000–$287,500from the description
Posted
Jun 30, 2026
Location
Santa Clara, CA
Last confirmed open
Jul 21, 2026

What this job asks for AI summary

This role centers on building and maintaining Topograph, an open-source, Kubernetes-native system that collects, aggregates, and normalizes cluster topology data from multiple sources to inform GPU infrastructure provisioning and workload scheduling across cloud providers. The work involves hands-on integration with cutting-edge NVIDIA hardware and networking fabrics to optimize GPU-to-GPU communication at scale. It suits a seasoned systems engineer with deep distributed systems and Kubernetes experience, ideally with a background in HPC or GPU cluster environments.

Senior level · 8+ years · Remote · Full-time

Must have (8)
Go, C++ or RustKubernetesLinuxdistributed systemsSlurm or SlinkycontainersRESTCI/CD
Nice to have (6)
GPU clustersNVLinkInfiniBandKubernetes schedulingDRAKueue

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

Roughly 23,800 people nationally plausibly meet what this posting asks for (software developers). range 10,000–35,800

Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.

What won't set you apart
Linux45%CI/CD45%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $138,970 (middle half $107,524–$175,762). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (8)

The JD requires a Bachelor's degree in CS or a related field 'or equivalent experience' — treated as no hard degree gate.

Go is listed as the primary language with 'or another systems language'; C++ and Rust are the most common systems-language alternatives in this context and are captured in alternatives.

Slurm and Slinky are listed together ('Slurm/Slinky') in the requirements section; Slinky is treated as an alternative since they serve overlapping workload-scheduling roles in HPC contexts.

The 'APIs' mention in the requirements section is general; captured as REST (the canonical API paradigm) rather than a specific product.

Networking, cluster topology, cloud infrastructure, and large-scale compute familiarity are listed under requirements but framed with 'familiarity with' — a softer gate; however they appear in the required section. Given the hedged language ('familiarity'), these are treated as context rather than hard-gated named technologies and are omitted from the skills list per the instruction to list only concrete, named tools.

The 'Ways to stand out from the crowd' section is treated as preferred/nice-to-have throughout.

Compensation is stated as base salary only; equity and benefits are also mentioned but not quantified.

Caller marked this a fully-remote role — scored against the national candidate pool.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗