Senior Systems Software Engineer, Observability and Telemetry Platform
NVIDIA Corporation · Santa Clara, CA
$184,000–$356,500from the description
Jul 13, 2026
Santa Clara, CA
Jul 21, 2026
What this job asks for AI summary
This role centers on building and operating a large-scale observability and telemetry platform for GPU cloud services, with a focus on reliability, performance, and automation. Day-to-day work spans system design, capacity management, incident response, and eliminating manual toil through tooling and automation. It suits engineers with deep experience in distributed systems, infrastructure platforms, and observability tooling such as Prometheus, Grafana, and OpenTelemetry.
Senior level · 8+ years · Remote · Full-time
“or” means any one of them counts — you don't need all of them.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
Roughly 7,700 people nationally plausibly meet what this posting asks for (software developers). range 4,650–11,600
Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.
Rare in this occupation — lead with these, and say what you built with them.
Most people in this occupation already list these. Still required — just not what gets you shortlisted.
What the occupation pays Median $138,970 (middle half $107,524–$175,762). This posting is about at that midpoint.
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.
Why we read it this way (11)
The compensation range spans two internal levels: Level 4 ($184,000–$287,500) and Level 5 ($224,000–$356,500). The bottom of the pay band/max here reflect the full combined range across both levels.
SOC classification is Medium confidence. The role blends software development (building automation, tools, platforms) with systems/infrastructure operations (SRE-style on-call, incident response, capacity management). The primary day-to-day emphasis on writing software and building platforms tips it to 15-1252 over 15-1244, but the operational SRE framing is a genuine competing signal.
Python, Go, Perl, and Ruby are listed together as 'one or more of the following' under the required qualifications section — emitted as a single skill with alternatives.
Linux, networking, and containers are stated as 'in depth knowledge' under the required section and are hard gates.
'Infrastructure automation' and 'distributed systems design' are role-level experience requirements (8+ years) rather than named tools; they are captured here as skills because the JD explicitly names them as required competencies, though they are not specific technologies.
Kubernetes, OpenStack, Docker, Grafana, OpenTelemetry, and Prometheus all appear under the 'Ways to stand out from the crowd' section (preferred/nice-to-have), not in the required qualifications block.
The 5+ years of observability platform experience is a role-level requirement stated separately from the 8+ years general infrastructure requirement; both are captured.
A BS degree in Computer Science or related field is listed but explicitly allows 'equivalent experience,' so the degree requirement is set to None.
Remote status set to true per caller declaration; no metro is inferred.
Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): infrastructure automation.
Caller marked this a fully-remote role — scored against the national candidate pool.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.