Principal Software Engineer, At-Scale Reliability and Fleet Intelligence CSP Engagements at Nvidia
$272,000–$431,250from the description
Jun 29, 2026
—
Jul 20, 2026
What this job asks for AI summary
Senior level · 15+ years · San Jose-Sunnyvale-Santa Clara, CA · Full-time
Advertised as Principal, but the requirements read as Senior.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
Rare in this occupation — lead with these, and say what you built with them.
What the occupation pays Median $217,796 (middle half $177,469–$231,052). This posting is about at that midpoint.
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.
Why we read it this way (7)
The JD accepts 'equivalent experience' in lieu of a BS/MS degree, so the degree requirement is set to None.
The role's actual scope — a senior individual contributor driving cross-functional reliability work streams for key customers — aligns with Senior on the career ladder. The 'Principal' title is the advertised level, but the requirements (even at 15+ years) describe a highly experienced Senior IC rather than an org-wide technical authority with Principal-level scope.
SOC classification is a genuine judgment call: the role is primarily systems software and reliability engineering (15-1252), but the heavy emphasis on statistical failure analysis, predictive modeling, and fleet telemetry analytics gives 15-2051 (Data Scientists) a credible claim as a runner-up.
Survival analysis and Weibull analysis appear in the 'What you'll be doing' responsibilities section rather than the 'What we need to see' requirements section; they are listed as preferred accordingly.
Skills under 'Ways to stand out from the crowd' (NVIDIA GPU error taxonomy, NVLink, health scoring models, telemetry pipelines) are explicitly framed as differentiators, not gates, and are marked preferred.
The posting lists Santa Clara, CA as the location with no remote option mentioned; on-site is assumed.
Ignored 9 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): multi-NUMA system software, statistical failure analysis, MTBF/MTBI analysis, Pareto analysis, burn-in/stress testing frameworks, predictive maintenance, survival analysis, Weibull analysis, NVIDIA GPU error taxonomy.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.