Principal Research & Engineering, Realtime Voice AI at Inflection AI
Palo Alto, CA
$400,000–$550,000from the description
Jun 29, 2026
Palo Alto, CA
Jul 21, 2026
What this job asks for AI summary
A technical leadership role focused on building and owning Inflection's real-time Voice AI stack end to end — spanning streaming speech recognition, text-to-speech, speech-to-speech models, turn-taking, and low-latency inference. The work involves setting the technical roadmap, making build-vs-buy-vs-train decisions, designing voice quality evaluation systems, and deploying voice agents for enterprise use cases, while also mentoring a team of speech researchers and audio engineers.
Staff level · San Jose-Sunnyvale-Santa Clara, CA · Full-time
Advertised as Principal, but the requirements read as Staff.
We read this from the posting text with AI. Skim the description below before ruling yourself out.
How this req sits in the market our data
Roughly 45 people in the San Jose-Sunnyvale-Santa Clara, CA area plausibly meet what this posting asks for (software developers). range 15–70
Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.
Rare in this occupation — lead with these, and say what you built with them.
What the occupation pays Median $217,796 (middle half $177,469–$231,052). This posting is about at that midpoint.
Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.
Why we read it this way (6)
The role is framed as 'principal Research and Engineering contributor' and 'hands-on technical leader' with roadmap ownership, build-vs-buy strategy, and team mentoring — consistent with a Staff-level IC rather than a Principal (org-wide authority) or a people-manager (11-3021). The title is not explicitly stated in the posting; 'principal contributor' is used descriptively, not as a formal title, so the level advertised in the title is set to Principal based on that framing.
The SOC is a close call between 15-1252 (Software Developers — production systems, realtime stack) and 15-2051 (Data Scientists — speech modeling, research). The role spans both, but the emphasis on building and owning a production Voice AI stack, streaming systems, and runtime tips it toward 15-1252.
The requirements section lists experience with 'one or more of' a set of speech/audio technologies (streaming ASR, TTS, speech-to-speech, speech LLMs, audio tokenization, multimodal models, barge-in, low-latency inference, realtime agents). These are presented as a menu rather than all being individually required; ASR, TTS, speech LLMs, realtime voice AI, and low-latency inference are marked required as the core gates, while the remaining items from that list are marked preferred.
The degree requirement states 'bachelor's degree or equivalent in a related field,' which explicitly accepts equivalent experience — treated as no hard degree gate.
Location is inferred as the San Francisco Bay Area based on the mention of 'Bay Area' in the visa support benefit; no explicit city is named for the office.
Ignored 3 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): realtime voice/speech AI, speech-to-speech systems, barge-in/interruption handling.
Read the full posting
The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.