Boston, MA

Salary
$226,000–$307,000
Posted
Jun 29, 2026
Location
Boston, MA
Last confirmed open
Jul 21, 2026

What this job asks for AI summary

This role sits within an autonomous vehicle perception team and centers on deploying large-scale ML models — including multimodal, LLM, and VLM architectures — onto power- and thermally-constrained vehicle hardware. Day-to-day work involves allocating CPU/GPU resources across on-vehicle inference engines, applying quantization and pruning techniques, building TensorRT compilation pipelines, and writing low-latency C++ and CUDA code for real-time edge execution. It suits engineers with deep experience in model compression and GPU-level performance optimization for production systems.

Senior level · San Jose-Sunnyvale-Santa Clara, CA · Full-time

Pay in the description: $226,000–$307,000

Must have (9)
model quantization (PTQ/QAT)mixed-precision inferenceCUDATensorRTC++PythonLLMs, Vlms or Foundation Modelsreal-time systemsedge deployment
Nice to have (4)
TensorRT-LLMLoRA or QloraLiDARRadar

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

Roughly 15 people in the San Jose-Sunnyvale-Santa Clara, CA area plausibly meet what this posting asks for (data scientists). range 3–30

Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.

What gives you an edge
CUDA4%

Rare in this occupation — lead with these, and say what you built with them.

What won't set you apart
Python88%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $185,080 (middle half $149,570–$219,460). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (7)

This role sits at the intersection of ML model optimization and systems/low-level software engineering, making it genuinely ambiguous between 15-2051 (Data Scientists, for the ML model compression/quantization/deployment work) and 15-1252 (Software Developers, for the production C++/CUDA kernel and inference engine work). The primary day-to-day emphasis on model efficiency, quantization, and on-device ML deployment tips the classification toward 15-2051, but 15-1252 is a strong runner-up.

Zoox is based in Foster City, CA (San Francisco Bay Area / San Jose CBSA). The posting does not explicitly state a location but Zoox's headquarters is in Foster City, CA.

The salary range ($226K–$307K/year) is described as covering 'a range of levels Zoox is considering,' so the actual level for a given hire may fall within a narrower band.

Model quantization techniques (PTQ, QAT) and mixed-precision formats (INT8, FP8, FP4, BF16/FP16) are listed under the required Qualifications section and are treated as hard gates. LoRA/QLoRA appear in the responsibilities narrative rather than the Qualifications section and are treated as preferred.

TensorRT appears in both the responsibilities (model conversion pipelines) and the required Qualifications (TensorRT Plugins), making it a hard gate. TensorRT-LLM appears only under Bonus Qualifications and is preferred.

Autonomous driving perception algorithms (BEV, 3D object detection, occupancy networks) and multi-modal sensor modalities (LiDAR, Radar) appear exclusively under Bonus Qualifications and are preferred.

Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): CPU/GPU performance optimization.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗