Foster City, MI

Salary
$242,000–$290,000
Posted
Jul 16, 2026
Location
Foster City, MI
Last confirmed open
Jul 21, 2026

What this job asks for AI summary

A senior deep learning engineering role focused on building 3D occupancy and segmentation models for an autonomous vehicle system. The work centers on designing multi-modal sensor fusion architectures combining lidar, camera, and radar data into voxel-based scene representations, with attention to temporal consistency and real-time on-vehicle inference. Outputs feed directly into tracking, prediction, and planning components. Suits candidates with strong 3D computer vision and multi-sensor fusion backgrounds.

Senior level · 6+ years · San Jose-Sunnyvale-Santa Clara, CA · Master's required · Full-time

Pay in the description: $242,000–$290,000

Must have (8)
PythonPyTorchC++3D Computer Visiondeep learningmulti-sensor fusionLiDARoccupancy networks, Nerf, Gaussian Splatting or Scene Flow Estimation
Nice to have (3)
TensorRT or Cudasparse convolutionsVision Language Models

“or” means any one of them counts — you don't need all of them.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

Roughly 100 people in the San Jose-Sunnyvale-Santa Clara, CA area plausibly meet what this posting asks for (data scientists). range 20–150

Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.

What won't set you apart
Python88%deep learning40%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $189,150 (middle half $152,859–$224,286). This posting is about at that midpoint.

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (7)

The role sits at the intersection of deep learning research and production engineering — it could reasonably be coded as Software Developers (15-1252) given the emphasis on architecting and optimizing models for on-vehicle inference in C++/TensorRT, but the core work (designing and training 3D perception networks, occupancy/segmentation/flow modeling) aligns more closely with Data Scientists (15-2051). Medium confidence reflects this genuine ambiguity.

The degree requirement is MS or PhD; because a Master's is the stated minimum and no equivalent-experience substitution is offered, the degree requirement is set to Masters.

Zoox is headquartered in Foster City, CA (San Jose-Sunnyvale-Santa Clara CBSA); the posting does not specify a city but Zoox's primary site is in that metro.

The 'occupancy networks, implicit representations (NeRF/Gaussian Splats), or scene flow estimation' requirement is listed as a single disjunctive gate in the required qualifications; NeRF, Gaussian Splatting, and scene flow estimation are captured as alternatives to occupancy networks.

TensorRT/CUDA optimization and sparse convolutions appear under 'Bonus Qualifications' and are marked preferred accordingly.

Vision Language Models, multi-modal 3D foundation models, World Models, and VLA all appear under 'Bonus Qualifications'; they are grouped under the Vision Language Models skill as the most concrete named technology in that list.

The compensation range ($242,000–$290,000/year) covers base salary only; the posting notes that RSUs, Zoox Stock Appreciation Rights, and a potential sign-on bonus are additional components.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗