Perplexity · San Francisco, CA

Salary
$250,000–$485,000
Posted
Jul 15, 2026
Location
San Francisco, CA
Last confirmed open
Jul 21, 2026

What this job asks for AI summary

This role is responsible for building and operating a unified, self-serve GPU compute platform that abstracts away the complexity of a multi-cloud fleet for inference engineers and researchers. Day-to-day work spans Kubernetes operator development, multi-cluster GPU provisioning and lifecycle management, cross-provider scheduling, and fault tolerance — covering both long-running distributed training jobs and latency-sensitive production inference services. It suits engineers with deep Kubernetes and GPU cluster experience who are comfortable owning infrastructure end-to-end with limited direction.

Senior level · National · Full-time

Must have (8)
KubernetesKubernetes operatorsCRDsNVIDIA GPUCUDAInfiniBand or RoceGo, Rust or C++AWS, GCP or Coreweave
Nice to have (4)
vLLM, Sglang or Tensorrt LlmSlurmTritonPrometheus or Weights Biases

“or” means any one of them counts — you don't need all of them.

Posted 2 times — it's one opening, so apply once.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

What gives you an edge
CUDA4%

Rare in this occupation — lead with these, and say what you built with them.

What the occupation pays Median $138,970 (middle half $107,524–$175,762).

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (7)

No location is specified in the posting; the CBSA and state fields are left blank. The role appears to be US-based given the company (Perplexity AI, San Francisco) but the JD itself does not state a location or remote policy.

The posting has no explicit title or seniority level. The scope — owning a GPU compute platform end-to-end, setting technical direction across teams, and operating a large multi-cloud fleet — is consistent with a Senior or Staff engineer. Sized as Senior because the JD does not explicitly grant org-wide architectural authority across multiple teams (a Staff signal); 'partner with' and 'set technical direction' language is present but falls short of a clear Staff mandate.

Go, Rust, and C++ are presented as interchangeable options for the required systems-level coding language ('Go, Rust or C++'); Go is listed first and is the most common choice in Kubernetes-adjacent platform work.

InfiniBand and RoCE appear together in the qualifications section as the required high-speed networking knowledge; RoCE is captured as an alternative. They also appear separately under 'Additional experience we value' (with RDMA added), but the primary gate is in the qualifications block.

The multi-cloud requirement names CoreWeave, AWS, and GCP as examples ('or similar'); AWS is used as the primary skill name with the others as alternatives.

Inference serving stacks (vLLM / SGLang / TensorRT-LLM), Slurm, GPU kernel work (CUDA/Triton), and observability tools (Prometheus / Grafana / Weights & Biases) all appear under the 'Additional experience we value' section and are treated as preferred. Note that CUDA also appears in the required qualifications block as part of GPU cluster management knowledge — that instance is captured as a separate must-have skill.

No compensation range is disclosed in the posting.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗