Plaud · Seattle, WAremote

Salary
Posted
Jul 21, 2026
Location
Seattle, WA
Last confirmed open
Jul 22, 2026

What this job asks for AI summary

A site reliability engineering role focused on keeping AI-driven cloud products stable and performant at scale. Day-to-day work covers designing and operating cloud-native infrastructure on major providers, managing Kubernetes-based distributed systems, building out observability tooling, and owning incident response and on-call practices. The role also involves defining SLOs and error budgets and collaborating with product and engineering teams on reliability design. Suited to experienced SRE or platform engineers with a background in distributed systems.

Senior level · 5+ years · San Francisco-Oakland-Berkeley, CA · Full-time

Must have (3)
AWS, GCP, Azure or Oracle Cloud InfrastructureKubernetesGo, Python or Java
Nice to have (3)
AI/ML platformsSLO/SLA frameworksmulti-region systems

“or” means any one of them counts — you don't need all of them.

Posted 2 times — it's one opening, so apply once.

We read this from the posting text with AI. Skim the description below before ruling yourself out.

How this req sits in the market our data

Roughly 1,350 people in the San Francisco-Oakland-Berkeley, CA area plausibly meet what this posting asks for (software developers). range 1,000–1,750

Applicant volume Moderate — A normal amount of company. The rare requirements below are what will separate a shortlisted application from the rest.

What won't set you apart
Go51%

Most people in this occupation already list these. Still required — just not what gets you shortlisted.

What the occupation pays Median $190,744 (middle half $167,095–$224,501).

Estimated from BLS employment for this occupation and area, per-skill prevalence across our listing corpus, and published wage benchmarks — as of Jul 28, 2026. It is a model, not a headcount.

Why we read it this way (7)

The role title is not stated in the posting; based on the responsibilities (SRE, reliability engineering, incident response, observability, SLOs), this is classified as an SRE/Platform Engineering role. The SOC code is Medium confidence: the JD emphasizes building reliability automation, observability tooling, and operating cloud-native systems — consistent with Software Developers (15-1252) — but the strong operational/incident-response framing also pulls toward Systems Administrators (15-1244).

Cloud platform requirement lists AWS, GCP, Azure, and OCI as interchangeable options; all four are captured as alternatives under AWS.

Programming language requirement lists Go, Python, and Java as interchangeable; all three are captured as alternatives under Go.

SLO/SLA frameworks, GPU cluster management, multi-region systems, and AI/ML platform experience all appear under 'Preferred Qualifications' and are marked accordingly.

No compensation figures are provided in the posting.

The hybrid work model requires a minimum of three in-office days per week at the San Francisco office, so the role is not fully remote.

Ignored 1 non-technology phrase(s) as skills (responsibilities/concepts, not named tools): GPU cluster management.

Read the full posting

The employer publishes the full description on their own site — read it there ↗. Or sign in to read it here — it's free, and it also lets you track this application.

Apply

Apply on employer site ↗