Member of Technical Staff, Site Reliablity Engineer
Vapi · Hybrid
MeritLog read this listing from Vapi's Ashby job board and last checked it on September 12, 2026.
Source: the employer's Ashby job board. Open the original listing for current details.
Job details
- Work model
- Hybrid
- Salary
- $200K - $270K
- Location
- San Francisco
- Company website
- www.daily.co
Hiring context
How this role compares at Vapi
Vapi has 35 live roles in MeritLog’s catalog across 8 job families, and 12 of them are in engineering. 34 of those listings publish a pay range, a disclosure rate of 97%.
This role's posted range of $200K - $270K sits above 70% of the 33 other Vapi roles quoted over the same currency and period.
Vapi concentrates this hiring in:
Counted across the job boards MeritLog tracks, at the time this page was served. Pay comparisons use only listings that publish a complete range in the same currency and period.
What the role asks for
What you'd do
- 30 Day: Join the oncall rotation. Walk the 15 stability-gap incidents and turn the patterns into a prioritized reliability backlog. Define the first set of SLOs for the call-completion path.
- 60 Day: Stand up error budgets and SLO-based alerting in Chronosphere/Prometheus for the highest-impact services. Run the first proper load test against provider rate limits and per-org concurrency. Tune autoscaling for wscaler / workerpool-cron-scaler.
- 90 Day: Ship a real platform service - capacity forecaster, auto-remediation, or oncall tooling - in Go or TypeScript. Own the postmortem process. Drive a measurable improvement in p99 call completion or MTTR.
What they're asking for
- You’ve run incident command and postmortem discipline at scale on a real oncall rotation.Skill
- You’ve operated SLOs and error budgets in Chronosphere, Prometheus, Grafana, or Datadog.Skill
- You’ve done capacity planning and load testing for production systems with real users.Skill
- You’re fluent in Kubernetes production ops: pod crash diagnosis, HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown.Skill
- You know backpressure and autoscaling patterns - KEDA, custom metrics scaling.Skill
- You ship code, not just scripts. You can build platform services in Go or TypeScript (matches Vapi’s cluster-manager, database-health, wscaler, incidentManager).SkillPreferred
- Real-time / latency-sensitive product background where degraded means a dropped call, not a slow dashboard.SkillPreferred
- Languages: Go and TypeScript (you ship code, not just scripts), Bash.SkillPreferred
- Observability: Chronosphere, Prometheus, Grafana, Datadog, OpenTelemetry.SkillPreferred
- Orchestration: Kubernetes on EKS - production ops (HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown, pod crash diagnosis).SkillPreferred
- Autoscaling and backpressure: KEDA, custom metrics scaling (matches Vapi’s wscaler and workerpool-cron-scaler).SkillPreferred
- Load testing: script-based load testing, provider rate-limit auditing, per-org concurrency auditing.SkillPreferred
- Vapi services you’ll touch or build: cluster-manager, database-health, wscaler, incidentManager.SkillPreferred
- A real-time / latency-sensitive product (Discord, Zoom, Mux, Twitch, Twilio, LiveKit, Cloudflare, a trading firm, a gaming backend), or a FAANG SRE / Production Engineer (Google, Uber, Twitter/X, Meta) who misses being hands-on.SkillPreferred
- Weak fit: SRE from analytics or CRM backends where “degraded” means a slow dashboard, not a dropped call. Anyone uncomfortable reading or writing code.SkillPreferred
- Generational impact: Build the human interface for every businessSkillPreferred
- Ownership culture: 70% of the company are previous foundersSkillPreferred
- Kind team: The founders, Jordan and Nikhil, are CanadiansSkillPreferred
- Tier-1 Investors: YC, KP seed, Bessemer Series ASkillPreferred
Parsed by MeritLog from the employer’s own posting. The full description follows below.
Job description
Vapi (/ˈVɑːpi/): - Voice AI that resolves, not transfers - Powering 1 billion calls for companies like Amazon Ring, Intuit, ServiceTitan, and New York Life - Trusted by 1 million developers building the future of voice agents - Backed by Peak XV, Bessemer, Kleiner Perkins, M12, Y Combinator, and more with $72M raised - Try talking to Vapi now! https://vapi.ai/ WHY WE’RE HIRING THIS ROLE: - 99.99% call completion is the number this role drives. Vapi runs live phone calls - a p99 spike means callers drop. We’ve had 15 stability-gap outages worth learning from, and we need someone who runs incident command, owns SLOs and error budgets, and builds the reliability culture from scratch. - This is not a bash-and-YAML role. You’ll ship code (Go or TypeScript) for services that monitor and manage the platform: auto-remediation, capacity forecasters, oncall tooling. Capacity planning, load testing, and KEDA-based autoscaling for Vapi’s wscaler and workerpool-cron-scaler are on your plate. WHAT YOU’LL DO: - 30 Day: Join the oncall rotation. Walk the 15 stability-gap incidents and turn the patterns into a prioritized reliability backlog. Define the first set of SLOs for the call-completion path. - 60 Day: Stand up error budgets and SLO-based alerting in Chronosphere/Prometheus for the highest-impact services. Run the first proper load test against provider rate limits and per-org concurrency. Tune autoscaling for wscaler / workerpool-cron-scaler. - 90 Day: Ship a real platform service - capacity forecaster, auto-remediation, or oncall tooling - in Go or TypeScript. Own the postmortem process. Drive a measurable improvement in p99 call completion or MTTR. WHO YOU ARE: Must-haves - You’ve run incident command and postmortem discipline at scale on a real oncall rotation. - You’ve operated SLOs and error budgets in Chronosphere, Prometheus, Grafana, or Datadog. - You’ve done capacity planning and load testing for production systems with real users. - You’re fluent in Kubernetes production ops: pod crash diagnosis, HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown. - You know backpressure and autoscaling patterns - KEDA, custom metrics scaling. Nice-to-haves - You ship code, not just scripts. You can build platform services in Go or TypeScript (matches Vapi’s cluster-manager, database-health, wscaler, incidentManager). - Real-time / latency-sensitive product background where degraded means a dropped call, not a slow dashboard. Tech stack you’ll work in - Languages: Go and TypeScript (you ship code, not just scripts), Bash. - Observability: Chronosphere, Prometheus, Grafana, Datadog, OpenTelemetry. - Orchestration: Kubernetes on EKS - production ops (HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown, pod crash diagnosis). - Autoscaling and backpressure: KEDA, custom metrics scaling (matches Vapi’s wscaler and workerpool-cron-scaler). - Load testing: script-based load testing, provider rate-limit auditing, per-org concurrency auditing. - Vapi services you’ll touch or build: cluster-manager, database-health, wscaler, incidentManager. Where you likely come from - A real-time / latency-sensitive product (Discord, Zoom, Mux, Twitch, Twilio, LiveKit, Cloudflare, a trading firm, a gaming backend), or a FAANG SRE / Production Engineer (Google, Uber, Twitter/X, Meta) who misses being hands-on. - Weak fit: SRE from analytics or CRM backends where “degraded” means a slow dashboard, not a dropped call. Anyone uncomfortable reading or writing code. WHY VAPI: - Generational impact: Build the human interface for every business - Ownership culture: 70% of the company are previous founders - Kind team: The founders, Jordan and Nikhil, are Canadians - Tier-1 Investors: YC, KP seed, Bessemer Series A WHAT WE OFFER: - Real stake: We offer a competitive salary and excellent equity ownership - Comprehensive health coverage: medical, dental, and vision plans - Team love: We love hanging out, and we do quarterly off-sites - Flexible time off: take what you need More: catered meals, transportation, gym, and a $10k annual L&D budget
Keep exploring
More Engineering roles
- Sales EngineerZscaler · Remote
- Sr. Software Development Engineer - Python Automation / Kubernetes / NetworkingZscaler · On-site
- Staff Software Development Engineer (Backend - Java/API)Zscaler · On-site
- Escalation Engineer - DLPZscaler · Hybrid
- Staff Software Development EngineerZscaler · Hybrid
- Senior Sales Engineer - Finance and StrategicZscaler · Hybrid