Back to search
Data & AnalyticsOn-site

Inference Systems Performance Architect

SambaNova · On-site

Apply
Last seen by MeritLog September 12, 2026Source: GreenhouseSource version: greenhouse-job-board-v1

MeritLog read this listing from SambaNova's Greenhouse job board and last checked it on September 12, 2026.

Source: the employer's Greenhouse job board. Open the original listing for current details.

Job details

Work model
On-site
Salary
Conflicting source ranges
Location
San Jose, California, United States

Hiring context

How this role compares at SambaNova

SambaNova has 55 live roles in MeritLog’s catalog across 9 job families, and 13 of them are in data & analytics. 0 of those listings publish a pay range, a disclosure rate of 0%.

Counted across the job boards MeritLog tracks, at the time this page was served. Pay comparisons use only listings that publish a complete range in the same currency and period.

What the role asks for

What you'd do

  • Define and drive the technical strategy for inference-systems performance including workload capture, benchmarking, modeling, and simulation, while developing and architecture that enables many potential futures
  • Build the workload-capture and agentic-benchmarking capability - capture representative production traffic and enforce the discipline of interrogating results, spotting artificial contention or misleadingly high cache-hit rates that never occur in real use
  • Own the performance-modeling and simulation practice - models that predict how a configuration change moves the output, informing capacity planning against customer SLOs and next-generation system and hardware planning
  • Attack the end-to-end profiling gap - drive tooling that produces accurate, actionable profiles of a distributed inference pipeline so bottlenecks can be localized across host, accelerator, and fabric
  • Serve as the senior technical voice across model-optimization, systems, hardware, and product, tying together multiple engineering activities and teams, and weighing trade-offs of reliability, scalability, operational cost, and ease of adoption
  • Act as a resource for the entire organization including representing SambaNova's performance story to customers and partners
  • Mentor and multiply by raising the capability of principal and senior engineers, building the systems, tools, and patterns that make everyone more productive
  • Drive the resolution of the most ambiguous, novel challenges that span organizational boundaries or have no established answer in the field yet

What they're asking for

  • B.S. in Computer Science, Computer Engineering, or Related FieldEducation
  • 12+ years of experience in performance engineering, with a demonstrated record of technical leadership on large-scale, complex systemsExperience
  • Deep expertise in end-to-end performance analysis of distributed systems with many moving parts and the ability to localize bottlenecks that others cannotSkill
  • Proven command of realistic workload generation and simulation and of performance modeling, including calibrating models against real, variable workloadsSkill
  • Demonstrated ability to enter an unfamiliar domain and apply core performance methods with transferable discipline expertiseSkill
  • Ability to lead cross-functional efforts, mentor senior engineers, and influence organizational directionSkill
  • Experience representing an organizations credibly to customers and partnersSkill
  • Track record of independently scoping and delivering high-complexity, high-ambiguity work with significant impact on products or roadmapSkill
  • M.S. or PHD in Computer Science, Computer Engineering, or Related FieldEducationPreferred
  • Direct experience with LLM inference serving - continuous batching, prompt/KV caching, prefill/decode disaggregation, tail-latency SLOsSkillPreferred
  • Familiarity with inference simulation frameworks or agentic benchmarking effortsSkillPreferred
  • A public technical voice - talks, writing, or community presence on systems performanceSkillPreferred

Parsed by MeritLog from the employer’s own posting. The full description follows below.

Job description

SambaNova is a leader in next-generation AI infrastructure, delivering a full-stack inference platform for customers worldwide. At the core of SambaNova's technology is the RDU (Reconfigurable Dataflow Unit) - a chip built on a dataflow architecture rather than the traditional GPU model. Its decode performance is especially strong for agentic workloads like multi-turn agents, code generation, and long-running applications. RDUs are packaged into SambaRack, rack-scale hardware that lets customers deploy state-of-the-art models with better performance, greater energy efficiency, and faster time to value. About the Team The Inference Systems Performance team answers how fast SambaNova's systems can serve large language models and what it takes to get there. We capture real production traffic, benchmark it faithfully, model and simulate configurations that do not exist yet, and profile the distributed serving pipeline across host, accelerator, and fabric. Our work feeds serving optimization today and hardware and capacity planning for the next generation. We sit close to model optimization, systems, hardware, and product, and the whole company depends on our numbers. About the Role You will be the architect for end-to-end inference performance: how a request moves through tokenization, prefill, decode, and the fabric between them, and how a deployment is sized against customer SLOs. The work has two coupled pillars. One is reproducible workload capture and benchmarking, building replayable representations of real and increasingly agentic traffic so measurements reflect production rather than a naive load script. The other is performance modeling and simulation, turning measurement into a what-if capability for configurations and hardware that do not exist yet. The frontier you will help define is heterogeneous, disaggregated inference, with GPU on prefill and the RDU on decode. That opens hard problems in networking, storage, prompt caching, and tail-latency-bound data movement. Inference systems performance is a young field and most answers are still being discovered, so you will spend your time on ambiguous problems with no established solution, and you will set the technical direction that others build on. Responsibilities • Define and drive the technical strategy for inference-systems performance including workload capture, benchmarking, modeling, and simulation, while developing and architecture that enables many potential futures • Build the workload-capture and agentic-benchmarking capability - capture representative production traffic and enforce the discipline of interrogating results, spotting artificial contention or misleadingly high cache-hit rates that never occur in real use • Own the performance-modeling and simulation practice - models that predict how a configuration change moves the output, informing capacity planning against customer SLOs and next-generation system and hardware planning • Attack the end-to-end profiling gap - drive tooling that produces accurate, actionable profiles of a distributed inference pipeline so bottlenecks can be localized across host, accelerator, and fabric • Serve as the senior technical voice across model-optimization, systems, hardware, and product, tying together multiple engineering activities and teams, and weighing trade-offs of reliability, scalability, operational cost, and ease of adoption • Act as a resource for the entire organization including representing SambaNova's performance story to customers and partners • Mentor and multiply by raising the capability of principal and senior engineers, building the systems, tools, and patterns that make everyone more productive • Drive the resolution of the most ambiguous, novel challenges that span organizational boundaries or have no established answer in the field yet Required qualifications • B.S. in Computer Science, Computer Engineering, or Related Field • 12+ years of experience in performance engineering, with a demonstrated record of technical leadership on large-scale, complex systems • Deep expertise in end-to-end performance analysis of distributed systems with many moving parts and the ability to localize bottlenecks that others cannot • Proven command of realistic workload generation and simulation and of performance modeling, including calibrating models against real, variable workloads • Demonstrated ability to enter an unfamiliar domain and apply core performance methods with transferable discipline expertise • Ability to lead cross-functional efforts, mentor senior engineers, and influence organizational direction • Experience representing an organizations credibly to customers and partners • Track record of independently scoping and delivering high-complexity, high-ambiguity work with significant impact on products or roadmap Preferred qualifications • M.S. or PHD in Computer Science, Computer Engineering, or Related Field • Direct experience with LLM inference serving - continuous batching, prompt/KV caching, prefill/decode disaggregation, tail-latency SLOs • Familiarity with inference simulation frameworks or agentic benchmarking efforts • A public technical voice - talks, writing, or community presence on systems performance Base Salary Range: Base Pay Range $245,000-$325,000 USD Submission Guidelines Please note that in order to be considered an applicant for any position at SambaNova Systems, you must submit an application form for each position for which you believe you are qualified. EEO Policy SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws. Benefits Summary for US-Based, Full-Time Employment Positions SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.

Keep exploring

More Data & Analytics roles

Search all jobs

Privacy choices

Analytics and advertising stay off unless you allow them. Private data stays out.

Read the privacy notice