Back to search
EngineeringOn-site

Software Engineer, Inference

Pulse · On-site

Apply
Last seen by MeritLog September 9, 2026Source: AshbySource version: ashby-public-job-posting-v1

MeritLog read this listing from Pulse's Ashby job board and last checked it on September 9, 2026.

Source: the employer's Ashby job board. Open the original listing for current details.

Job details

Work model
On-site
Salary
$150K - $230K
Location
San Francisco
Occupation
Software Developers(O*NET 15-1252.00)
Company website
www.runpulse.com

Hiring context

How this role compares at Pulse

Pulse has 10 live roles in MeritLog’s catalog across 4 job families, and 6 of them are in engineering. 7 of those listings publish a pay range, a disclosure rate of 70%.

Counted across the job boards MeritLog tracks, at the time this page was served. Pay comparisons use only listings that publish a complete range in the same currency and period.

What the role asks for

What you'd do

  • Build inference services with smart batching and caching
  • Optimize kernels, tokenization, and model graphs
  • Evaluate vLLM, TensorRT LLM, and Triton tradeoffs
  • Implement autoscaling and admission control with clear SLOs
  • Own performance dashboards and capacity planning

What they're asking for

  • 3+ years in performance engineering or ML systemsExperience
  • Strong Python, plus C++ or CUDA exposureSkill
  • Experience with GPU profiling and model servingSkill
  • Experience reducing p95 and cost in production ML systemsSkillPreferred

Parsed by MeritLog from the employer’s own posting. The full description follows below.

Job description

Overview Pulse is tackling one of the most persistent challenges in data infrastructure: extracting accurate, structured information from complex documents at scale. We have a breakthrough approach to document understanding that combines intelligent schema mapping with fine-tuned extraction models where legacy OCR and other parsing tools consistently fail. We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises, YC startups, public investment firms, and growth-stage companies. We are backed by tier 1 investors and growing quickly. What makes our tech special https://www.runpulse.com/blog/pulse-approach-to-document-intelligence is our multi-stage architecture: - Layout understanding with specialized component detection models - Low-latency OCR models for targeted extraction - Advanced reading-order algorithms for complex structures - Proprietary table structure recognition and parsing - Fine-tuned vision-language models for charts, tables, and figures If you are passionate about the intersection of computer vision, NLP, and data infrastructure, your work at Pulse will directly impact customers and shape the future of document intelligence. What we are looking for - 5 days in-office at our San Francisco office - Eager to learn and adapt quickly - Prior startup or founding experience is a plus What we are looking for - 5 days in-office at our San Francisco office - Eager to learn and adapt quickly - Prior startup or founding experience is a plus About the Role Specialize in low-latency, high-throughput inference for OCR and multimodal models. Own profiling, batching, and autoscaling across single-tenant and multi-tenant environments. Responsibilities - Build inference services with smart batching and caching - Optimize kernels, tokenization, and model graphs - Evaluate vLLM, TensorRT LLM, and Triton tradeoffs - Implement autoscaling and admission control with clear SLOs - Own performance dashboards and capacity planning Requirements - 3+ years in performance engineering or ML systems - Strong Python, plus C++ or CUDA exposure - Experience with GPU profiling and model serving Nice to have - Experience reducing p95 and cost in production ML systems Sponsorship Sponsorship available. Compensation and benefits Competitive base salary plus equity, performance-based bonus, relocation assistance for Bay Area moves, daily meal stipend, medical, vision, and dental coverage.

Privacy choices

Analytics and advertising stay off unless you allow them. Private data stays out.

Read the privacy notice