Back to search
Listing unavailableEngineeringHybrid

Senior/Staff AI Engineer

DDN · Hybrid

This listing is no longer verified as available.

MeritLog keeps this source-backed description for reference. Availability is not verified, and there is no application link here.

Last seen by MeritLog September 12, 2026Source: AshbySource version: ashby-public-job-posting-v1

Source: the employer's Ashby job board. Open the original listing for current details. Availability is not verified for this retained page.

Job details

Work model
Hybrid
Salary
Not listed by source
Location
Remote - California

What the role asks for

What you'd do

  • Build and optimize LLM serving and inference systems for production environments
  • Improve performance across GPU and CPU pathways
  • Work on KV cache, memory, storage, and throughput bottlenecks
  • Design and scale systems that support RAG and retrieval-heavy AI workloads
  • Contribute to infrastructure where storage architecture and systems efficiency materially affect AI performance
  • Solve engineering problems at the intersection of AI, high-performance systems, and distributed infrastructure
  • An engineer who has spent meaningful time building or optimizing production AI systems, not just experimenting with models
  • Someone who understands how inference performance is shaped by the interaction between compute, memory, storage, and serving architecture
  • Deep hands-on experience working close to the systems layer - for example, improving how workloads run across GPU and CPU resources, reducing bottlenecks, or tuning infrastructure for better throughput and latency
  • Evidence of real ownership in areas like model serving, retrieval, caching, storage, or distributed performance, rather than purely application-layer AI work
  • The ability to move comfortably between architecture decisions and hands-on implementation, especially in environments where efficiency and scale matter
  • A background that suggests you can operate in technically demanding environments, whether that comes from AI infrastructure, high-performance systems, storage platforms, or adjacent distributed systems work
  • PhD preferred, but far less important than having built serious systems in the real world
  • This is not a “prompt engineering” job.
  • This is not an “AI wrapper” job.
  • This is not a generic backend role with AI sprinkled on top.
  • This is a chance to work on the infrastructure that determines whether modern AI systems are fast, scalable, efficient, and commercially viable.
  • If you want to work on the real mechanics of AI performance - serving, retrieval, compute efficiency, memory behavior, storage architecture, and inference at scale - this is where that work happens.
  • Engineers who enjoy deep systems problems
  • Builders who care about performance, scale, and architecture
  • People who want to work where AI meets infrastructure
  • Candidates who would rather solve hard technical bottlenecks than ship surface-level AI features
  • Purely academic researchers without meaningful production ownership
  • Generic software engineers without clear AI systems or inference depth
  • Candidates focused mainly on prompt engineering or lightweight application integrations
  • MLOps generalists who have not worked deeply on serving, storage, or performance-critical AI systems

Parsed by MeritLog from the employer’s own posting. The full description follows below.

Job description

WHAT YOU’LL DO - Build and optimize LLM serving and inference systems for production environments - Improve performance across GPU and CPU pathways - Work on KV cache, memory, storage, and throughput bottlenecks - Design and scale systems that support RAG and retrieval-heavy AI workloads - Contribute to infrastructure where storage architecture and systems efficiency materially affect AI performance - Solve engineering problems at the intersection of AI, high-performance systems, and distributed infrastructure WHAT WE’RE LOOKING FOR - An engineer who has spent meaningful time building or optimizing production AI systems, not just experimenting with models - Someone who understands how inference performance is shaped by the interaction between compute, memory, storage, and serving architecture - Deep hands-on experience working close to the systems layer - for example, improving how workloads run across GPU and CPU resources, reducing bottlenecks, or tuning infrastructure for better throughput and latency - Evidence of real ownership in areas like model serving, retrieval, caching, storage, or distributed performance, rather than purely application-layer AI work - The ability to move comfortably between architecture decisions and hands-on implementation, especially in environments where efficiency and scale matter - A background that suggests you can operate in technically demanding environments, whether that comes from AI infrastructure, high-performance systems, storage platforms, or adjacent distributed systems work - PhD preferred, but far less important than having built serious systems in the real world WHY THIS ROLE IS COMPELLING - This is not a “prompt engineering” job. - This is not an “AI wrapper” job. - This is not a generic backend role with AI sprinkled on top. - This is a chance to work on the infrastructure that determines whether modern AI systems are fast, scalable, efficient, and commercially viable. - If you want to work on the real mechanics of AI performance - serving, retrieval, compute efficiency, memory behavior, storage architecture, and inference at scale - this is where that work happens. WHO WILL LOVE THIS ROLE - Engineers who enjoy deep systems problems - Builders who care about performance, scale, and architecture - People who want to work where AI meets infrastructure - Candidates who would rather solve hard technical bottlenecks than ship surface-level AI features WHO SHOULD NOT APPLY This role is not for: - Purely academic researchers without meaningful production ownership - Generic software engineers without clear AI systems or inference depth - Candidates focused mainly on prompt engineering or lightweight application integrations - MLOps generalists who have not worked deeply on serving, storage, or performance-critical AI systems -

Keep exploring

Available Engineering roles

These current listings are available to explore now.

Search all jobs

Privacy choices

Analytics and advertising stay off unless you allow them. Private data stays out.

Read the privacy notice