Back to search
EngineeringOn-site

Intermediate/ Senior Software Engineer - Cortex LLM Training Platform

Snowflake · On-site

Apply
Last seen by MeritLog September 9, 2026Source: AshbySource version: ashby-public-job-posting-v1

MeritLog read this listing from Snowflake's Ashby job board and last checked it on September 9, 2026.

Source: the employer's Ashby job board. Open the original listing for current details.

Job details

Work model
On-site
Salary
$200K - $287.5K
Location
US-WA-Bellevue
Company website
app.snowflake.com

Hiring context

How this role compares at Snowflake

Snowflake has 366 live roles in MeritLog’s catalog across 10 job families, and 153 of them are in engineering. 248 of those listings publish a pay range, a disclosure rate of 68%.

This role's posted range of $200K - $287.5K sits above 64% of the 241 other Snowflake roles quoted over the same currency and period.

Counted across the job boards MeritLog tracks, at the time this page was served. Pay comparisons use only listings that publish a complete range in the same currency and period.

What the role asks for

What you'd do

  • Design and build across the full stack - from the public training APIs and SDK through the control plane to the GPU data plane.
  • Scale the distributed systems that make GPU compute serverless - multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools, with fault tolerance built in.
  • Drive end-to-end performance at scale - keep the training, inference, and RL loops fast and the data plane responsive under heavy concurrent load, with GPUs kept saturated.
  • Productionize research building blocks - partner with Snowflake Research to turn state-of-the-art training and inference techniques into reliable, composable components customers can run at enterprise scale.

What they're asking for

  • 3 + years (Intermediate) | 6+ years (Senior) building and shipping production ML systemsExperience
  • Strong distributed systems and infrastructure foundation - designing scalable, fault-tolerant services and operating them on Kubernetes in production.Skill
  • Familiarity with GPU and LLM infrastructure - e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM; able to debug across the data, infrastructure, and GPU layers.Skill
  • Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency.Skill
  • BS in Computer Science or a related field (MS/PhD a plus).EducationPreferred
  • (Bonus) Hands-on LLM post-training / modeling experience - the strongest candidates pair deep infra skills with real post-training intuition.SkillPreferred

Parsed by MeritLog from the employer’s own posting. The full description follows below.

Job description

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset - who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. SENIOR SOFTWARE ENGINEER - CORTEX TRAINING The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: - Design and build across the full stack - from the public training APIs and SDK through the control plane to the GPU data plane. - Scale the distributed systems that make GPU compute serverless - multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools, with fault tolerance built in. - Drive end-to-end performance at scale - keep the training, inference, and RL loops fast and the data plane responsive under heavy concurrent load, with GPUs kept saturated. - Productionize research building blocks - partner with Snowflake Research to turn state-of-the-art training and inference techniques into reliable, composable components customers can run at enterprise scale. QUALIFICATIONS: - 3 + years (Intermediate) | 6+ years (Senior) building and shipping production ML systems - Strong distributed systems and infrastructure foundation - designing scalable, fault-tolerant services and operating them on Kubernetes in production. - Familiarity with GPU and LLM infrastructure - e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM; able to debug across the data, infrastructure, and GPU layers. - Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency. - BS in Computer Science or a related field (MS/PhD a plus). - (Bonus) Hands-on LLM post-training / modeling experience - the strongest candidates pair deep infra skills with real post-training intuition. Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake. How do you want to make your impact? For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com http://careers.snowflake.com

Keep exploring

More Engineering roles

Search all jobs

Privacy choices

Analytics and advertising stay off unless you allow them. Private data stays out.

Read the privacy notice