Member of Technical Staff - ML Performance
Modal · On-site
MeritLog read this listing from Modal's Ashby job board and last checked it on September 12, 2026.
Source: the employer's Ashby job board. Open the original listing for current details.
Job details
- Work model
- On-site
- Salary
- $200K - $350K
- Location
- New York
- Company website
- modal.com
Hiring context
How this role compares at Modal
Modal has 31 live roles in MeritLog’s catalog across 6 job families, and 8 of them are in data & analytics. 28 of those listings publish a pay range, a disclosure rate of 90%.
This role's posted range of $200K - $350K sits above 88% of the 26 other Modal roles quoted over the same currency and period.
Modal concentrates this hiring in:
Counted across the job boards MeritLog tracks, at the time this page was served. Pay comparisons use only listings that publish a complete range in the same currency and period.
What the role asks for
What they're asking for
- 5+ years of experience writing high-quality, high-performance code.Experience
- Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT).Skill
- Familiarity with Nvidia GPU architecture and CUDA.Skill
- Experience with ML performance engineering (tell us a story about boosting GPU performance - debugging SM occupancy issues, rewriting an algorithm to be compute-bound, eliminating host overhead, etc).Skill
- Nice-to-have: familiarity with low-level operating system foundations (Linux kernel, file systems, containers, etc).Skill
Parsed by MeritLog from the employer’s own posting. The full description follows below.
Job description
ABOUT US: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable https://modal.com/blog/lovable-case-study, Ramp https://modal.com/blog/how-ramp-built-a-full-context-background-coding-agent-on-modal, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C https://modal.com/blog/modal-series-c at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g.,Seaborn https://github.com/mwaskom/seaborn,Luigi https://github.com/spotify/luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. THE ROLE: We are looking for strong engineers with experience in making ML systems performant at scale. If you are interested in contributing to open-source projects and Modal’s container runtime to push language and diffusion models towards higher throughput and lower latency, we’d love to hear from you! REQUIREMENTS: - 5+ years of experience writing high-quality, high-performance code. - Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT). - Familiarity with Nvidia GPU architecture and CUDA. - Experience with ML performance engineering (tell us a story about boosting GPU performance - debugging SM occupancy issues, rewriting an algorithm to be compute-bound, eliminating host overhead, etc). - Nice-to-have: familiarity with low-level operating system foundations (Linux kernel, file systems, containers, etc).
Keep exploring