Senior / Principal Infrastructure Engineer - ML Platform
Roblox · Not provided by source
MeritLog read this listing from Roblox's Greenhouse job board and last checked it on September 12, 2026.
Source: the employer's Greenhouse job board. Open the original listing for current details.
Job details
- Work model
- Not provided by source
- Salary
- $278,530 – $345,040 per year
- Location
- San Mateo, CA, United States
Hiring context
How this role compares at Roblox
Roblox has 232 live roles in MeritLog’s catalog across 11 job families, and 44 of them are in data & analytics. 210 of those listings publish a pay range, a disclosure rate of 91%.
This role's posted range of $278,530 – $345,040 per year sits above 68% of the 196 other Roblox roles quoted over the same currency and period.
Roblox concentrates this hiring in:
Counted across the job boards MeritLog tracks, at the time this page was served. Pay comparisons use only listings that publish a complete range in the same currency and period.
What the role asks for
What you'd do
- Bootstrap and maintain Kubernetes and Cloud infrastructure for ML Platform components--Serving Layer, Metadata Store, Model Registry, and Pipeline Orchestrator.
- Set technical strategy and oversee development of high scale and reliable infrastructure systems.
- Propose and implement new platform tooling to improve time to production for MLEs and Data Scientists across the full ML lifecycle.
- Work on infrastructure projects such as GPU fleet management, hybrid-cloud orchestration, and writing custom Kubernetes controllers and resources.
- Stay abreast of industry trends in machine learning and infrastructure to ensure the adoption of leading-edge technologies and practices.
- Partner across organizations to build tooling, interfaces, and visualizations that make the ML@Roblox a delight to use.
- 6+ years of professional experience and a tool chest of system design experience upon which to draw to build scalable, reliable platforms.
- Deep experience with Kubernetes (K8s) and cluster management at scale - e.g., managing 100s–1000s of nodes, serving 100k+ QPS, and ideally having experience writing custom Kubernetes controllers.
- Strong proficiency in Infrastructure as Code (IaC), specifically using Terraform to bootstrap, manage, and automate cloud infrastructure across AWS, GCP, or similar environments.
- Bachelor's degree in Computer Science, Computer Engineering, Data Science, or a similar technical field or equivalent practical experience.
- Proficient in DevOps tooling such as Docker, Kubernetes, CI/CD systems, and bootstrapping cloud infrastructure (AWS, GCP, etc.)
- Experienced with the end-to-end ML model lifecycle such as model serving, training, model CI/CD, and GPU resources management, and have built ML platform features that are delightful to use.
- An automation advocate: you're passionate about infrastructure-as-code and automating painful manual processes.
- A reliability nut: you love digging into tricky postmortems and identifying weaknesses in complicated systems.
- Passionate about supporting internal partners (data scientists and ML Engineers) to meet and understand their needs.
Parsed by MeritLog from the employer’s own posting. The full description follows below.
Job description
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. ML Platform @ Roblox today supports hundreds of ML use cases and billions of inferences per day across Discovery, Safety, Engine, and much more. As an Infrastructure Engineer on the ML Platform team, you will design, scale, and maintain the foundational infrastructure powering our entire machine learning ecosystem. We are looking for accomplished engineers to spearhead the development of our next-generation ML tooling and platform capabilities. You will: • Bootstrap and maintain Kubernetes and Cloud infrastructure for ML Platform components--Serving Layer, Metadata Store, Model Registry, and Pipeline Orchestrator. • Set technical strategy and oversee development of high scale and reliable infrastructure systems. • Propose and implement new platform tooling to improve time to production for MLEs and Data Scientists across the full ML lifecycle. • Work on infrastructure projects such as GPU fleet management, hybrid-cloud orchestration, and writing custom Kubernetes controllers and resources. • Stay abreast of industry trends in machine learning and infrastructure to ensure the adoption of leading-edge technologies and practices. • Partner across organizations to build tooling, interfaces, and visualizations that make the ML@Roblox a delight to use. You have: • 6+ years of professional experience and a tool chest of system design experience upon which to draw to build scalable, reliable platforms. • Deep experience with Kubernetes (K8s) and cluster management at scale - e.g., managing 100s–1000s of nodes, serving 100k+ QPS, and ideally having experience writing custom Kubernetes controllers. • Strong proficiency in Infrastructure as Code (IaC), specifically using Terraform to bootstrap, manage, and automate cloud infrastructure across AWS, GCP, or similar environments. • Bachelor's degree in Computer Science, Computer Engineering, Data Science, or a similar technical field or equivalent practical experience. You are: • Proficient in DevOps tooling such as Docker, Kubernetes, CI/CD systems, and bootstrapping cloud infrastructure (AWS, GCP, etc.) • Experienced with the end-to-end ML model lifecycle such as model serving, training, model CI/CD, and GPU resources management, and have built ML platform features that are delightful to use. • An automation advocate: you're passionate about infrastructure-as-code and automating painful manual processes. • A reliability nut: you love digging into tricky postmortems and identifying weaknesses in complicated systems. • Passionate about supporting internal partners (data scientists and ML Engineers) to meet and understand their needs. For roles that are based at our headquarters in San Mateo, CA: The starting base pay for this position is as shown below. The actual base pay is dependent upon a variety of job-related factors such as professional background, training, work experience, location, business needs and market demand. Therefore, in some circumstances, the actual salary could fall outside of this expected range. This pay range is subject to change and may be modified in the future. All full-time employees are also eligible for equity compensation and for benefits as described on this page. Annual Salary Range $278,530-$345,040 USD Roles that are based in an office are onsite Tuesday, Wednesday, and Thursday, with optional presence on Monday and Friday (unless otherwise noted). Roblox provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. Roblox also provides reasonable accommodations to candidates with qualifying disabilities or religious beliefs during the recruiting process. For US based roles only, please note the Company may not be able to employ candidates for this role who have United States work authorization related to certain U.S. visa categories, or support future H-1B sponsorship at this time.
Keep exploring
More Data & Analytics roles
- IT Support ApprenticeORCA Service Technologies · On-site
- Head of Commercial - DevelpNimbus · Remote
- Data Scientist (AI Data & LLM Specialist)Eclipse · Remote
- 한국 시장 KOL & Affiliate BD 매니저WOO X · Remote
- Regional Affiliate BD ManagerWOO X · Not provided by source
- Investment Analyst - Summer 2027UVIMCO · Not provided by source