Senior ML Ops Engineer
Confido · On-site
MeritLog read this listing from Confido's Ashby job board and last checked it on September 11, 2026.
Source: the employer's Ashby job board. Open the original listing for current details.
Job details
- Work model
- On-site
- Salary
- $210K - $300K
- Location
- NYC Office
- Company website
- wooded-maxilla-fc0.notion.site
Hiring context
How this role compares at Confido
Confido has 33 live roles in MeritLog’s catalog across 8 job families, and 11 of them are in data & analytics. 33 of those listings publish a pay range, a disclosure rate of 100%.
This role's posted range of $210K - $300K sits above 78% of the 32 other Confido roles quoted over the same currency and period.
Counted across the job boards MeritLog tracks, at the time this page was served. Pay comparisons use only listings that publish a complete range in the same currency and period.
What the role asks for
What you'd do
- Own ML pipelines end to end - experimentation to production - and the infrastructure behind training, inference, and agentic workloads
- Give the AI/ML team a paved road: reproducible environments and fast paths from prototype to production, so they can try new models and agents without fighting the infra
- Stand up the cloud foundation as Infrastructure as Code and the CI/CD that ships ML safely
- Serve and optimize inference and forecasting workloads - latency, throughput, and cost - and the data streams feeding them (e.g. turning a heavy synchronous model call into an async, parallelized one)
- Own the data interface with data engineering: serve the right data to models and agents, and write their outputs back into the platform's data systems for the rest of Confido to use
- Make reliability, observability, security, and privacy the default - and keep model and agent quality measurable in production through online evals and human-in-the-loop review, not just uptime
- 5+ years in MLOps, ML platform, AI infrastructure, or platform engineering - on production ML systems, not pipelines on paper
- You live at the seam of software and infrastructure: equally at home writing production code and standing up cloud infra.
- You've driven a real pipeline end to end and can walk through it: the architecture, the security and cost trade-offs, and what you'd change
- Deep cloud infrastructure understanding, distributed data systems, and IaC - you can boot an environment from scratch, wire CI/CD, and run containerized workloads in production without hand-holding
- Strong Python and comfort in a production app codebase (Ruby, Java) monitoring, security, and cost are instincts, not afterthoughts
- High ownership in a fast-moving startup, and experience productionizing what research/AI teams build
What they're asking for
- LLMOps tooling - tracing, prompt/version management, eval harnessesSkillPreferred
- Inference optimization (vLLM, ONNX, TensorRT) and GPU / spot-instance economicsSkillPreferred
- ML platform and orchestration tooling (MLflow, BentoML, Ray, Airflow)SkillPreferred
- Large-scale data systems (Snowflake, Kafka) and vector databasesSkillPreferred
- Managed ML services (Bedrock, SageMaker, Vertex AI)SkillPreferred
- Multimodal or generative AI in productionSkillPreferred
Parsed by MeritLog from the employer’s own posting. The full description follows below.
Job description
Confido is the AI infrastructure powering modern CPG - the platform that 200+ brands like OLIPOP, Simple Mills, Dr. Squatch, and Tropicana use to run everything from deductions to production planning. Finance, accounting, sales, and operations, unified in one system for the first time. We're growing 5x year over year with a small team in New York City; the people who join now will shape the product, the culture, and the company itself. If you want your work on shelves everywhere - and outsized ownership while you build - we'd love to meet you. THE ROLE Be the first dedicated owner of Confido's ML platform. Our AI/ML team already ships document-understanding, forecasting, and agentic systems into production - on infrastructure we've stood up by hand. You'll own that layer: the pipelines, serving, and cloud foundation that turn models and agents into reliable, cost-efficient production systems, at the scale of hundreds of thousands of documents and heavy LLM/VLM workloads. Location: New York, NY (Relocation supported) WHAT YOU'LL DO - Own ML pipelines end to end - experimentation to production - and the infrastructure behind training, inference, and agentic workloads - Give the AI/ML team a paved road: reproducible environments and fast paths from prototype to production, so they can try new models and agents without fighting the infra - Stand up the cloud foundation as Infrastructure as Code and the CI/CD that ships ML safely - Serve and optimize inference and forecasting workloads - latency, throughput, and cost - and the data streams feeding them (e.g. turning a heavy synchronous model call into an async, parallelized one) - Own the data interface with data engineering: serve the right data to models and agents, and write their outputs back into the platform's data systems for the rest of Confido to use - Make reliability, observability, security, and privacy the default - and keep model and agent quality measurable in production through online evals and human-in-the-loop review, not just uptime WHAT WE'RE LOOKING FOR Required - 5+ years in MLOps, ML platform, AI infrastructure, or platform engineering - on production ML systems, not pipelines on paper - You live at the seam of software and infrastructure: equally at home writing production code and standing up cloud infra. - You've driven a real pipeline end to end and can walk through it: the architecture, the security and cost trade-offs, and what you'd change - Deep cloud infrastructure understanding, distributed data systems, and IaC - you can boot an environment from scratch, wire CI/CD, and run containerized workloads in production without hand-holding - Strong Python and comfort in a production app codebase (Ruby, Java) monitoring, security, and cost are instincts, not afterthoughts - High ownership in a fast-moving startup, and experience productionizing what research/AI teams build Nice to have - LLMOps tooling - tracing, prompt/version management, eval harnesses - Inference optimization (vLLM, ONNX, TensorRT) and GPU / spot-instance economics - ML platform and orchestration tooling (MLflow, BentoML, Ray, Airflow) - Large-scale data systems (Snowflake, Kafka) and vector databases - Managed ML services (Bedrock, SageMaker, Vertex AI) - Multimodal or generative AI in production Our stack: Python · Ruby/Rails · AWS · Terraform · Kubernetes · GitHub Actions · Snowflake · Aurora/RDS · Redis · Kafka - with more of the above added as we scale. Learn more about AI @ Confido here https://wooded-maxilla-fc0.notion.site/AI-ML-Engineering-31b061e1396480bb8b19fef50ad197e5 🌴 PERKS + BENEFITS - Equity - own a meaningful piece of the company you’re helping build - Fully paid health coverage through Aetna - we cover 100% of employee premiums - Top-tier dental and vision coverage through Guardian - Carrot Fertility Pro - comprehensive fertility and family-forming support - 12 weeks of paid parental leave - Unlimited PTO - plus regular 4-day holiday weekends we actually take - 401(k) through Vestwell - Paid relocation support - we’ll help you make the move to NYC - Fully equipped workspace from day one - laptop, monitor, keyboard, and a $200 stipend to personalize your setup - Team perks - catered Friday lunches, team dinners, and unlimited coffee + snacks featuring products from the brands we work with Confido provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.