Mid-Level Site Reliability Engineer
INSTALILY.AI · Not provided by source
MeritLog read this listing from INSTALILY.AI's Greenhouse job board and last checked it on September 12, 2026.
Source: the employer's Greenhouse job board. Open the original listing for current details.
Job details
- Work model
- Not provided by source
- Salary
- $150,000 – $190,000 per year
- Location
- New York, New York
Hiring context
How this role compares at INSTALILY.AI
INSTALILY.AI has 18 live roles in MeritLog’s catalog across 7 job families, and 9 of them are in engineering. 15 of those listings publish a pay range, a disclosure rate of 83%.
This role's posted range of $150,000 – $190,000 per year sits above 73% of the 11 other INSTALILY.AI roles quoted over the same currency and period.
INSTALILY.AI concentrates this hiring in:
Counted across the job boards MeritLog tracks, at the time this page was served. Pay comparisons use only listings that publish a complete range in the same currency and period.
What the role asks for
What you'd do
- Build out Instalily’s Internal Developer Platform - including the developer portal, golden paths, and one-click developer workflows.
- Help design and stand up our Kubernetes platform; operate clusters and workloads as services migrate, focusing on networking, autoscaling, RBAC, and reliability.
- Design, implement, and maintain cloud infrastructure across AWS, GCP, and/or Azure to support the AI agent platform.
- Develop and maintain Infrastructure as Code (IaC) using OpenTofu, focusing on modularity and safe rollouts.
- Build and improve CI/CD pipelines and GitOps workflows (e.g., ArgoCD, Flux) for seamless deployment.
- Implement platform security best practices, including IAM, network segmentation, and policy-as-code.
- Implement and tune logging, monitoring, and alerting using tools such as Datadog, Prometheus, and OpenTelemetry.
- Optimize platform environments for cost, performance, and reliability; participate in on-call rotation.
- Treat developers as customers - gather feedback and measure platform adoption to iterate on the developer experience.
- Mentor more junior engineers and contribute to architectural discussions and technical reviews.
What they're asking for
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.Education
- 3 to 5 years of experience as a platform, cloud, infrastructure, or DevOps engineer.Experience
- Production experience operating Kubernetes (managing clusters, upgrades, and reliability), with greenfield build-out experience being a strong plus.Skill
- Strong hands-on experience with at least one major cloud platform (AWS, GCP, or Azure); multi-cloud experience is preferred.SkillPreferred
- Experience contributing to or building Internal Developer Platforms (golden paths, paved roads) is a strong plus.Skill
- Proficiency with Infrastructure as Code (OpenTofu or Terraform) and GitOps workflows (e.g., ArgoCD, Flux).Skill
- Experience with CI/CD tools and practices such as GitHub Actions, Jenkins, or GitLab CI.Skill
- Solid understanding of cloud networking concepts (VPCs, load balancers, DNS, CDNs, service meshes).Skill
- Working knowledge of cloud security principles, IAM, and compliance frameworks (SOC 2, HIPAA, ISO 27001).Skill
- A product-minded approach to internal tooling, focusing on adoption and feedback loops rather than just uptime.Skill
- Strong problem-solving skills and ability to work in fast-paced, collaborative environments.Skill
- Excellent communication skills to engage effectively with technical and non-technical teams.Skill
- Interest in AI and machine learning infrastructure (GPU workloads, model serving, vector databases) is a plus.SkillPreferred
- Proven product: Customers are live; this isn't a bet on an unproven thesisSkill
- AI-native: In how we build, how we work, and what we sellSkill
- Stage: Early enough to shape how the company scalesSkill
- Global: Based in New York with offices in SF and LondonSkill
- Culture: Sharp, low-ego team that keeps raising the barSkill
- Growth: The learning curve is steepSkill
Parsed by MeritLog from the employer’s own posting. The full description follows below.
Job description
About InstaLILY InstaLILY is an AI products and infrastructure company that puts execution at the frontier of enterprise AI. That work begins with Lily™, the world's first AI Forward Deployed Engineer, which learns how a business works, builds the software it needs, and goes live in days. It does not leave when the work ships; it stays and keeps the software working as the business changes. Lily runs wherever the work happens, in the cloud, on-premise, or at the edge, through InstaLILY's Small Data Center, built with NVIDIA technology. Founded in 2023 by Amit Shah and Sumantro Das, InstaLILY has raised nearly $100 million from Energize Capital, Insight Partners, and Home Depot Ventures. Headquartered in New York, with offices in San Francisco, London, and Toronto, InstaLILY serves leading companies across construction, industrial distribution, logistics, healthcare, and other operationally intensive industries. Learn more at https://instalily.ai/. The Traction Revenue grew 5x over the past year, and Lily has driven over $200M in new annual sales for a single customer. We serve some of the largest operators in our industries, including SRS Distribution (part of The Home Depot family), United Rentals, and Henry Schein, and we work closely with the Google DeepMind and NVIDIA ecosystems. How We Work We work in small teams with real ownership: clear problems, direct access to the customers whose work you're changing, and room to ship. Your code runs in live production systems inside billion-dollar operations, so you see your impact directly. People who do well here want that proximity to the work. We're growing fast, and the people who join now shape what this company becomes. Everything runs on three principles: Customers, Culture, and Code. Role Overview Instalily, a cutting-edge AI startup, is seeking a curious and highly skilled Site Reliability Engineer to help build the Internal Developer Platform (IDP) that powers our AI agent platform. We are redefining how organizations leverage AI using vertical agents, and we’re looking for engineers who think of developer experience as a product. You will build the paved roads, golden paths, and self-service tooling that allow every team at Instalily to ship AI products with speed and confidence. As a Mid-Level Site Reliability Engineer, you will help build and own meaningful pieces of our IDP and the multi-cloud, Kubernetes-based infrastructure beneath it. You’ll partner closely with AI and Software Engineers to turn rough edges into self-service abstractions. You will benefit from world-class mentorship from highly-regarded executive leaders at Internet Retailer 100 brands. Responsibilities • Build out Instalily’s Internal Developer Platform - including the developer portal, golden paths, and one-click developer workflows. • Help design and stand up our Kubernetes platform; operate clusters and workloads as services migrate, focusing on networking, autoscaling, RBAC, and reliability. • Design, implement, and maintain cloud infrastructure across AWS, GCP, and/or Azure to support the AI agent platform. • Develop and maintain Infrastructure as Code (IaC) using OpenTofu, focusing on modularity and safe rollouts. • Build and improve CI/CD pipelines and GitOps workflows (e.g., ArgoCD, Flux) for seamless deployment. • Implement platform security best practices, including IAM, network segmentation, and policy-as-code. • Implement and tune logging, monitoring, and alerting using tools such as Datadog, Prometheus, and OpenTelemetry. • Optimize platform environments for cost, performance, and reliability; participate in on-call rotation. • Treat developers as customers - gather feedback and measure platform adoption to iterate on the developer experience. • Mentor more junior engineers and contribute to architectural discussions and technical reviews. Requirements • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience. • 3 to 5 years of experience as a platform, cloud, infrastructure, or DevOps engineer. • Production experience operating Kubernetes (managing clusters, upgrades, and reliability), with greenfield build-out experience being a strong plus. • Strong hands-on experience with at least one major cloud platform (AWS, GCP, or Azure); multi-cloud experience is preferred. • Experience contributing to or building Internal Developer Platforms (golden paths, paved roads) is a strong plus. • Proficiency with Infrastructure as Code (OpenTofu or Terraform) and GitOps workflows (e.g., ArgoCD, Flux). • Experience with CI/CD tools and practices such as GitHub Actions, Jenkins, or GitLab CI. • Solid understanding of cloud networking concepts (VPCs, load balancers, DNS, CDNs, service meshes). • Working knowledge of cloud security principles, IAM, and compliance frameworks (SOC 2, HIPAA, ISO 27001). • A product-minded approach to internal tooling, focusing on adoption and feedback loops rather than just uptime. • Strong problem-solving skills and ability to work in fast-paced, collaborative environments. • Excellent communication skills to engage effectively with technical and non-technical teams. • Interest in AI and machine learning infrastructure (GPU workloads, model serving, vector databases) is a plus. What You'll Get • Proven product: Customers are live; this isn't a bet on an unproven thesis • AI-native: In how we build, how we work, and what we sell • Stage: Early enough to shape how the company scales • Global: Based in New York with offices in SF and London • Culture: Sharp, low-ego team that keeps raising the bar • Growth: The learning curve is steep Compensation and Benefits • Salary Range: $150,000–$190,000 per year, commensurate with experience • Equity: Stock options awards, and refreshers for top performers • Benefits: Medical, Dental, Vision, 401K, in-office Lunch reimbursement, Wellbeing Stipend, Generous Parental leave, PTO and 10 US Federal Holidays, and more! Quality Over Quantity To ensure a focused, high-quality hiring experience, we kindly ask candidates to limit their applications to 3 open requisitions at any given time. Applying strategically to roles that best align with your skills and career goals gives you the highest chance of standing out. Have you interviewed with us in the past 12 months? We encourage you to reach out directly to your previous interviewer rather than submitting a new application. InstaLILY is committed to providing an inclusive and barrier-free recruitment process. If you require an accommodation, please let us know, and we will work with you to meet your needs.
Keep exploring
More Engineering roles
- Software Engineer, Growth InfrastructureReplit · Hybrid
- Staff Software Engineer, ProductReplit · Hybrid
- Software Engineer, GrowthReplit · Hybrid
- Partnerships Lead, GSIsReplit · Hybrid
- Staff Software Engineer, Money PartnershipsReplit · Hybrid
- Product Security Engineer (PSIRT - Product Security Incident Response Team)Replit · Hybrid