Senior SRE Engineer
Kronos Research · On-site
This listing is no longer verified as available.
MeritLog keeps this source-backed description for reference. Availability is not verified, and there is no application link here.
Source: the employer's Greenhouse job board. Open the original listing for current details. Availability is not verified for this retained page.
Job details
- Work model
- On-site
- Salary
- Not listed by source
- Location
- Taiwan
What the role asks for
What you'd do
- Manage large-scale Linux environments: troubleshooting and root-cause analysis
- Write maintainable, hand-off-ready Bash / Ansible / Python automation
- On-call for infrastructure, CI/CD, and production service incidents
- Operate HPC clusters (Slurm) along with usage analytics, auditing, and monitoring tools
- Maintain and plan storage for compute environments (Lustre, NAS)
- Manage multi-cloud environments (AWS, Alibaba Cloud, GCP) with Terraform / AWS CDK
- Build and operate Docker (ECS) / Kubernetes (EKS) environments and their deployment workflows
- Operate self-hosted GitLab server and Runner fleet
- Operate CI/CD systems and design deployment pipelines for research and other projects
- Build internal AI platforms (LangChain / LangGraph / Bedrock, Elasticsearch RAG)
- Develop MCP servers, chatbots, AI agents, and similar services
What they're asking for
- **5+ years** of hands-on Linux systems administration and infrastructure operations experienceExperience
- Solid Linux internals knowledge (process / memory / filesystem / networking / systemd / cgroup); able to localize issues even without complete logsSkill
- Strong Bash / Shell scripting skills - able to write maintainable scripts that others can pick upSkill
- Programming ability for data processing, CLI tools, and API services; Python proficiency preferredSkillPreferred
- Solid storage fundamentals with hands-on experience: RAID levels and rebuild trade-offs, filesystem selection, snapshot and backup planning; NAS / shared storage (NFS / SMB) operations experienceSkill
- Experience with at least one major public cloud (AWS / GCP / Alibaba Cloud) and IaC tooling (Terraform / CDK / Ansible)Skill
- Familiar with containerization and orchestration (Docker, Kubernetes)Skill
- CI/CD pipeline design and operations experience (GitLab CI / Jenkins / Airflow)Skill
- Able to own a cross-service subsystem end-to-end: design, implementation, documentation, handoffSkill
- **Strong autonomy**: can drive a problem from discovery, root-cause investigation, decision-making, to delivery with minimal supervision; able to make judgment calls under incomplete information and proactively communicate progress, risks, and rationaleSkill
- **Self-directed**: doesn't wait for tickets - identifies problems worth solving and prioritizes them independentlySkill
- HPC scheduler experience (Slurm / PBS / LSF)SkillPreferred
- Parallel filesystem operations experience (Lustre / GPFS / BeeGFS)SkillPreferred
- Advanced Linux performance analysis (perf, eBPF, ftrace) and kernel parameter tuningSkillPreferred
- DB operations experience (MySQL, ClickHouse)SkillPreferred
- Low-latency network tuning and cross-datacenter link optimizationSkillPreferred
- LLM application development (LangChain, RAG, Agent, MCP)SkillPreferred
- Self-managed Kubernetes experience (Kubespray, kubeadm)SkillPreferred
- GPU server operations (single-node): NVIDIA driver / CUDA toolkit version management, `nvidia-smi` / DCGM monitoring, nvidia-container-toolkit integration, troubleshooting XID / ECC errors and thermal throttlingSkillPreferred
Parsed by MeritLog from the employer’s own posting. The full description follows below.
Job description
Responsibilities Linux Systems & Automation (Core) - Manage large-scale Linux environments: troubleshooting and root-cause analysis - Write maintainable, hand-off-ready Bash / Ansible / Python automation - On-call for infrastructure, CI/CD, and production service incidents HPC Cluster & Storage - Operate HPC clusters (Slurm) along with usage analytics, auditing, and monitoring tools - Maintain and plan storage for compute environments (Lustre, NAS) Cloud & Hybrid Infrastructure - Manage multi-cloud environments (AWS, Alibaba Cloud, GCP) with Terraform / AWS CDK - Build and operate Docker (ECS) / Kubernetes (EKS) environments and their deployment workflows CI/CD & Developer Experience - Operate self-hosted GitLab server and Runner fleet - Operate CI/CD systems and design deployment pipelines for research and other projects GenAI / Internal Platform - Build internal AI platforms (LangChain / LangGraph / Bedrock, Elasticsearch RAG) - Develop MCP servers, chatbots, AI agents, and similar services Requirements - **5+ years** of hands-on Linux systems administration and infrastructure operations experience - Solid Linux internals knowledge (process / memory / filesystem / networking / systemd / cgroup); able to localize issues even without complete logs - Strong Bash / Shell scripting skills - able to write maintainable scripts that others can pick up - Programming ability for data processing, CLI tools, and API services; Python proficiency preferred - Solid storage fundamentals with hands-on experience: RAID levels and rebuild trade-offs, filesystem selection, snapshot and backup planning; NAS / shared storage (NFS / SMB) operations experience - Experience with at least one major public cloud (AWS / GCP / Alibaba Cloud) and IaC tooling (Terraform / CDK / Ansible) - Familiar with containerization and orchestration (Docker, Kubernetes) - CI/CD pipeline design and operations experience (GitLab CI / Jenkins / Airflow) - Able to own a cross-service subsystem end-to-end: design, implementation, documentation, handoff - **Strong autonomy**: can drive a problem from discovery, root-cause investigation, decision-making, to delivery with minimal supervision; able to make judgment calls under incomplete information and proactively communicate progress, risks, and rationale - **Self-directed**: doesn't wait for tickets - identifies problems worth solving and prioritizes them independently Nice to Have - HPC scheduler experience (Slurm / PBS / LSF) - Parallel filesystem operations experience (Lustre / GPFS / BeeGFS) - Advanced Linux performance analysis (perf, eBPF, ftrace) and kernel parameter tuning - DB operations experience (MySQL, ClickHouse) - Low-latency network tuning and cross-datacenter link optimization - LLM application development (LangChain, RAG, Agent, MCP) - Self-managed Kubernetes experience (Kubespray, kubeadm) - GPU server operations (single-node): NVIDIA driver / CUDA toolkit version management, `nvidia-smi` / DCGM monitoring, nvidia-container-toolkit integration, troubleshooting XID / ECC errors and thermal throttling - Experience or familiarity with integrating GPU resources into Slurm: GRES configuration, cgroup-based GPU isolation, user/job-level resource limits
Keep exploring
Available Engineering roles
These current listings are available to explore now.
- Facilities Engineer, Electrical - MemphisSpaceXAI · Not provided by source
- Operations Engineer, Facility Operations - MemphisSpaceXAI · On-site
- Network Security EngineerSpaceXAI · On-site
- Software Engineer - Platform Infrastructure (Rust, C++)SpaceXAI · On-site
- Exceptional Software EngineerSpaceXAI · On-site
- Security Engineer - Detection & Response (Japan, X Money)SpaceXAI · On-site