Back to search
Listing unavailableEngineeringOn-site

Senior SRE Engineer

Kronos Research · On-site

This listing is no longer verified as available.

MeritLog keeps this source-backed description for reference. Availability is not verified, and there is no application link here.

Last seen by MeritLog September 7, 2026Source: GreenhouseSource version: greenhouse-job-board-v1

Source: the employer's Greenhouse job board. Open the original listing for current details. Availability is not verified for this retained page.

Job details

Work model
On-site
Salary
Not listed by source
Location
Taiwan

What the role asks for

What you'd do

  • Manage large-scale Linux environments: troubleshooting and root-cause analysis
  • Write maintainable, hand-off-ready Bash / Ansible / Python automation
  • On-call for infrastructure, CI/CD, and production service incidents
  • Operate HPC clusters (Slurm) along with usage analytics, auditing, and monitoring tools
  • Maintain and plan storage for compute environments (Lustre, NAS)
  • Manage multi-cloud environments (AWS, Alibaba Cloud, GCP) with Terraform / AWS CDK
  • Build and operate Docker (ECS) / Kubernetes (EKS) environments and their deployment workflows
  • Operate self-hosted GitLab server and Runner fleet
  • Operate CI/CD systems and design deployment pipelines for research and other projects
  • Build internal AI platforms (LangChain / LangGraph / Bedrock, Elasticsearch RAG)
  • Develop MCP servers, chatbots, AI agents, and similar services

What they're asking for

  • **5+ years** of hands-on Linux systems administration and infrastructure operations experienceExperience
  • Solid Linux internals knowledge (process / memory / filesystem / networking / systemd / cgroup); able to localize issues even without complete logsSkill
  • Strong Bash / Shell scripting skills - able to write maintainable scripts that others can pick upSkill
  • Programming ability for data processing, CLI tools, and API services; Python proficiency preferredSkillPreferred
  • Solid storage fundamentals with hands-on experience: RAID levels and rebuild trade-offs, filesystem selection, snapshot and backup planning; NAS / shared storage (NFS / SMB) operations experienceSkill
  • Experience with at least one major public cloud (AWS / GCP / Alibaba Cloud) and IaC tooling (Terraform / CDK / Ansible)Skill
  • Familiar with containerization and orchestration (Docker, Kubernetes)Skill
  • CI/CD pipeline design and operations experience (GitLab CI / Jenkins / Airflow)Skill
  • Able to own a cross-service subsystem end-to-end: design, implementation, documentation, handoffSkill
  • **Strong autonomy**: can drive a problem from discovery, root-cause investigation, decision-making, to delivery with minimal supervision; able to make judgment calls under incomplete information and proactively communicate progress, risks, and rationaleSkill
  • **Self-directed**: doesn't wait for tickets - identifies problems worth solving and prioritizes them independentlySkill
  • HPC scheduler experience (Slurm / PBS / LSF)SkillPreferred
  • Parallel filesystem operations experience (Lustre / GPFS / BeeGFS)SkillPreferred
  • Advanced Linux performance analysis (perf, eBPF, ftrace) and kernel parameter tuningSkillPreferred
  • DB operations experience (MySQL, ClickHouse)SkillPreferred
  • Low-latency network tuning and cross-datacenter link optimizationSkillPreferred
  • LLM application development (LangChain, RAG, Agent, MCP)SkillPreferred
  • Self-managed Kubernetes experience (Kubespray, kubeadm)SkillPreferred
  • GPU server operations (single-node): NVIDIA driver / CUDA toolkit version management, `nvidia-smi` / DCGM monitoring, nvidia-container-toolkit integration, troubleshooting XID / ECC errors and thermal throttlingSkillPreferred

Parsed by MeritLog from the employer’s own posting. The full description follows below.

Job description

Responsibilities Linux Systems & Automation (Core) - Manage large-scale Linux environments: troubleshooting and root-cause analysis - Write maintainable, hand-off-ready Bash / Ansible / Python automation - On-call for infrastructure, CI/CD, and production service incidents HPC Cluster & Storage - Operate HPC clusters (Slurm) along with usage analytics, auditing, and monitoring tools - Maintain and plan storage for compute environments (Lustre, NAS) Cloud & Hybrid Infrastructure - Manage multi-cloud environments (AWS, Alibaba Cloud, GCP) with Terraform / AWS CDK - Build and operate Docker (ECS) / Kubernetes (EKS) environments and their deployment workflows CI/CD & Developer Experience - Operate self-hosted GitLab server and Runner fleet - Operate CI/CD systems and design deployment pipelines for research and other projects GenAI / Internal Platform - Build internal AI platforms (LangChain / LangGraph / Bedrock, Elasticsearch RAG) - Develop MCP servers, chatbots, AI agents, and similar services Requirements - **5+ years** of hands-on Linux systems administration and infrastructure operations experience - Solid Linux internals knowledge (process / memory / filesystem / networking / systemd / cgroup); able to localize issues even without complete logs - Strong Bash / Shell scripting skills - able to write maintainable scripts that others can pick up - Programming ability for data processing, CLI tools, and API services; Python proficiency preferred - Solid storage fundamentals with hands-on experience: RAID levels and rebuild trade-offs, filesystem selection, snapshot and backup planning; NAS / shared storage (NFS / SMB) operations experience - Experience with at least one major public cloud (AWS / GCP / Alibaba Cloud) and IaC tooling (Terraform / CDK / Ansible) - Familiar with containerization and orchestration (Docker, Kubernetes) - CI/CD pipeline design and operations experience (GitLab CI / Jenkins / Airflow) - Able to own a cross-service subsystem end-to-end: design, implementation, documentation, handoff - **Strong autonomy**: can drive a problem from discovery, root-cause investigation, decision-making, to delivery with minimal supervision; able to make judgment calls under incomplete information and proactively communicate progress, risks, and rationale - **Self-directed**: doesn't wait for tickets - identifies problems worth solving and prioritizes them independently Nice to Have - HPC scheduler experience (Slurm / PBS / LSF) - Parallel filesystem operations experience (Lustre / GPFS / BeeGFS) - Advanced Linux performance analysis (perf, eBPF, ftrace) and kernel parameter tuning - DB operations experience (MySQL, ClickHouse) - Low-latency network tuning and cross-datacenter link optimization - LLM application development (LangChain, RAG, Agent, MCP) - Self-managed Kubernetes experience (Kubespray, kubeadm) - GPU server operations (single-node): NVIDIA driver / CUDA toolkit version management, `nvidia-smi` / DCGM monitoring, nvidia-container-toolkit integration, troubleshooting XID / ECC errors and thermal throttling - Experience or familiarity with integrating GPU resources into Slurm: GRES configuration, cgroup-based GPU isolation, user/job-level resource limits

Keep exploring

Available Engineering roles

These current listings are available to explore now.

Search all jobs

Privacy choices

Analytics and advertising stay off unless you allow them. Private data stays out.

Read the privacy notice