Senior SRE Engineer
Manage large-scale Linux environments, HPC clusters, storage, multi-cloud infrastructure, CI/CD systems, and internal AI platforms while supporting production operations.
Responsibilities
- Manage large-scale Linux environments, troubleshooting, and root-cause analysis.
- Write maintainable Bash, Ansible, and Python automation.
- Participate in on-call support for infrastructure, CI/CD, and production incidents.
- Operate HPC clusters using Slurm and related analytics, auditing, and monitoring tools.
- Maintain and plan Lustre and NAS storage for compute environments.
- Manage AWS, Alibaba Cloud, and GCP infrastructure using Terraform and AWS CDK.
- Build and operate Docker/ECS and Kubernetes/EKS environments and deployment workflows.
- Operate self-hosted GitLab servers and Runner fleets.
- Design and operate CI/CD systems and deployment pipelines.
- Build internal AI platforms using LangChain, LangGraph, Bedrock, and Elasticsearch RAG.
- Develop MCP servers, chatbots, AI agents, and similar services.
Requirements
- 5+ years of hands-on Linux systems administration and infrastructure operations experience.
- Strong Linux internals knowledge across processes, memory, filesystems, networking, systemd, and cgroups.
- Strong Bash or shell scripting skills.
- Programming ability for data processing, CLI tools, and API services; Python preferred.
- Storage fundamentals including RAID, filesystems, snapshots, backups, NFS, and SMB.
- Experience with a major public cloud and IaC tooling such as Terraform, CDK, or Ansible.
- Familiarity with Docker and Kubernetes.
- CI/CD pipeline design and operations experience with GitLab CI, Jenkins, or Airflow.
- Ability to own cross-service subsystems end-to-end.
- Strong autonomy and problem-solving ability.
- Self-directed approach to identifying and prioritizing problems.