Senior SRE Engineer

Manage large-scale Linux environments, HPC clusters, storage, multi-cloud infrastructure, CI/CD systems, and internal AI platforms while supporting production operations.

Responsibilities

  • Manage large-scale Linux environments, troubleshooting, and root-cause analysis.
  • Write maintainable Bash, Ansible, and Python automation.
  • Participate in on-call support for infrastructure, CI/CD, and production incidents.
  • Operate HPC clusters using Slurm and related analytics, auditing, and monitoring tools.
  • Maintain and plan Lustre and NAS storage for compute environments.
  • Manage AWS, Alibaba Cloud, and GCP infrastructure using Terraform and AWS CDK.
  • Build and operate Docker/ECS and Kubernetes/EKS environments and deployment workflows.
  • Operate self-hosted GitLab servers and Runner fleets.
  • Design and operate CI/CD systems and deployment pipelines.
  • Build internal AI platforms using LangChain, LangGraph, Bedrock, and Elasticsearch RAG.
  • Develop MCP servers, chatbots, AI agents, and similar services.

Requirements

  • 5+ years of hands-on Linux systems administration and infrastructure operations experience.
  • Strong Linux internals knowledge across processes, memory, filesystems, networking, systemd, and cgroups.
  • Strong Bash or shell scripting skills.
  • Programming ability for data processing, CLI tools, and API services; Python preferred.
  • Storage fundamentals including RAID, filesystems, snapshots, backups, NFS, and SMB.
  • Experience with a major public cloud and IaC tooling such as Terraform, CDK, or Ansible.
  • Familiarity with Docker and Kubernetes.
  • CI/CD pipeline design and operations experience with GitLab CI, Jenkins, or Airflow.
  • Ability to own cross-service subsystems end-to-end.
  • Strong autonomy and problem-solving ability.
  • Self-directed approach to identifying and prioritizing problems.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available