Site Reliability Engineer (SRE)

You will help ensure the reliability, scalability, and performance of Unlimit's core platform and services. You will work closely with Engineering and other stakeholders to design, build, and operate cloud-based infrastructure and distributed systems while improving automation, observability, and incident response.

Responsibilities

  • Ensure platform and service availability, resilience, and performance.
  • Own incident management, troubleshooting, escalation handling, and SLA follow-ups.
  • Participate in an on-call rotation and drive reliability improvements.
  • Design, deploy, configure, and manage Linux-based system architecture.
  • Build and support AWS and cloud infrastructure implementations.
  • Design and implement complex technology projects through production rollout and handover.
  • Support Kubernetes workloads and platform components.
  • Automate recurring operational tasks and improve deployment repeatability.
  • Use Terraform and Ansible for infrastructure provisioning and configuration management.
  • Manage CI/CD pipelines across more than 20 repositories.
  • Improve build, release, deployment, monitoring, logging, and alerting reliability.
  • Use metrics and incident learnings to reduce noise and improve detection and recovery times.
  • Produce configuration standards, troubleshooting runbooks, and infrastructure documentation.
  • Contribute to internal standards for consistency, security, and operational maturity.

Requirements

  • 5+ years of Linux systems administration or engineering experience in production environments.
  • Strong knowledge of Linux, Kubernetes, GitLab, Terraform, and Ansible or equivalents.
  • Experience working in Agile Scrum teams.
  • Experience with AWS and/or Google Cloud Platform.
  • Experience with distributed systems design, maintenance, and troubleshooting.
  • Strong scripting or coding ability in Python, Golang, or Bash.
  • Experience with Zabbix, Splunk, Prometheus, Grafana, and PagerDuty or equivalents.
  • Strong English communication and collaboration skills.
  • Working knowledge of PostgreSQL, MongoDB, RabbitMQ, Apache, and Nginx.
  • Ability to learn quickly, work independently, make sound decisions, and collaborate effectively.
  • Strong AI-driven mindset and curiosity about emerging AI technologies.
  • Hands-on experience using AI tools to improve productivity or system performance.
  • Nice to have: experience with highly available, high-volume web services, operational toil reduction, and SLOs, SLIs, error budgets, or formal reliability practices.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available