Senior Platform SRE

Own and improve the reliability platform used by engineering teams. Implement observability with OpenTelemetry and distributed tracing, maintain SLOs and error budgets, build deployment and self-healing automation, run chaos experiments, develop CI/CD pipelines, contribute to AI incident tooling, mentor engineers, define SRE standards, guide system design and capacity planning, facilitate post-incident reviews, and track remediation actions.

Responsibilities

  • Implement monitoring and observability with OpenTelemetry and distributed tracing
  • Maintain SLOs, error budgets, and burn-rate tracking
  • Establish 24/7 operational readiness
  • Automate deployments and zero-downtime patching
  • Engineer auto-remediation and automated traffic rerouting
  • Build reliability-focused automation tools and CI/CD pipelines
  • Contribute to the SRE AI agent
  • Mentor junior SREs and Reliability Champions
  • Author and evolve SRE standards
  • Guide teams in system design, capacity planning, and architectural reviews
  • Design SLOs around customer journeys
  • Facilitate blameless post-incident reviews
  • Maintain the Lessons Register
  • Track remediation actions to closure
  • Surface incident patterns quarterly

Requirements

  • 6+ years of experience in the stated technical areas
  • Hands-on OpenTelemetry experience
  • Production experience with Honeycomb, Datadog, Dynatrace, or Grafana
  • Ability to instrument Java or Python services
  • Experience designing SLIs and setting error budgets
  • Experience with multi-window burn-rate alerts
  • Experience building CI/CD pipelines with blue/green or canary releases and automated rollback
  • Kubernetes experience required
  • HashiCorp Nomad experience is advantageous
  • Understanding of cloud networking and infrastructure as code
  • Production-quality Java or Python coding
  • Strong understanding of distributed systems
  • On-call production experience
  • Experience facilitating blameless post-incident reviews
  • Chaos engineering experience
  • Experience writing RFCs and building engineering standards
  • Experience in high-throughput production environments
  • Strong distributed-systems troubleshooting skills

Benefits

  • Hybrid working with 3 days in the office
  • Tailored development programs
  • Mentoring opportunities
  • Clear career progression
  • Sports and social clubs
  • Extra time off for volunteering and community work

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available