Senior Platform SRE
Own and improve the reliability platform used by engineering teams. Implement observability with OpenTelemetry and distributed tracing, maintain SLOs and error budgets, build deployment and self-healing automation, run chaos experiments, develop CI/CD pipelines, contribute to AI incident tooling, mentor engineers, define SRE standards, guide system design and capacity planning, facilitate post-incident reviews, and track remediation actions.
Responsibilities
- Implement monitoring and observability with OpenTelemetry and distributed tracing
- Maintain SLOs, error budgets, and burn-rate tracking
- Establish 24/7 operational readiness
- Automate deployments and zero-downtime patching
- Engineer auto-remediation and automated traffic rerouting
- Build reliability-focused automation tools and CI/CD pipelines
- Contribute to the SRE AI agent
- Mentor junior SREs and Reliability Champions
- Author and evolve SRE standards
- Guide teams in system design, capacity planning, and architectural reviews
- Design SLOs around customer journeys
- Facilitate blameless post-incident reviews
- Maintain the Lessons Register
- Track remediation actions to closure
- Surface incident patterns quarterly
Requirements
- 6+ years of experience in the stated technical areas
- Hands-on OpenTelemetry experience
- Production experience with Honeycomb, Datadog, Dynatrace, or Grafana
- Ability to instrument Java or Python services
- Experience designing SLIs and setting error budgets
- Experience with multi-window burn-rate alerts
- Experience building CI/CD pipelines with blue/green or canary releases and automated rollback
- Kubernetes experience required
- HashiCorp Nomad experience is advantageous
- Understanding of cloud networking and infrastructure as code
- Production-quality Java or Python coding
- Strong understanding of distributed systems
- On-call production experience
- Experience facilitating blameless post-incident reviews
- Chaos engineering experience
- Experience writing RFCs and building engineering standards
- Experience in high-throughput production environments
- Strong distributed-systems troubleshooting skills
Benefits
- Hybrid working with 3 days in the office
- Tailored development programs
- Mentoring opportunities
- Clear career progression
- Sports and social clubs
- Extra time off for volunteering and community work