Staff Site Reliability Engineer - Release Engineering
Define and scale reliability practices across product engineering by architecting SLO and error-budget programs, promoting progressive delivery and automated safety gates, guiding teams toward production readiness, building self-service deployment features, leading critical incident response, and improving release safety for high-velocity development.
Responsibilities
- Lead the expansion of reliability standards across product engineering
- Architect and manage SLO and error-budget frameworks
- Promote progressive delivery and automated safety gates
- Guide product teams toward production readiness
- Collaborate with Platform and Infrastructure teams on self-service platform features
- Direct critical incident response and post-mortem improvements
- Scale safety systems for increased code-change volume
Requirements
- Over 8 years of professional experience in backend systems, SRE, or platform engineering
- Experience designing reliability programs such as service maturity models or SLI frameworks
- Experience building or operating canary rollout systems, metric-gated analysis, or automated rollback infrastructure
- Technical proficiency in software development, preferably with Go or similar systems languages
- Ability to drive organizational change without formal authority
- Technical judgment in high-stakes production scenarios
- Exposure to Kubernetes, service mesh technologies, Prometheus, or ArgoCD
Benefits
- Equity
- Medical, dental, and vision coverage
- 401(k)