Director, Site Reliability Engineering
This senior engineering leadership role reports to the CTO and sets the vision, operating model, and culture for SRE. The role owns cloud foundations, Kubernetes, CI/CD, observability, infrastructure automation, service ownership standards, reliability practices, and operational maturity across engineering.
Responsibilities
- Lead, coach, and develop a distributed SRE team.
- Define and roll out a Service Ownership & Maturity Framework across engineering.
- Own cloud foundations, Kubernetes, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation.
- Improve service ownership through standards, dashboards, runbooks, alerting, escalation paths, and deployment practices.
- Measure reliability, operational maturity, infrastructure health, and developer productivity.
- Improve deployment automation, resilience, self-healing, disaster recovery, and service reliability.
- Mature incident response, escalation, postmortems, and on-call health.
- Build paved paths and self-service infrastructure to reduce toil and cognitive load.
- Partner with Security, Compliance, Legal, Finance, Procurement, and Corporate IT.
- Evaluate AI-assisted and agentic workflows for infrastructure operations and developer productivity.
Requirements
- 10+ years of experience in SRE, infrastructure, platform, cloud infrastructure, production operations, or related engineering roles.
- 5+ years leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers.
- Experience defining team charters, operating models, roadmaps, success measures, and engineering practices.
- Deep technical judgment across cloud infrastructure, production operations, distributed systems, reliability, automation, and operational risk.
- 3+ years with AWS, GCP, or similar cloud infrastructure.
- 3+ years with Kubernetes, container orchestration, infrastructure-as-code, CI/CD, and deployment safety.
- Strong experience with observability, monitoring, alerting, logging, dashboards, SLOs/SLIs, incident response, postmortems, and on-call practices.
- Experience improving service ownership, operational readiness, and production accountability.
- Pragmatic tooling judgment and ability to operate effectively in a small or mid-sized engineering organization.
- Clear executive communication and ability to partner with a CTO and senior engineering leaders.
Benefits
- Competitive health, dental, and vision coverage.
- Flexible time off and 15 company holidays.
- Paid parental leave and pregnancy disability leave.
- $80 monthly gym reimbursement.
- Life and AD&D insurance up to $50,000.
- Short- and long-term disability coverage.
- 401(k) with 4% match.
- Health and dependent care FSA accounts.
- $250 monthly commuter contribution.
- Health savings account contributions.
- Family building benefits through Kindbody.
- Wellbeing benefits including One Medical, Rightway, and Headspace.
- $1,500 annual learning and development budget.
- Daily lunch and snacks in office.
- Company retreats.
- Lumen-denominated grants.