Director, Site Reliability Engineering

This senior engineering leadership role reports to the CTO and sets the vision, operating model, and culture for SRE. The role owns cloud foundations, Kubernetes, CI/CD, observability, infrastructure automation, service ownership standards, reliability practices, and operational maturity across engineering.

Responsibilities

  • Lead, coach, and develop a distributed SRE team.
  • Define and roll out a Service Ownership & Maturity Framework across engineering.
  • Own cloud foundations, Kubernetes, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation.
  • Improve service ownership through standards, dashboards, runbooks, alerting, escalation paths, and deployment practices.
  • Measure reliability, operational maturity, infrastructure health, and developer productivity.
  • Improve deployment automation, resilience, self-healing, disaster recovery, and service reliability.
  • Mature incident response, escalation, postmortems, and on-call health.
  • Build paved paths and self-service infrastructure to reduce toil and cognitive load.
  • Partner with Security, Compliance, Legal, Finance, Procurement, and Corporate IT.
  • Evaluate AI-assisted and agentic workflows for infrastructure operations and developer productivity.

Requirements

  • 10+ years of experience in SRE, infrastructure, platform, cloud infrastructure, production operations, or related engineering roles.
  • 5+ years leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers.
  • Experience defining team charters, operating models, roadmaps, success measures, and engineering practices.
  • Deep technical judgment across cloud infrastructure, production operations, distributed systems, reliability, automation, and operational risk.
  • 3+ years with AWS, GCP, or similar cloud infrastructure.
  • 3+ years with Kubernetes, container orchestration, infrastructure-as-code, CI/CD, and deployment safety.
  • Strong experience with observability, monitoring, alerting, logging, dashboards, SLOs/SLIs, incident response, postmortems, and on-call practices.
  • Experience improving service ownership, operational readiness, and production accountability.
  • Pragmatic tooling judgment and ability to operate effectively in a small or mid-sized engineering organization.
  • Clear executive communication and ability to partner with a CTO and senior engineering leaders.

Benefits

  • Competitive health, dental, and vision coverage.
  • Flexible time off and 15 company holidays.
  • Paid parental leave and pregnancy disability leave.
  • $80 monthly gym reimbursement.
  • Life and AD&D insurance up to $50,000.
  • Short- and long-term disability coverage.
  • 401(k) with 4% match.
  • Health and dependent care FSA accounts.
  • $250 monthly commuter contribution.
  • Health savings account contributions.
  • Family building benefits through Kindbody.
  • Wellbeing benefits including One Medical, Rightway, and Headspace.
  • $1,500 annual learning and development budget.
  • Daily lunch and snacks in office.
  • Company retreats.
  • Lumen-denominated grants.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available