Site Reliability Team Lead
Summary
Lead a global SRE team across Israel and Ukraine, combining hands-on technical leadership with people management. Drive reliability, scalability, and operational excellence of a Kubernetes-based cloud platform on GCP and AWS, with focus on automation, CI/CD, observability (Datadog/Prometheus/Grafana), and modern deployment strategies like canary and blue/green.
We are looking for an SRE Team Lead to join our global SRE organization. In this role, you'll lead a team of SREs while partnering with engineering teams globally to improve the reliability, scalability, and operational excellence of our cloud platform.
As an SRE Team Lead, you'll combine hands-on technical leadership with people management, driving engineering-focused reliability initiatives while growing and mentoring a team of SREs across automation, deployment processes, observability, and operational excellence in production.
Responsibilities
- Team Leadership: Manage, mentor, and grow a global team of SRE engineers across Israel and Ukraine, running regular 1:1s, setting goals, and supporting career development. Own hiring,
onboarding, and performance management for the team.
- Distributed Team Operations: Keep the team working as one unit across sites, with shared standards, consistent handoffs, and clear ownership so that reliability work is not fragmented by location or timezone.
- Technical Direction: Set the technical roadmap for the team's reliability, automation, and observability initiatives, and stay hands on enough to guide design decisions and unblock complex problems.
- Reliability Engineering: Guide the design and implementation of solutions that improve the reliability, availability, and scalability of our production platform, and ensure the team proactively identifies and eliminates operational risks.
- Automation and Platform Engineering: Prioritize and oversee the build of internal tools and automation that eliminate manual operational work, improve engineering productivity, and streamline production workflows.
- Observability: Drive the team's roadmap for monitoring, alerting, dashboards, and production visibility, reducing alert fatigue and strengthening operational insight across services.
- Production Rollouts: Oversee the team's work on deployment processes using modern release strategies such as Canary, Blue/Green, and Feature Flags, ensuring safe and reliable releases.
- Production Reliability and Incident Response: Act as an escalation point for critical production incidents, guide root cause analysis, and ensure long term preventive improvements are implemented and tracked.
- On-Call and Operational Excellence: Own the team's on-call rotation and coverage across sites,participate as needed, and drive continuous improvements that reduce operational toil and prevent future incidents.
- Cloud and Infrastructure: Oversee the team's work on our Kubernetes based cloud platform,CI/CD pipelines, and production infrastructure running on GCP and AWS.
- Cross Team Partnership: Represent the SRE team in planning and decision making with Software Engineering, DevOps, DBA, and Product leadership, and align the team's priorities with broader engineering goals.
Requirements
- 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Infrastructure Engineering, including some experience leading or mentoring other engineers.
- Hands on experience operating Kubernetes in production environments.
- Strong experience working with public cloud platforms (GCP or AWS).
- Strong programming and scripting skills (Python preferred, Go or Bash are a plus).
- Experience designing and building automation and internal engineering tools.
- Experience working with CI/CD pipelines and modern deployment methodologies.
- Hands on experience with observability platforms such as Datadog, Prometheus, or Grafana.
- Strong understanding of Linux, networking, distributed systems, and cloud native architectures.
- Excellent troubleshooting, debugging, and root cause analysis skills.
- Strong communication skills and proven experience working with remote colleagues across sites and timezones, with a genuine interest in growing into people leadership.
- Fluent English, written and spoken, as the team works across multiple countries
Advantages
- Prior formal people management or team lead experience.
- Experience leading or coordinating engineers who are not co-located.
- Experience with Infrastructure as Code (Terraform, Ansible, etc.).
- Experience with messaging and distributed technologies such as Kafka, Pub/Sub, or Redis.
- Experience supporting modern deployment strategies such as Canary, Blue/Green, or Feature Flags.
- Familiarity with OpenTelemetry and modern observability tooling.
- Understanding of Site Reliability Engineering principles, including SLIs, SLOs, and error budgets.
- Experience working in large scale SaaS production environments.
- Relevant cloud or Kubernetes certifications (GCP, AWS, CKA, CKAD).