Site Reliability Team Lead

Summary

Lead a global SRE team across Israel and Ukraine, combining hands-on technical leadership with people management. Drive reliability, scalability, and operational excellence of a Kubernetes-based cloud platform on GCP and AWS, with focus on automation, CI/CD, observability (Datadog/Prometheus/Grafana), and modern deployment strategies like canary and blue/green.

We are looking for an SRE Team Lead to join our global SRE organization. In this role, you'll lead a team of SREs while partnering with engineering teams globally to improve the reliability, scalability, and operational excellence of our cloud platform.
As an SRE Team Lead, you'll combine hands-on technical leadership with people management, driving engineering-focused reliability initiatives while growing and mentoring a team of SREs across automation, deployment processes, observability, and operational excellence in production.

Responsibilities

  • Team Leadership: Manage, mentor, and grow a global team of SRE engineers across Israel and Ukraine, running regular 1:1s, setting goals, and supporting career development. Own hiring,

onboarding, and performance management for the team.

  • Distributed Team Operations: Keep the team working as one unit across sites, with shared standards, consistent handoffs, and clear ownership so that reliability work is not fragmented by location or timezone.
  • Technical Direction: Set the technical roadmap for the team's reliability, automation, and observability initiatives, and stay hands on enough to guide design decisions and unblock complex problems.
  • Reliability Engineering: Guide the design and implementation of solutions that improve the reliability, availability, and scalability of our production platform, and ensure the team proactively identifies and eliminates operational risks.
  • Automation and Platform Engineering: Prioritize and oversee the build of internal tools and automation that eliminate manual operational work, improve engineering productivity, and streamline production workflows.
  • Observability: Drive the team's roadmap for monitoring, alerting, dashboards, and production visibility, reducing alert fatigue and strengthening operational insight across services.
  • Production Rollouts: Oversee the team's work on deployment processes using modern release strategies such as Canary, Blue/Green, and Feature Flags, ensuring safe and reliable releases.
  • Production Reliability and Incident Response: Act as an escalation point for critical production incidents, guide root cause analysis, and ensure long term preventive improvements are implemented and tracked.
  • On-Call and Operational Excellence: Own the team's on-call rotation and coverage across sites,participate as needed, and drive continuous improvements that reduce operational toil and prevent future incidents.
  • Cloud and Infrastructure: Oversee the team's work on our Kubernetes based cloud platform,CI/CD pipelines, and production infrastructure running on GCP and AWS.
  • Cross Team Partnership: Represent the SRE team in planning and decision making with Software Engineering, DevOps, DBA, and Product leadership, and align the team's priorities with broader engineering goals.

Requirements

  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Infrastructure Engineering, including some experience leading or mentoring other engineers.
  • Hands on experience operating Kubernetes in production environments.
  • Strong experience working with public cloud platforms (GCP or AWS).
  • Strong programming and scripting skills (Python preferred, Go or Bash are a plus).
  • Experience designing and building automation and internal engineering tools.
  • Experience working with CI/CD pipelines and modern deployment methodologies.
  • Hands on experience with observability platforms such as Datadog, Prometheus, or Grafana.
  • Strong understanding of Linux, networking, distributed systems, and cloud native architectures.
  • Excellent troubleshooting, debugging, and root cause analysis skills.
  • Strong communication skills and proven experience working with remote colleagues across sites and timezones, with a genuine interest in growing into people leadership.
  • Fluent English, written and spoken, as the team works across multiple countries

Advantages

  • Prior formal people management or team lead experience.
  • Experience leading or coordinating engineers who are not co-located.
  • Experience with Infrastructure as Code (Terraform, Ansible, etc.).
  • Experience with messaging and distributed technologies such as Kafka, Pub/Sub, or Redis.
  • Experience supporting modern deployment strategies such as Canary, Blue/Green, or Feature Flags.
  • Familiarity with OpenTelemetry and modern observability tooling.
  • Understanding of Site Reliability Engineering principles, including SLIs, SLOs, and error budgets.
  • Experience working in large scale SaaS production environments.
  • Relevant cloud or Kubernetes certifications (GCP, AWS, CKA, CKAD).

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available