Site Reliability Engineer
Operate and scale production infrastructure with a focus on reliability, automation, and security. Manage Kubernetes clusters, build declarative Terraform infrastructure, design GitOps CI/CD workflows, implement observability, troubleshoot networking and storage, automate operational workflows, participate in incident response, and improve reliability through postmortems and threat modeling.
Responsibilities
- Operate production Kubernetes clusters and manage system components
- Build and maintain declarative infrastructure using Terraform or similar tools
- Design and manage CI/CD workflows for infrastructure and applications with GitOps tools
- Implement and maintain observability using metrics, logs, and dashboards
- Diagnose and troubleshoot networking and storage issues in distributed systems
- Automate operational workflows using Python, Go, or Bash
- Participate in on-call rotations, respond to incidents, and drive postmortems
- Apply security-minded design, including least privilege and threat modeling
Requirements
- Experience operating production Kubernetes clusters
- Experience with Terraform or similar infrastructure-as-code tools
- Experience with GitOps and ArgoCD, including ApplicationSets or similar patterns
- Experience designing CI/CD pipelines
- Experience with Prometheus, Loki, Mimir, Grafana, or CloudWatch
- Proficiency in Linux and shell scripting
- Programming experience in Python or Go
- Experience with AWS, GCP, or Azure
- Experience with on-call rotations and incident response
- Experience designing secure infrastructure and contributing to threat models
Benefits
- Remote-first global workforce and New York office
- Annual company offsite and team onsites
- Professional reimbursement program
- Medical, dental, and vision coverage in the US and some other countries
- 401k retirement plan with company match in the US
- Wellness stipend
- Home office setup and ergonomic equipment program