Site Reliability Engineer
Join zerohash's Platform team to co-own reliable production services, deliver and operate services, improve operational performance through metrics and benchmarks, maintain CI/CD tooling, participate in weekly on-call rotations, and manage scalable infrastructure.
Responsibilities
- Co-own production services and ensure reliable, scalable operation
- Deliver new features and services while operating existing services
- Drive operational improvements through metric-driven collection and analysis
- Develop and maintain application performance benchmarks
- Improve operational efficiency through code releases and performance monitoring
- Maintain tooling, automation, monitoring, workflow management, and CI/CD
- Participate in the weekly on-call rotation
- Manage and scale infrastructure systems
- Improve CI/CD pipelines and AWS infrastructure
- Implement blue/green and canary deployments
Requirements
- Extensive AWS infrastructure deployment, management, and troubleshooting experience
- Production container deployment lifecycle experience using self-managed Kubernetes, ECS, or EKS
- Experience writing custom tools and teaching others how to use them
- CI/CD knowledge and custom production deployment tooling experience
- Distributed Linux systems troubleshooting and request tracing across applications, systems, and networks
- Automation experience and proficiency in at least two programming languages
- Strong written and spoken communication skills
- Ability to collaborate effectively and work independently
Benefits
- Equity opportunity
- Maternity leave
- Paternity leave
- WeWork Membership
- WFH yearly stipend
- L&D stipend after 6 months