Site Reliability Engineer/Infrastructure
Engineer the reliability of Busha's Platform API by collaborating with DevOps and backend engineering teams. The role combines distributed systems expertise with hands-on Go development to define SLOs, lead incident response, build automation, and embed resilience patterns into production services.
Responsibilities
- Act as Incident Commander during major incidents affecting payments, trading, or compliance systems.
- Design observability strategies using Grafana, Sentry, CloudWatch, Prometheus, and OpenTelemetry.
- Instrument Go services with metrics and distributed traces.
- Define, measure, and defend SLOs and Error Budgets with product and engineering teams.
- Build internal tooling, automation platforms, and self-healing mechanisms.
- Contribute circuit breakers, retries, and backpressure patterns to backend services.
- Partner with backend teams to build reliable, scalable, and observable services.
- Analyze traffic and performance to model future capacity needs.
- Conduct load testing and chaos engineering experiments for financial and compliance workflows.
Requirements
- At least 4 years of SRE or Backend Engineering experience with strong Go proficiency.
- Deep understanding of distributed systems architecture and design patterns.
- Strong command of microservices and event-driven architectures.
- Hands-on AWS or GCP experience and infrastructure as code knowledge.
- Experience running production workloads and troubleshooting infrastructure issues.
- Experience designing observability strategies and implementing monitoring and alerting.
- Familiarity with PostgreSQL, ClickHouse, RabbitMQ, and Kafka in high-throughput environments.
Benefits
- Progressive hybrid work policy
- Competitive salary
- Learning and development plan
- Health insurance and pension
- Work tools and gadget that works for you