Site Reliability Engineer/Infrastructure

Engineer the reliability of Busha's Platform API by collaborating with DevOps and backend engineering teams. The role combines distributed systems expertise with hands-on Go development to define SLOs, lead incident response, build automation, and embed resilience patterns into production services.

Responsibilities

  • Act as Incident Commander during major incidents affecting payments, trading, or compliance systems.
  • Design observability strategies using Grafana, Sentry, CloudWatch, Prometheus, and OpenTelemetry.
  • Instrument Go services with metrics and distributed traces.
  • Define, measure, and defend SLOs and Error Budgets with product and engineering teams.
  • Build internal tooling, automation platforms, and self-healing mechanisms.
  • Contribute circuit breakers, retries, and backpressure patterns to backend services.
  • Partner with backend teams to build reliable, scalable, and observable services.
  • Analyze traffic and performance to model future capacity needs.
  • Conduct load testing and chaos engineering experiments for financial and compliance workflows.

Requirements

  • At least 4 years of SRE or Backend Engineering experience with strong Go proficiency.
  • Deep understanding of distributed systems architecture and design patterns.
  • Strong command of microservices and event-driven architectures.
  • Hands-on AWS or GCP experience and infrastructure as code knowledge.
  • Experience running production workloads and troubleshooting infrastructure issues.
  • Experience designing observability strategies and implementing monitoring and alerting.
  • Familiarity with PostgreSQL, ClickHouse, RabbitMQ, and Kafka in high-throughput environments.

Benefits

  • Progressive hybrid work policy
  • Competitive salary
  • Learning and development plan
  • Health insurance and pension
  • Work tools and gadget that works for you

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available