Senior Site Reliability Engineer
Join the Hexagate team at Chainalysis to build a real-time on-chain detection and response platform. Own reliability as a product capability, operate resilient and observable systems, modernize infrastructure, improve developer experience, and partner with engineering teams to raise operational standards.
Responsibilities
- Define and evolve SLOs, alerting standards, incident response practices, and production readiness expectations.
- Build and improve CI/CD, deployment workflows, and infrastructure automation.
- Design and operate resilient, observable systems for real-time ingestion, detection, and response workloads.
- Lead improvements in scalability, performance, and operational maturity.
- Modernize Kubernetes-based infrastructure, infrastructure as code, and operational patterns.
- Create internal tooling, paved roads, and clear operational standards.
- Partner with backend and platform engineers to reduce toil and prevent recurring failures.
- Participate in incident management, root cause analysis, and systemic fixes.
Requirements
- 5+ years of experience in SRE, infrastructure engineering, platform engineering, or a related role.
- Strong production experience with cloud-native systems.
- Practical experience with Kubernetes, AWS, Terraform, or Pulumi.
- Understanding of observability, monitoring, debugging, and performance tuning for distributed systems.
- Experience with CI/CD, deployment tooling, and operational automation.
- Coding ability in Python, Go, or Rust.
- Sound judgment around reliability trade-offs, failure modes, and safe delivery.
- Collaborative mindset with experience enabling teams and mentoring engineers.
- High ownership and a bias toward durable solutions.
Benefits
- Real problems, real attackers, and real impact.
- Small team with high ownership and close customer proximity.
- Startup pace with the backing and data advantage of a leading blockchain intelligence company.