Lead SRE (Site Reliability Engineer)
Design, build, and optimize high-availability blockchain data infrastructure while owning CI/CD, orchestration, infrastructure as code, monitoring, alerting, incident management, on-call operations, reliability, observability, automation, fault tolerance, and cost efficiency.
Responsibilities
- Design, build, and optimize a high-availability blockchain data ingestion pipeline.
- Own CI/CD, orchestration, and infrastructure-as-code layers.
- Identify and implement tools for running and monitoring blockchain nodes.
- Assess infrastructure trade-offs to optimize performance, reliability, and cost.
- Build and maintain a public status page and incident management process.
- Define and maintain SRE metrics, logging, and alerting.
- Improve observability, automation, and fault tolerance.
- Contribute to incident response, troubleshooting, and on-call rotations.
Requirements
- 3+ years of experience as an SRE, DevOps Engineer, or similar role.
- Experience running production services against real SLAs, including on-call, incident response, and post-mortems.
- Experience defining and implementing metrics, logging, and alerting.
- Proficiency in Kubernetes, Terraform, Prometheus, Grafana, or equivalent monitoring tools.
- Strong understanding of distributed systems and streaming data pipelines.
- Deep knowledge of AWS, GCP, or bare-metal infrastructure and cost optimization.
- Experience balancing performance, reliability, and cost.
- Programming skills in Python, Go, Rust, or Bash.
- Experience monitoring and running blockchain nodes or working with node providers.
- Willingness to learn blockchain node internals, EVM/SVM data, and new chain integration.
- Previous Web3 experience preferred.
Benefits
- Competitive salary plus token incentives
- Fully remote work
- Flexible hours
- High-impact role with ownership
- Opportunity to build operational foundations for an AI/Web3 company