Blockchain Site Reliability Engineer
This role focuses on ensuring the reliability, availability, and performance of blockchain nodes and related infrastructure. The engineer will monitor and troubleshoot production systems, develop automation, manage incidents, participate in on-call rotations, and collaborate with protocol engineers and open-source communities on upgrades and long-term stability.
Responsibilities
- Deploy, monitor, and maintain blockchain nodes across multiple networks.
- Manage incidents, troubleshoot node failures, and maintain system reliability and uptime.
- Develop automation and maintenance tools using Golang, Shell, Python, and related technologies.
- Build and maintain monitoring, alerting, and logging systems.
- Collaborate with engineering teams and solution architects on reliability improvements.
- Participate in the on-call rotation for incident response and resolution.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience.
- Strong Linux system administration skills, including networking, performance tuning, debugging, and security.
- Strong programming skills in at least one mainstream language such as Golang, Python, JavaScript, or Rust.
- Experience with monitoring and alerting tools such as Prometheus, Grafana, or ELK.
- Strong problem-solving skills and ability to respond quickly under pressure.
- Solid technical documentation skills.
- Hands-on experience with blockchain node deployment, maintenance, and upgrades.
- Familiarity with mainstream blockchain protocols such as Ethereum, Cosmos, Polkadot, or Solana.
- Experience with Docker or Kubernetes.
- Knowledge of smart contracts, Web3 RPC, or Solidity is a plus.