Senior Site Reliability Engineer

Ensure Hyperbolic's GPU marketplace and AI infrastructure operate with high reliability, performance, and security. Define and maintain SLOs, design monitoring and alerting, build capacity-management automation, implement progressive rollouts and rollbacks, lead incident response and post-mortems, and develop infrastructure security and compliance practices.

Responsibilities

  • Define and maintain service level objectives and service level agreements
  • Build and operate incident response systems and lead on-call rotations
  • Manage capacity planning and resource allocation across distributed GPU networks
  • Implement progressive rollouts, canary deployments, feature flags, and automated rollbacks
  • Design monitoring and alerting systems to provide deep infrastructure visibility
  • Automate capacity management and resource allocation
  • Lead post-mortem processes and drive improvements to reduce MTTR
  • Harden infrastructure security including tenant and workload isolation
  • Implement secrets and key management and certificate rotation
  • Develop compliance frameworks and security best practices

Requirements

  • Expertise in site reliability engineering with experience defining, monitoring, and maintaining SLOs and SLAs
  • Strong background in capacity planning, forecasting, resource allocation, and cost optimization for distributed systems
  • Experience in incident response, on-call rotations, and post-mortem processes with measurable MTTR improvements
  • Knowledge of deployment systems including progressive rollouts, canary deployments, feature flags, and automated rollback mechanisms
  • Proficiency with observability tools and practices including metrics, logging, tracing, and alerting using Prometheus, Grafana, ELK, or similar
  • Strong understanding of infrastructure security including tenant isolation, workload isolation, and network segmentation
  • Experience with secrets management, key management systems, certificate management, and secure credential rotation
  • Knowledge of compliance frameworks and security best practices for cloud platforms such as SOC 2 and ISO 27001
  • Experience with infrastructure-as-code, configuration management, and CI/CD pipelines
  • Preferred experience with GPU infrastructure, AI/ML platforms, distributed systems, container security, chaos engineering, and cost optimization

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available