Senior Staff Software Engineer, DC Infrastructure
Develop software that manages GPU servers and data centers, focusing on diagnostics, observability, automation, repair, and reliability for high-performance GPU clusters. Build tooling and AI agents for hardware diagnosis, remediation, validation, facilities management, power, and liquid cooling, while owning deployment and operational support.
Responsibilities
- Develop deep-level diagnostics for GPU hardware faults.
- Build troubleshooting and automation tooling for GPU platforms.
- Develop AI agents for component diagnosis and hardware remediation.
- Develop tooling for critical-environment management.
- Build post-repair validation and testing tools.
- Own deployment, monitoring, and operational support of tooling.
- Develop facilities-management automation for power and liquid-cooling systems.
Requirements
- Software engineering experience.
- Ability to rapidly develop and ship scalable solutions.
- Expertise in distributed systems, reliability, and cloud platforms.
- Strength in Go, Python, Java, or Rust.
- Experience with Kubernetes, infrastructure as code, and GCP.
- Strong analytical and problem-solving skills.
- Experience with Temporal and Kubernetes preferred.
- Experience with large-scale GPU fleet operations or hyperscale data centers preferred.
Benefits
- Industry competitive pay
- Restricted Stock Units
- Health insurance with HDHP and PPO options
- Vision insurance
- Dental insurance
- Employer contributions to HSA accounts
- Paid parental leave
- Paid life insurance
- Short-term and long-term disability
- Teladoc
- 401(k) with 100% match up to 4% of salary
- Paid time off
- Paid holidays
- Cell phone reimbursement
- Tuition reimbursement
- Calm app subscription
- MetLife Legal
- Company-paid commuter benefit of $300 per month