Senior Staff Software Engineer, DC Infrastructure

Develop software that manages GPU servers and data centers, focusing on diagnostics, observability, automation, repair, and reliability for high-performance GPU clusters. Build tooling and AI agents for hardware diagnosis, remediation, validation, facilities management, power, and liquid cooling, while owning deployment and operational support.

Responsibilities

  • Develop deep-level diagnostics for GPU hardware faults.
  • Build troubleshooting and automation tooling for GPU platforms.
  • Develop AI agents for component diagnosis and hardware remediation.
  • Develop tooling for critical-environment management.
  • Build post-repair validation and testing tools.
  • Own deployment, monitoring, and operational support of tooling.
  • Develop facilities-management automation for power and liquid-cooling systems.

Requirements

  • Software engineering experience.
  • Ability to rapidly develop and ship scalable solutions.
  • Expertise in distributed systems, reliability, and cloud platforms.
  • Strength in Go, Python, Java, or Rust.
  • Experience with Kubernetes, infrastructure as code, and GCP.
  • Strong analytical and problem-solving skills.
  • Experience with Temporal and Kubernetes preferred.
  • Experience with large-scale GPU fleet operations or hyperscale data centers preferred.

Benefits

  • Industry competitive pay
  • Restricted Stock Units
  • Health insurance with HDHP and PPO options
  • Vision insurance
  • Dental insurance
  • Employer contributions to HSA accounts
  • Paid parental leave
  • Paid life insurance
  • Short-term and long-term disability
  • Teladoc
  • 401(k) with 100% match up to 4% of salary
  • Paid time off
  • Paid holidays
  • Cell phone reimbursement
  • Tuition reimbursement
  • Calm app subscription
  • MetLife Legal
  • Company-paid commuter benefit of $300 per month

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available