Infrastructure Engineer: GPU Fleet (HPC)
Architect and own the lifecycle of a high-density GPU fleet supporting production AI training workloads, including firmware validation, bare-metal provisioning, liquid-cooled infrastructure, telemetry, automated remediation, NetBox integration, vendor operations, enterprise infrastructure reviews, and 24/7 incident response.
Responsibilities
- Own the end-to-end health and lifecycle of H200, B200, and B300 GPU nodes.
- Lead operational oversight of high-density liquid-cooled environments.
- Monitor CDU health, secondary loop telemetry, and GPU thermals.
- Architect telemetry using Prometheus, Grafana, and NVIDIA DCGM.
- Trigger automated node draining, reboots, and health validation.
- Migrate inventory to NetBox DCIM.
- Build API integrations for asset tracking, IPAM, and cabling.
- Serve as the primary technical interface for facility operators and MSPs.
- Set SLA and KPI compliance standards.
- Lead technical post-mortems and manage cluster-level outage escalations.
- Support enterprise deal cycles with capacity planning and infrastructure reviews.
- Participate in a 24/7 on-call rotation.
- Own fleet availability and incident response.
Requirements
- Extensive experience managing large-scale HPC environments or production GPU fleets at a hyperscaler, neocloud, or top-tier research facility.
- Deep hands-on experience with H200, B200, or B300 systems.
- Expert knowledge of 400G/800G InfiniBand, ConnectX-7 NDR, ConnectX-8 XDR, NVLink, and NVSwitch architectures.
- Strong Linux internals knowledge.
- Proven proficiency building infrastructure automation using Python or Go.
- Deep experience deploying and scaling DCGM-based telemetry and SNMP-based environmental monitoring.
- Direct experience with Direct-to-Chip systems, coolant chemistry management, or immersion cooling.
- Familiarity with NVIDIA Mission Control.
- Expertise in Intel TDX or NVIDIA RIM attestation flows.
- Prior experience as an initial infrastructure hire.