AI Compute Hardware and Platform Lead
Own the selection, qualification, and lifecycle management of GPU compute platforms, turning accelerator systems into dependable production infrastructure at scale.
Responsibilities
- Own server qualification for AI compute platforms
- Lead procurement strategy and OEM relationships
- Lead BOM validation and platform bring-up
- Manage firmware lifecycles, repair strategies, sparing models, burn-in, acceptance testing, and end-of-life planning
- Establish fleet strategy across compute, interconnect, power, cooling, and management systems
- Build operating models for FRU inventories, RMA execution, depot workflows, and field replacement
- Create service strategies for board swaps, firmware rollback, and spare-part pooling
- Influence vendor roadmaps through technical engagement
- Build lifecycle decision frameworks for hardware expansion, sustainment, refresh, and retirement
Requirements
- Extensive experience building and operating large accelerator fleets
- Experience in hyperscale, cloud, HPC, or advanced systems environments
- Deep knowledge of server architecture, board-level integration, firmware risk, and fleet reliability engineering
- Track record with NVIDIA HGX or DGX, TPU infrastructure, custom accelerator programs, or dense rack-scale compute systems
- Strong understanding of manufacturing quality, service logistics, and infrastructure deployment at scale
- Ability to bridge lab-based engineering with global production operations
- Ability to mentor senior engineers on hardware systems thinking and fleet management