Member of Technical Staff - GPU Infrastructure
Design, deploy, optimize, and support large-scale GPU infrastructure for customers, including GPU clusters, orchestration, high-performance networking, parallel filesystems, system performance, infrastructure troubleshooting, documentation, and operational support.
Responsibilities
- Partner with clients to understand workload requirements and design GPU cluster architectures
- Create technical proposals and capacity plans for clusters ranging from 100 to 10,000+ GPUs
- Develop deployment strategies for LLM training, inference, and HPC workloads
- Present architectural recommendations to technical and executive stakeholders
- Deploy and configure SLURM and Kubernetes
- Implement InfiniBand, RoCE, and NVLink networking
- Optimize GPU utilization, memory management, and inter-node communication
- Configure Lustre, BeeGFS, and GPFS filesystems
- Tune kernel and CUDA configurations
- Resolve customer infrastructure issues
- Implement monitoring, alerting, and automated remediation
- Provide 24/7 on-call support for critical customer deployments
- Create runbooks and documentation
Requirements
- 3+ years of hands-on experience with GPU clusters and HPC environments
- Deep expertise with SLURM and Kubernetes in production GPU settings
- Experience with InfiniBand configuration and troubleshooting
- Strong understanding of NVIDIA GPU architecture, CUDA, and drivers
- Experience with Ansible and Terraform
- Proficiency in Python, Bash, and systems programming
- Customer-facing technical leadership experience
- Experience with NVIDIA drivers, Fabric Manager, and DCGM
- Experience configuring Docker, Containerd, and Enroot for GPUs
- Linux kernel tuning and performance optimization
- AI workload network topology design
- Knowledge of power and cooling requirements for high-density GPU deployments
Benefits
- Equity incentives