AI Infrastructure Engineer
Architect, deploy, and maintain infrastructure for large-scale AI compute environments in a high-density AI data center, including GPU clusters, high-performance networking, storage, provisioning automation, and infrastructure performance optimization.
Responsibilities
- Deploy and manage large-scale GPU clusters using Kubernetes or Slurm.
- Optimize high-speed, low-latency networking for distributed compute.
- Plan and monitor rack density across AI infrastructure.
- Implement and maintain high-throughput storage systems for GPU-intensive workloads.
- Automate infrastructure provisioning and configuration using Terraform, Ansible, or other Infrastructure as Code tools.
- Troubleshoot and optimize compute, networking, and storage performance in a mission-critical environment.
Requirements
- Degree in Computer Science, Data Engineering, or a related technical field; a master's degree is preferred.
- Strong experience with Linux administration, containerization, and GPU infrastructure.
- Experience with Kubernetes or Slurm.
- Familiarity with CUDA, NCCL, and Triton Inference Server.
- Understanding of NVIDIA GB300 and VR NVL72 Scalable Units.
- Experience in HPC, AI infrastructure, or large-scale distributed compute environments.
- Experience with InfiniBand, RoCE v2, or high-performance networking architectures.
- Familiarity with Lustre, BeeGFS, or WekaIO.
- Experience with Terraform, Ansible, or other Infrastructure as Code frameworks.
- NVIDIA, Kubernetes, or cloud infrastructure certifications.