AI Infrastructure Engineer

Architect, deploy, and maintain infrastructure for large-scale AI compute environments in a high-density AI data center, including GPU clusters, high-performance networking, storage, provisioning automation, and infrastructure performance optimization.

Responsibilities

  • Deploy and manage large-scale GPU clusters using Kubernetes or Slurm.
  • Optimize high-speed, low-latency networking for distributed compute.
  • Plan and monitor rack density across AI infrastructure.
  • Implement and maintain high-throughput storage systems for GPU-intensive workloads.
  • Automate infrastructure provisioning and configuration using Terraform, Ansible, or other Infrastructure as Code tools.
  • Troubleshoot and optimize compute, networking, and storage performance in a mission-critical environment.

Requirements

  • Degree in Computer Science, Data Engineering, or a related technical field; a master's degree is preferred.
  • Strong experience with Linux administration, containerization, and GPU infrastructure.
  • Experience with Kubernetes or Slurm.
  • Familiarity with CUDA, NCCL, and Triton Inference Server.
  • Understanding of NVIDIA GB300 and VR NVL72 Scalable Units.
  • Experience in HPC, AI infrastructure, or large-scale distributed compute environments.
  • Experience with InfiniBand, RoCE v2, or high-performance networking architectures.
  • Familiarity with Lustre, BeeGFS, or WekaIO.
  • Experience with Terraform, Ansible, or other Infrastructure as Code frameworks.
  • NVIDIA, Kubernetes, or cloud infrastructure certifications.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available