Senior AI Infrastructure Engineer

You will design, build, and operate the infrastructure that transforms raw GPUs into a programmable, orchestrated pool for AI workloads. You will implement bare-metal provisioning and lifecycle management, develop GPU scheduling and placement strategies, automate provisioning with infrastructure as code, integrate storage solutions for training data, design APIs and cloud-init workflows for automated configuration, optimize GPU compute with CUDA, and work directly with hardware vendors to troubleshoot and improve integrations.

Responsibilities

  • Build and scale a multi-tenant GPU cloud marketplace
  • Design and implement multi-tenancy provisioning and virtualization solutions
  • Transform raw GPUs into a programmable, orchestrated resource pool
  • Implement bare-metal provisioning and lifecycle management
  • Develop GPU scheduling, placement strategies, and fragmentation minimization
  • Automate infrastructure using Terraform or Pulumi and CI/CD pipelines
  • Implement secrets management, configuration management, and observability
  • Design APIs and cloud-init workflows for automated provisioning
  • Integrate and operate storage solutions for AI/ML workloads
  • Collaborate with hardware vendors to troubleshoot and optimize integrations

Requirements

  • Bare-metal provisioning and lifecycle management, including IPMI, Redfish, BMC, PXE, and automated OS deployment
  • GPU scheduling and orchestration with GPU type awareness, memory, topology, placement, and fragmentation minimization
  • Terraform or Pulumi and CI/CD for infrastructure
  • Secrets management and configuration management
  • Observability stack implementation
  • Storage and data infrastructure for AI/ML, including object storage, high-IOPS block storage, and distributed file systems
  • API design and cloud-init for automated provisioning
  • GPU architecture, CUDA, and GPU compute optimization
  • Experience building and scaling cloud infrastructure or distributed systems in production
  • Ability to work with hardware vendors and vendor engineering teams
  • Strong communication skills
  • Preferred: InfiniBand, RoCE, and distributed storage systems such as Ceph, Weka, or VAST Data

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available