Data Center Site Manager

Manage daily data center operations and lead six shift Operations Engineers while ensuring continuous infrastructure availability. Oversee AI and HPC systems, incident resolution, maintenance, staffing, vendors, operational performance, compliance, and hands-on support during critical events.

Responsibilities

  • Lead daily data center site operations
  • Supervise and manage six shift Operations Engineers
  • Plan manpower and schedule shifts
  • Assign tasks and manage performance
  • Ensure 24x7 operational coverage
  • Escalate incidents and coordinate issue resolution
  • Oversee AI and HPC infrastructure operation and maintenance
  • Establish and improve SOPs, EOPs, and preventive maintenance programs
  • Monitor site health, KPIs, incidents, and infrastructure performance
  • Coordinate hardware installation, rack and stack activities, commissioning, expansion, and lifecycle management
  • Review and approve maintenance activities, change requests, incident reports, and handover records
  • Ensure policy, safety, security, and operational compliance
  • Coordinate with engineering, network, facilities, and vendor teams
  • Participate in on-call duties and provide hands-on operational support

Requirements

  • Bachelor's degree or above in a relevant technical discipline
  • Minimum 5 years of data center operations, IT infrastructure, or HPC/AI infrastructure management experience
  • Minimum 2 years of team leadership or people management experience
  • Experience with 24x7 shift operations in a mission-critical environment
  • Experience with large-scale AI or HPC clusters
  • Knowledge of NVIDIA GB200 and GB300 clusters, GPU servers, x86 servers, storage systems, Ethernet, and InfiniBand networking
  • Familiarity with NVIDIA GPU architecture, NVLink, NVSwitch, and AI cluster deployment
  • Server hardware troubleshooting, firmware management, and hardware lifecycle management
  • Structured cabling knowledge including optical fiber, MPO/LC connectors, DAC, and AOC cabling
  • Linux administration, diagnostics, log analysis, network troubleshooting, scripting, and automation
  • Shift scheduling, incident management, performance management, SOP/EOP development, vendor coordination, communication, and decision-making
  • Willingness to provide hands-on support and participate in on-call duties

Benefits

  • Culture valuing authenticity and diversity of thoughts and backgrounds
  • Inclusive environment with open workspaces and startup spirit
  • Opportunity to network with industrial pioneers and enthusiasts
  • Direct impact on the future of the digital asset industry
  • Involvement in new projects and process or system development
  • Personal accountability, autonomy, fast growth, and learning opportunities
  • Training, mentoring, welfare benefits, and developmental opportunities

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available