Senior AI Platform Engineer

Operate and harden Kubernetes-based MaaS production environments across CPU nodes, edge ingress, and regional GPU tiers. Own SLOs, observability, runbooks, incident response, rollout safety, capacity planning, automation, and incident debugging from the public API edge to model workers.

Responsibilities

  • Operate and harden Kubernetes-based MaaS production environments
  • Define and own SLOs, alerting, dashboards, runbooks, and incident response
  • Improve rollout safety with canaries, fallback, health-aware routing, maintenance mode, and rollback
  • Plan capacity for GPU utilization, burst traffic, quotas, rate limits, latency, and customer growth
  • Automate operations with Helm, Argo CD, operators, scripts, and self-healing workflows
  • Debug incidents from the public API edge to model workers
  • Partner with runtime and performance engineers on incident resolution

Requirements

  • 6+ years in SRE, platform engineering, or infrastructure engineering for production cloud services
  • Deep Kubernetes experience
  • Experience with Helm, Argo CD, GitOps, CNI, ingress, secrets, storage, and workload scheduling
  • GPU, AI infrastructure, or HPC workload experience strongly preferred
  • Observability experience with Prometheus, VictoriaMetrics, OpenTelemetry, logs, and traces
  • Go, Python, Bash, Linux networking, and production automation experience
  • Ability to design reliable systems with SLOs and operational ownership

Benefits

  • Welfare benefits
  • Inclusive work environment
  • Training and mentoring
  • Developmental opportunities
  • Autonomy and fast growth opportunities

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available