Senior AI Platform Engineer
Operate and harden Kubernetes-based MaaS production environments across CPU nodes, edge ingress, and regional GPU tiers. Own SLOs, observability, runbooks, incident response, rollout safety, capacity planning, automation, and incident debugging from the public API edge to model workers.
Responsibilities
- Operate and harden Kubernetes-based MaaS production environments
- Define and own SLOs, alerting, dashboards, runbooks, and incident response
- Improve rollout safety with canaries, fallback, health-aware routing, maintenance mode, and rollback
- Plan capacity for GPU utilization, burst traffic, quotas, rate limits, latency, and customer growth
- Automate operations with Helm, Argo CD, operators, scripts, and self-healing workflows
- Debug incidents from the public API edge to model workers
- Partner with runtime and performance engineers on incident resolution
Requirements
- 6+ years in SRE, platform engineering, or infrastructure engineering for production cloud services
- Deep Kubernetes experience
- Experience with Helm, Argo CD, GitOps, CNI, ingress, secrets, storage, and workload scheduling
- GPU, AI infrastructure, or HPC workload experience strongly preferred
- Observability experience with Prometheus, VictoriaMetrics, OpenTelemetry, logs, and traces
- Go, Python, Bash, Linux networking, and production automation experience
- Ability to design reliable systems with SLOs and operational ownership
Benefits
- Welfare benefits
- Inclusive work environment
- Training and mentoring
- Developmental opportunities
- Autonomy and fast growth opportunities