Member of Technical Staff - Training Platform
Prime Intellect is hiring a Member of Technical Staff to develop hosted training infrastructure and platform experiences spanning Kubernetes orchestration, GPU scheduling, autoscaling, observability, backend services, monitoring tools, and frontend product interfaces.
Responsibilities
- Design and operate Kubernetes-based training and inference orchestration across multi-cluster, multi-cloud GPU fleets
- Build and maintain Helm charts for reproducible training stacks
- Develop Python control-plane agents that watch pods and synchronize cluster state
- Implement scheduling and autoscaling for heterogeneous GPU hardware
- Operate GitOps workflows and build model caches, checkpoint pipelines, and shared storage
- Operate observability systems and improve GPU cluster debugging
- Build job submission, live monitoring, logging, metrics, and model management surfaces
- Develop FastAPI backend services and REST APIs
- Build real-time monitoring and debugging tools
- Ship product interfaces with Next.js, React, and TypeScript
- Interface with trainers, inference servers, and environment servers
- Productize new training capabilities, model architectures, and reinforcement learning modes
Requirements
- Strong knowledge of open model families and fine-tuning techniques including LoRA, QLoRA, full fine-tuning, RLHF, and RLAIF
- Familiarity with inference engines such as vLLM, SGLang, and TensorRT-LLM
- Understanding of GPU hardware tradeoffs and distributed training fundamentals
- Strong Kubernetes operations experience with Helm, CRDs, operators, KEDA, gang scheduling, and GPU operators
- Production cluster debugging experience
- Cloud platform experience, preferably GCP
- Infrastructure automation experience with Helm, Terraform, and Ansible
- Observability experience with Prometheus, Grafana, Loki, OpenTelemetry, and DCGM
- Linux networking, namespaces, and performance tuning fundamentals
- Strong Python backend development with FastAPI, async programming, and SQLAlchemy
- Experience building Python agents that interact with Kubernetes APIs
- Modern frontend development with TypeScript, React or Next.js, Tailwind, and shadcn
- REST and tRPC API design experience
- Experience building developer tools, dashboards, and live-monitoring UIs
Benefits
- Cash compensation of $150K–$300K
- Significant equity
- Flexible work arrangement with remote or San Francisco office options
- Full visa sponsorship
- Relocation support
- Professional development budget for courses and conferences
- Regular team off-sites
- Conference attendance
- Opportunity to shape decentralized AI development