Member of Technical Staff - Compute Platform
Build platform software and infrastructure for managing and monitoring AI workloads, including web interfaces, Python APIs, backend services, debugging tools, distributed training infrastructure, automation pipelines, cloud resources, container orchesation, and hardware scheduling systems.
Responsibilities
- Build web interfaces for AI workload management and monitoring
- Develop REST APIs and backend services in Python
- Create real-time monitoring and debugging tools
- Implement user-facing features for resource management and job control
- Design distributed training infrastructure in Rust
- Build high-performance networking and coordination components
- Create infrastructure automation pipelines with Ansible
- Manage cloud resources and container orchestration
- Implement scheduling systems for CPU, GPU, and TPU hardware
- Integrate backend features into existing infrastructure
Requirements
- Strong Python backend development with FastAPI and async
- Modern frontend development with TypeScript, React/Next.js, and Tailwind
- Experience building developer tools and dashboards
- RESTful API design and implementation
- Systems programming experience with Rust
- Infrastructure automation with Ansible and Terraform
- Container orchestration with Kubernetes
- Cloud platform expertise, preferably GCP
- Observability tools such as Prometheus and Grafana
- GPU computing or ML infrastructure experience