Staff / Principal Machine Learning Engineer, Serving
Own the full lifecycle of research models by containerizing, optimizing, and operating their serving infrastructure in production. Build sub-second multimodal inference systems using model acceleration techniques and distributed infrastructure capable of handling thousands of concurrent connections.
Responsibilities
- Take models from the research team to production
- Containerize and optimize model serving
- Optimize inference using quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
- Build distributed serving infrastructure across multiple GPUs and nodes
- Handle thousands of concurrent connections reliably
- Profile code and optimize performance on NVIDIA GPUs
- Design benchmarks and prototypes to resolve open questions
- Contribute to open-source projects and technical write-ups
Requirements
- Deep understanding of modern serving frameworks like vLLM or TRT-LLM
- Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
- Proficiency in C++, CUDA, Rust, or highly optimized Python
- Experience with Kubernetes, Ray, custom load balancing, and multi-GPU/multi-node inference
- Public work such as open-source contributions or technical write-ups
- Full-cycle ownership experience taking models to production
- PhD in CS, Physics, Math, or equivalent practical experience
- Professional fluency in English
- Legal right to work in Switzerland
Benefits
- Remote work
- Full U.S. visa and relocation support may be available for those interested in relocating to the San Francisco Bay Area