Staff / Principal Machine Learning Engineer, Serving
Take models from the research team, containerize and optimize their serving, and ensure reliable production operation. Work on inference optimization, model acceleration, and high-performance distributed systems powering realtime multimodal AI at scale.
Responsibilities
- Optimize realtime inference and serving frameworks
- Containerize research models and ensure reliable production deployment
- Apply quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
- Profile code and optimize performance on NVIDIA GPUs
- Handle multi-GPU and multi-node inference and thousands of concurrent connections
- Design benchmarks and prototypes to validate technical decisions
- Share work and contribute to open-source projects
Requirements
- Deep understanding of modern serving frameworks such as vLLM or TRT-LLM
- Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
- Proficiency in C++, CUDA, Rust, or highly optimized Python
- Experience with Kubernetes, Ray, custom load balancing, multi-GPU or multi-node inference, and thousands of concurrent connections
- Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups
- Full-cycle ownership from research model to production serving
- PhD in CS, Physics, or Math, or equivalent practical experience building backend or ML systems
- Legal right to work in the United Kingdom
Benefits
- Equity
- Benefits