Senior / Lead Machine Learning Engineer, Serving
Work on optimizing realtime inference and serving of state-of-the-art voice models at massive scale. Take models from research, containerize and optimize their serving, ensure reliable production operation, and improve latency, throughput, and performance across distributed NVIDIA GPU systems.
Responsibilities
- Optimize realtime inference and model serving at scale
- Take models from the research team, containerize them, and optimize their serving
- Ensure models run reliably in production
- Apply model acceleration techniques including quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
- Profile code and optimize performance on NVIDIA GPUs
- Handle distributed systems and scaling using Kubernetes, Ray, and custom load balancing
- Manage multi-GPU and multi-node inference while handling thousands of concurrent connections
- Collaborate with US-based leadership and engineering teams
Requirements
- Deep understanding of modern serving frameworks such as vLLM or TRT-LLM
- Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
- Proficiency in C++, CUDA, Rust, or highly optimized Python
- Experience with Kubernetes, Ray, custom load balancing, and multi-GPU or multi-node inference
- Experience reliably handling thousands of concurrent connections
- Non-trivial systems programming projects, major inference-engine open-source contributions, or deep-dive technical write-ups
- Ability to take a model from research to production, containerize it, and optimize its serving
- PhD in CS, Physics, or Math, or equivalent practical experience building backend or ML systems
- Professional fluency in written and spoken English
Benefits
- Full U.S. visa and relocation support may be available for candidates interested in relocating to the San Francisco Bay Area, subject to business needs and applicable legal and work authorization requirements