Staff / Principal Machine Learning Engineer, Serving
Optimize sub-second multimodal inference for realtime voice models, taking models from research into production through containerization, serving optimization, and reliable operation. Apply model acceleration techniques and build high-performance distributed systems capable of handling thousands of concurrent connections.
Responsibilities
- Optimize realtime inference serving with frameworks such as vLLM and TRT-LLM.
- Apply quantization, distillation, caching, continuous batching, paged attention, and speculative decoding.
- Build high-performance systems in C++, CUDA, Rust, or optimized Python and profile GPU performance.
- Design and scale distributed systems using Kubernetes, Ray, custom load balancing, and multi-GPU or multi-node inference.
- Handle thousands of concurrent connections reliably.
- Take models from research to production, including containerization and serving optimization.
- Contribute to open-source projects and publish technical write-ups.
Requirements
- Deep understanding of modern serving frameworks such as vLLM or TRT-LLM.
- Hands-on experience with model acceleration techniques.
- Proficiency in C++, CUDA, Rust, or highly optimized Python.
- Experience profiling code and optimizing NVIDIA GPU performance.
- Experience with Kubernetes, Ray, custom load balancing, and multi-GPU or multi-node inference.
- Experience handling thousands of concurrent connections reliably.
- Non-trivial systems programming projects, open-source contributions, or deep-dive technical write-ups.
- Ability to take models from research to production.
- PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems.
Benefits
- Relocation assistance
- Bonus
- Equity
- Benefits