Staff / Principal Machine Learning Engineer, Serving

Take models from the research team, containerize and optimize their serving, and ensure reliable production operation. Work on inference optimization, model acceleration, and high-performance distributed systems powering realtime multimodal AI at scale.

Responsibilities

  • Optimize realtime inference and serving frameworks
  • Containerize research models and ensure reliable production deployment
  • Apply quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
  • Profile code and optimize performance on NVIDIA GPUs
  • Handle multi-GPU and multi-node inference and thousands of concurrent connections
  • Design benchmarks and prototypes to validate technical decisions
  • Share work and contribute to open-source projects

Requirements

  • Deep understanding of modern serving frameworks such as vLLM or TRT-LLM
  • Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
  • Proficiency in C++, CUDA, Rust, or highly optimized Python
  • Experience with Kubernetes, Ray, custom load balancing, multi-GPU or multi-node inference, and thousands of concurrent connections
  • Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups
  • Full-cycle ownership from research model to production serving
  • PhD in CS, Physics, or Math, or equivalent practical experience building backend or ML systems
  • Legal right to work in the United Kingdom

Benefits

  • Equity
  • Benefits

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available