Staff / Principal Machine Learning Engineer, Serving

Optimize sub-second multimodal inference for realtime voice models, taking models from research into production through containerization, serving optimization, and reliable operation. Apply model acceleration techniques and build high-performance distributed systems capable of handling thousands of concurrent connections.

Responsibilities

  • Optimize realtime inference serving with frameworks such as vLLM and TRT-LLM.
  • Apply quantization, distillation, caching, continuous batching, paged attention, and speculative decoding.
  • Build high-performance systems in C++, CUDA, Rust, or optimized Python and profile GPU performance.
  • Design and scale distributed systems using Kubernetes, Ray, custom load balancing, and multi-GPU or multi-node inference.
  • Handle thousands of concurrent connections reliably.
  • Take models from research to production, including containerization and serving optimization.
  • Contribute to open-source projects and publish technical write-ups.

Requirements

  • Deep understanding of modern serving frameworks such as vLLM or TRT-LLM.
  • Hands-on experience with model acceleration techniques.
  • Proficiency in C++, CUDA, Rust, or highly optimized Python.
  • Experience profiling code and optimizing NVIDIA GPU performance.
  • Experience with Kubernetes, Ray, custom load balancing, and multi-GPU or multi-node inference.
  • Experience handling thousands of concurrent connections reliably.
  • Non-trivial systems programming projects, open-source contributions, or deep-dive technical write-ups.
  • Ability to take models from research to production.
  • PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems.

Benefits

  • Relocation assistance
  • Bonus
  • Equity
  • Benefits

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available