Staff / Principal Machine Learning Engineer, Serving

Own the full lifecycle of research models by containerizing, optimizing, and operating their serving infrastructure in production. Build sub-second multimodal inference systems using model acceleration techniques and distributed infrastructure capable of handling thousands of concurrent connections.

Responsibilities

  • Take models from the research team to production
  • Containerize and optimize model serving
  • Optimize inference using quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
  • Build distributed serving infrastructure across multiple GPUs and nodes
  • Handle thousands of concurrent connections reliably
  • Profile code and optimize performance on NVIDIA GPUs
  • Design benchmarks and prototypes to resolve open questions
  • Contribute to open-source projects and technical write-ups

Requirements

  • Deep understanding of modern serving frameworks like vLLM or TRT-LLM
  • Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
  • Proficiency in C++, CUDA, Rust, or highly optimized Python
  • Experience with Kubernetes, Ray, custom load balancing, and multi-GPU/multi-node inference
  • Public work such as open-source contributions or technical write-ups
  • Full-cycle ownership experience taking models to production
  • PhD in CS, Physics, Math, or equivalent practical experience
  • Professional fluency in English
  • Legal right to work in Switzerland

Benefits

  • Remote work
  • Full U.S. visa and relocation support may be available for those interested in relocating to the San Francisco Bay Area

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available