Senior Software Engineer - Model Performance

You will make the inference stack faster and more efficient by implementing optimization techniques, experimenting with novel approaches, profiling GPU workloads, and bringing performant model architectures into production.

Responsibilities

  • Implement and productionize quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving.
  • Debug and improve vLLM, SGLang, TensorRT-LLM, and underlying libraries.
  • Profile CUDA kernels and optimize GPU utilization across serving infrastructure.
  • Add support for new model architectures and ensure they meet production performance standards.
  • Experiment with novel inference techniques and productionize successful approaches.
  • Build tooling and benchmarks to track inference performance across the fleet.
  • Collaborate with applied ML engineers to ensure trained models can be served efficiently.

Requirements

  • 2+ years of experience in ML systems, inference optimization, or GPU programming.
  • Strong proficiency in Python and familiarity with C++.
  • Hands-on experience with LLM inference frameworks such as vLLM, SGLang, TensorRT-LLM, or similar.
  • Deep understanding of GPU architecture and experience profiling GPU workloads.
  • Familiarity with quantization, speculative decoding, continuous batching, and KV cache management.
  • Experience with PyTorch and understanding of model execution on hardware.
  • Track record of measurably improving system performance.
  • Nice-to-have experience with CUDA programming, non-LLM model serving, distributed inference, open-source inference frameworks, Docker, and Kubernetes.

Benefits

  • Equity
  • Comprehensive benefits

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available