Senior Software Engineer - Model Performance
You will make the inference stack faster and more efficient by implementing optimization techniques, experimenting with novel approaches, profiling GPU workloads, and bringing performant model architectures into production.
Responsibilities
- Implement and productionize quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving.
- Debug and improve vLLM, SGLang, TensorRT-LLM, and underlying libraries.
- Profile CUDA kernels and optimize GPU utilization across serving infrastructure.
- Add support for new model architectures and ensure they meet production performance standards.
- Experiment with novel inference techniques and productionize successful approaches.
- Build tooling and benchmarks to track inference performance across the fleet.
- Collaborate with applied ML engineers to ensure trained models can be served efficiently.
Requirements
- 2+ years of experience in ML systems, inference optimization, or GPU programming.
- Strong proficiency in Python and familiarity with C++.
- Hands-on experience with LLM inference frameworks such as vLLM, SGLang, TensorRT-LLM, or similar.
- Deep understanding of GPU architecture and experience profiling GPU workloads.
- Familiarity with quantization, speculative decoding, continuous batching, and KV cache management.
- Experience with PyTorch and understanding of model execution on hardware.
- Track record of measurably improving system performance.
- Nice-to-have experience with CUDA programming, non-LLM model serving, distributed inference, open-source inference frameworks, Docker, and Kubernetes.
Benefits
- Equity
- Comprehensive benefits