Soniox
1 job
2 days ago
Software Engineer, LLM Inference
HybridEuropeSenior
Slovenia
Engineer optimizing Soniox's vLLM-based LLM inference stack for low latency, high throughput, and maximum GPU utilization at production scale — working on CUDA/Triton kernels, KV-cache management, distributed GPU execution, and deep performance profiling for their voice AI platform.
ai cuda distributed-systems llm pytorch +4 skills