Forward Deployed Infrastructure Engineer
Serve as the technical point of contact during customer trials, running standardized and custom benchmarks to validate performance. Design, run, and analyze performance tests across customer workloads, diagnose GPU and NCCL issues, optimize container and cluster configurations, produce clear reports and handoffs, and maintain benchmarking scripts, containers, and environments.
Responsibilities
- Serve as the technical point of contact during customer trials
- Design and run performance and benchmark tests across customer workloads
- Diagnose performance issues and recommend fixes
- Package results into clear reports and handoffs
- Maintain benchmarking scripts, containers, and environments
- Identify performance gaps and optimize cluster configurations
Requirements
- Experience running infrastructure performance tests or ML model benchmarks, including training or inference
- Strong knowledge of GPU cloud infrastructure and workload bottlenecks
- Clear and fast written communication
- Ability to manage multiple trials and projects concurrently
- Familiarity with AWS, Lambda, CoreWeave, Runpod, and similar GPU cloud providers
- Prior customer-facing experience in a startup or devtools setting is preferred
- Background as an ML engineer, solutions architect, or technical account manager is preferred