Machine Learning Engineer
Design and build infrastructure that powers quantitative research at scale, including experiment tracking, job orchestration, reproducibility systems, simulation tools, GPU utilization monitoring, distributed training diagnostics, financial data pipelines, and feature storage and retrieval patterns.
Responsibilities
- Design and build experiment tracking, job orchestration, and reproducibility infrastructure
- Create tools for all stages of the simulation lifecycle, including historical back-tests and production monitoring
- Own visibility into GPU cluster utilization, track allocation, and surface bottlenecks
- Diagnose and resolve performance issues across training pipelines
- Build and maintain data pipelines that move financial data into training workflows with correctness and versioning guarantees
- Develop feature storage and retrieval patterns for fast, reproducible access to training data at scale
- Work directly with researchers to reduce workflow friction
- Collaborate with infrastructure engineers on capacity planning and tooling decisions
- Stay current with ML infrastructure developments and bring valuable ideas into the stack
Requirements
- 5+ years of experience in ML engineering, research infrastructure, or HPC environments
- Strong Python engineering skills with clean, maintainable, well-tested code
- Experience building or operating distributed training infrastructure and knowledge of collective communication libraries such as NCCL or Horovod
- Practical experience with experiment tracking systems
- Comfort working across the Linux systems stack, including storage, networking, and job scheduling
- Excellent communication skills and ability to work closely with researchers and engineers
- Exposure to C++ in a performance-sensitive context is a plus
- Experience with on-prem compute environments and Slurm is desired
- Familiarity with GPU profiling tools such as NSight Systems and PyTorch Profiler is desired
- Experience with Parquet, Arrow, and Polars is desired
- Familiarity with Prefect, Dagster, or similar workflow orchestration tools is desired