- Analyze distributed GPU performance on RoCE and Infiniband
- Build NCCL and PyTorch performance tuners and benchmarks
- Develop AI framework and trainer for large scale distributed deep learning
- Enable distributed GPU communication for multi GPU multi node training
- Implement responsible ethical AI practices
- Improve distributed machine learning reliability and performance
- Lead collective communication library development
- Support data parallel and model parallel training