Staff AI Engineer Model Post-Training and Alignment
Design, execute, and optimize post-training pipelines for large language models across data strategy, reward modeling, reinforcement learning, alignment, evaluation, and production inference.
Responsibilities
- Lead and execute LLM post-training pipelines
- Design DPO and GRPO training paradigms
- Develop domain-specific data recipes and augmentation pipelines
- Train specialized small models from scratch
- Build and refine reward models
- Design and implement RLAIF closed-loop systems
- Optimize inference efficiency and deploy models
- Evaluate model performance with benchmarks and feedback loops
- Collaborate to productionize training and deployment workflows
Requirements
- Bachelor's degree in Computer Science, AI, Machine Learning, or a related field
- 8+ years of industry experience
- Strong experience with large model post-training
- Preference learning and alignment techniques including DPO and GRPO
- Reinforcement learning fundamentals
- Domain-specific data strategies
- Training specialized small models from scratch
- Reward modeling and RLAIF
- Low-latency production deployment using vLLM, SGLang, or similar
Benefits
- L&D programs
- Education subsidy
- Team building programs
- Company events
- Wellness allowances
- Meal allowances
- Comprehensive healthcare schemes for employees and dependants
- Competitive total compensation package