Staff AI Engineer Model Post-Training and Alignment

Design, execute, and optimize post-training pipelines for large language models across data strategy, reward modeling, reinforcement learning, alignment, evaluation, and production inference.

Responsibilities

  • Lead and execute LLM post-training pipelines
  • Design DPO and GRPO training paradigms
  • Develop domain-specific data recipes and augmentation pipelines
  • Train specialized small models from scratch
  • Build and refine reward models
  • Design and implement RLAIF closed-loop systems
  • Optimize inference efficiency and deploy models
  • Evaluate model performance with benchmarks and feedback loops
  • Collaborate to productionize training and deployment workflows

Requirements

  • Bachelor's degree in Computer Science, AI, Machine Learning, or a related field
  • 8+ years of industry experience
  • Strong experience with large model post-training
  • Preference learning and alignment techniques including DPO and GRPO
  • Reinforcement learning fundamentals
  • Domain-specific data strategies
  • Training specialized small models from scratch
  • Reward modeling and RLAIF
  • Low-latency production deployment using vLLM, SGLang, or similar

Benefits

  • L&D programs
  • Education subsidy
  • Team building programs
  • Company events
  • Wellness allowances
  • Meal allowances
  • Comprehensive healthcare schemes for employees and dependants
  • Competitive total compensation package

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available