AI Research Scientist, Reinforcement Learning (LLM) and Post-Training

Summary

Research scientist developing reinforcement learning methods for post-training large language models and code models. Day-to-day work includes designing reward models and training curricula, running on/off-policy RL experiments, scaling RL infrastructure, and publishing at top academic venues.

- Analyze failure modes reward hacking and instability - Collaborate to scale training with RL infrastructure - Define interfaces for rollout generation and logging - Design reward models and training curricula - Develop reinforcement learning methods for post training large language models and code models - Publish research at top academic venues - Run off policy and on policy training experiments

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available