Data Scientist, Agent

You will own the measurement and improvement of an AI agent by defining quality metrics, building evaluation and experimentation systems, analyzing agent traces and telemetry, identifying regressions, and working with engineering to improve success rates, task completion, and error rates.

Responsibilities

  • Define and own metrics for agent quality, including success, completion, and error rates.
  • Build evaluation systems and experiment frameworks to determine whether agent changes should ship.
  • Analyze agent traces and telemetry to identify concrete fixes with the agent engineering team.
  • Build tooling and agents that produce continuous evaluations as the agent evolves.
  • Set standards for judging agent behavior when there is no clean answer key.

Requirements

  • Experience or strong interest in LLM evaluation and observability.
  • Strong SQL and Python skills.
  • Applied statistics and experimentation experience, including A/B testing.
  • Ability to design experiments with noisy outcomes.
  • Ability to measure agent behavior without a clean answer key.
  • Entrepreneurial approach and comfort with ambiguity.
  • Ability to work closely with agent engineers.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available