Data Scientist, Agent
You will own the measurement and improvement of an AI agent by defining quality metrics, building evaluation and experimentation systems, analyzing agent traces and telemetry, identifying regressions, and working with engineering to improve success rates, task completion, and error rates.
Responsibilities
- Define and own metrics for agent quality, including success, completion, and error rates.
- Build evaluation systems and experiment frameworks to determine whether agent changes should ship.
- Analyze agent traces and telemetry to identify concrete fixes with the agent engineering team.
- Build tooling and agents that produce continuous evaluations as the agent evolves.
- Set standards for judging agent behavior when there is no clean answer key.
Requirements
- Experience or strong interest in LLM evaluation and observability.
- Strong SQL and Python skills.
- Applied statistics and experimentation experience, including A/B testing.
- Ability to design experiments with noisy outcomes.
- Ability to measure agent behavior without a clean answer key.
- Entrepreneurial approach and comfort with ambiguity.
- Ability to work closely with agent engineers.