Member of Technical Staff - Evals
You will play a key role in ensuring the quality and reliability of AI-powered features by designing evaluation frameworks, building automated test harnesses using real accounting data, defining metrics and benchmarks, creating evaluation tooling, and collaborating with Applied AI and Agent Engineering teams.
Responsibilities
- Design and maintain evaluation frameworks to measure accuracy, reliability, and regression behavior of AI capabilities
- Build automated test harnesses that operate on real accounting data to identify potential failures before they impact customers
- Define metrics and benchmarks to provide quantitative insights into model and system performance
- Create tooling that enables engineers to write, execute, and interpret evaluation tasks within their development workflow
- Collaborate with Applied AI and Agent Engineering teams to define and verify standards for shipped capabilities
Requirements
- Experience writing and reviewing evaluation tasks used by AI researchers to assess and improve model performance
- Background in quality assurance across software, client deliverables, or financial analysis
- Exceptional written communication skills, with the ability to clearly articulate steps to achieve desired outcomes
- Familiarity with accounting, financial systems, tax preparation, or payment technologies preferred