メインコンテンツへスキップ

𝜏-bench: benchmarking AI agents for the real-world

Sierra’s AI research team is on a mission to advance the frontier of conversational AI agents. In this research paper, we present a new benchmark for evaluating AI agents' performance and reliability in real-world settings, with dynamic user and tool interaction.

ダウンロード
Tau Bench cover

Sierraにできること

SierraがAIを活用してより良い成果の実現をどのように支援できるかをご覧ください。