Skip to main content

๐œ-bench: benchmarking AI agents for the real-world

Sierraโ€™s AI research team is on a mission to advance the frontier of conversational AI agents. In this research paper, we present a new benchmark for evaluating AI agents' performance and reliability in real-world settings, with dynamic user and tool interaction.

Download
Tau Bench cover

Sierra๊ฐ€ ์ œ๊ณตํ•˜๋Š” ๊ธฐ๋Šฅ์„ ์•Œ์•„๋ณด์„ธ์š”

AI๋กœ ์„ฑ๊ณผ๋ฅผ ๋†’์ด๋Š” Sierra์˜ ๋ฐฉ๋ฒ•์„ ์•Œ์•„๋ณด์„ธ์š”.