Skip to main content

๐œยฒ-bench: evaluating conversational agents in a dual-control environment

๐œยฒ-bench challenges AI agents not just to reason and act, but to coordinate, guide, and assist a user in achieving a shared objective. This leap from solo operation to co-ownership of a task pushes agents into a much more demanding space. And, critically, it reflects the kinds of tasks AI agents are increasingly being asked to perform in the real world.

Download
Tau Bench cover

Sierra๊ฐ€ ์ œ๊ณตํ•˜๋Š” ๊ธฐ๋Šฅ์„ ์•Œ์•„๋ณด์„ธ์š”

AI๋กœ ์„ฑ๊ณผ๋ฅผ ๋†’์ด๋Š” Sierra์˜ ๋ฐฉ๋ฒ•์„ ์•Œ์•„๋ณด์„ธ์š”.