Home Projects tau2-bench
tau2-bench
Python

tau2-bench

τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

by sierra-research · GitHub
Stars
Forks
License
Created
Last commit
Category
Language
aibenchmarkconversational-agentsMITPython
View on GitHub
In plain words

Evaluate how well AI agents perform tasks in real-world scenarios with a benchmarking tool.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
tau2-bench — GitHub preview card
📈 Star history
1.9k1.6k
2026-07-202026-08-31
📈 Track tau2-bench

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

tau2-bench has 1.9k stars on GitHub. It has been forked 481 times. tau2-bench is written mainly in Python. It has been in active development since 2025. tau2-bench is available under the MIT license. Its main topics are ai, benchmark, conversational-agents, language-model-agent.

Frequently asked questions

What is tau2-bench?

τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Is tau2-bench open source?

tau2-bench is an open-source project. It is released under the MIT license.

Is tau2-bench free?

Yes. tau2-bench is free and open source — you can use, modify and self-host it.

What license does tau2-bench use?

tau2-bench is available under the MIT license.

What language is tau2-bench written in?

tau2-bench is written mainly in Python.

🏅 Maintainer of this project?
olud.ai badge — tau2-bench

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=sierra-research-tau2-bench)](https://olud.ai/project/sierra-research-tau2-bench.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.