A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
Explore a curated collection of resources for building and evaluating AI agents, including papers and tools.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytimeA curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
awesome-evals has 850 stars on GitHub. It has been forked 89 times. It has been in active development since 2026. Its main topics are agent-evaluation, ai-agents, awesome, awesome-list.
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
awesome-evals is an open-source project.
Yes. awesome-evals is free and open source — you can use, modify and self-host it.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/benchflow-ai-awesome-evals.html)
Measured from GitHub topics shared by both projects, weighted by how rare each topic is.