Home Projects evals
evals
Python

evals

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

by openai · GitHub
Stars
Forks
Created
Last commit
Language
Python
View on GitHub
In plain words

Evaluate and test large language models with existing benchmarks or create your own custom tests.

From the README

Install
pip install evals

Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
evals — GitHub preview card
📈 Star history
19.3k18.9k
2026-07-122026-08-31
📈 Track evals

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

evals has 19.3k stars on GitHub. It has been forked 3.1k times. evals is written mainly in Python. It has been in active development since 2023.

Frequently asked questions

What is evals?

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

Is evals open source?

evals is an open-source project.

Is evals free?

Yes. evals is free and open source — you can use, modify and self-host it.

What language is evals written in?

evals is written mainly in Python.

🏅 Maintainer of this project?
olud.ai badge — evals

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=openai-evals)](https://olud.ai/project/openai-evals.html)
More badge options →
🧬 Related projects