llm-evaluation

37 projets partagent ce topic GitHub

llm-evaluation — langfuse ★32.1kllm-evaluationmlflow — ★27.3kpromptfoo — ★23.7kopik — ★21kdeepeval — ★17.3kphoenix — ★10.8kgarak — ★8.6kchinese-llm-benchmark — ★6.3khelicone — ★6kgiskard-oss — ★5.7kLLM-Engineers-Handbook — ★5.3kAutoRAG — ★5kagenta — ★4.4klmms-eval — ★4.3kAI-Infra-Guard — ★4.3kiFixAi — ★3.5ktrulens — ★3.5klmnr — ★3.1kgenerative-ai — ★2.6kFuzzyAI — ★1.5kfuture-agi — ★1.5kprompty — ★1.2kjudgeval — ★1kFinSight-AI — ★1kawesome-evals — ★765Awesome-LLM-Eval — ★654agent-skills-eval — ★641Awesome-LLM-in-Social-Science — ★639ClawBench — ★536continuous-eval — ★515EnterpriseRAG-Bench — ★491llm-leaderboard — ★359palico-ai — ★343UltraEval-Audio — ★311athina-evals — ★301coder_eval — ★107ai-reliability-copilot — ★102mlflow★ 27.3kpromptfoo★ 23.7kopik★ 21kdeepeval★ 17.3kphoenix★ 10.8kgarak★ 8.6kchinese-llm-benchmark★ 6.3khelicone★ 6kgiskard-oss★ 5.7kLLM-Engineers-Handbook★ 5.3kAutoRAG★ 5kagenta★ 4.4klmms-eval★ 4.3kAI-Infra-Guard★ 4.3kiFixAi★ 3.5ktrulens★ 3.5klmnr★ 3.1kgenerative-ai★ 2.6kFuzzyAI★ 1.5kfuture-agi★ 1.5kprompty★ 1.2kjudgeval★ 1kFinSight-AI★ 1kawesome-evals★ 765Awesome-LLM-Eval★ 654agent-skills-eval★ 641Awesome-LLM-in-Social-Sc…★ 639ClawBench★ 536continuous-eval★ 515EnterpriseRAG-Bench★ 491llm-leaderboard★ 359palico-ai★ 343UltraEval-Audio★ 311athina-evals★ 301coder_eval★ 107ai-reliability-copilot★ 102

Les traits relient les membres réellement apparentés entre eux. La taille des points suit les étoiles.

🧬 Membres
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground,…
★ 32.1k
mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to…
★ 27.3k
promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare…
★ 23.7k
opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive…
★ 21k
deepeval
The LLM Evaluation Framework
★ 17.3k
phoenix
AI Observability & Evaluation
★ 10.8k
garak
the LLM vulnerability scanner
★ 8.6k
chinese-llm-benchmark
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个…
★ 6.3k
helicone
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23…
★ 6k
giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents
★ 5.7k
LLM-Engineers-Handbook
The LLM's practical guide: From the fundamentals to deploying advanced LLM and RAG apps to AWS using LLMOps…
★ 5.3k
AutoRAG
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
★ 5k
agenta
The open-source workspace for building and running AI agents. Build agents through chat, share them with your…
★ 4.4k
lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3k
AI-Infra-Guard
A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, Skills…
★ 4.3k
iFixAi
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in…
★ 3.5k
trulens
Evaluation and Tracking for LLM Experiments and AI Agents
★ 3.5k
lmnr
Laminar - open-source observability platform purpose-built for AI agents. YC S24.
★ 3.1k
generative-ai
Comprehensive resources on Generative AI, including a detailed roadmap, projects, use cases, interview…
★ 2.6k
FuzzyAI
A powerful tool for automated LLM fuzzing. It is designed to help developers and security researchers…
★ 1.5k
future-agi
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications.…
★ 1.5k
prompty
Prompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty…
★ 1.2k
judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and…
★ 1k
FinSight-AI
AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports,…
★ 1k
awesome-evals
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs,…
★ 765
Awesome-LLM-Eval
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models,…
★ 654
agent-skills-eval
A test runner for agentskills.io-style AI agent skills
★ 641
Awesome-LLM-in-Social-Science
Awesome papers involving LLMs in Social Science.
★ 639
ClawBench
Open-source benchmark for browser AI agents on daily tasks.
★ 536
continuous-eval
Data-Driven Evaluation for LLM-Powered Applications
★ 515
EnterpriseRAG-Bench
Dataset and benchmark for RAG on company internal documents.
★ 491
llm-leaderboard
A comprehensive set of LLM benchmark scores and provider prices. (deprecated, read more in README)
★ 359
palico-ai
Build, Improve Performance, and Productionize your AI Application
★ 343
UltraEval-Audio
Your faithful, impartial partner for audio evaluation — know yourself, know your rivals.…
★ 311
athina-evals
Python SDK for running evaluations on LLM generated responses
★ 301
coder_eval
Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for…
★ 107
ai-reliability-copilot
Turn a production incident into a structured 9-section LLM response (severity, root cause, mitigation,…
★ 102
🔗 Familles voisines

Mesuré à partir des topics GitHub communs aux deux projets, pondérés par leur rareté.