llm-evaluation

42 projetos partilham este topic do GitHub

llm-evaluation — langfuse ★34kllm-evaluationmlflow — ★27.8kpromptfoo — ★24.7kopik — ★21.7kdeepeval — ★18kiFixAi — ★12.1kphoenix — ★11.2kgarak — ★9.1kchinese-llm-benchmark — ★6.4khelicone — ★6.1kAI-Infra-Guard — ★6.1kgiskard-oss — ★5.8kouroboros — ★5.4kLLM-Engineers-Handbook — ★5.3kAutoRAG — ★5.1klmms-eval — ★4.4kAI-Engineer-Headquarters — ★3.7ktrulens — ★3.5klmnr — ★3.2kgenerative-ai — ★2.6kfuture-agi — ★1.9kFuzzyAI — ★1.6kprompty — ★1.3kTracely-ai — ★1.2kjudgeval — ★1.1kFinSight-AI — ★1kawesome-evals — ★852Tracely — ★833agent-skills-eval — ★718Awesome-LLM-in-Social-Science — ★647ClawBench — ★614EnterpriseRAG-Bench — ★537continuous-eval — ★516rhesis — ★381aaabench — ★377llm-leaderboard — ★360KADATH — ★273EvoTrace — ★196ai-runtime-lab — ★124ai-eval-platform — ★117ai-reliability-copilot — ★99mlflow★ 27.8kpromptfoo★ 24.7kopik★ 21.7kdeepeval★ 18kiFixAi★ 12.1kphoenix★ 11.2kgarak★ 9.1kchinese-llm-benchmark★ 6.4khelicone★ 6.1kAI-Infra-Guard★ 6.1kgiskard-oss★ 5.8kouroboros★ 5.4kLLM-Engineers-Handbook★ 5.3kAutoRAG★ 5.1klmms-eval★ 4.4kAI-Engineer-Headquarters★ 3.7ktrulens★ 3.5klmnr★ 3.2kgenerative-ai★ 2.6kfuture-agi★ 1.9kFuzzyAI★ 1.6kprompty★ 1.3kTracely-ai★ 1.2kjudgeval★ 1.1kFinSight-AI★ 1kawesome-evals★ 852Tracely★ 833agent-skills-eval★ 718Awesome-LLM-in-Social-Sc…★ 647ClawBench★ 614EnterpriseRAG-Bench★ 537continuous-eval★ 516rhesis★ 381aaabench★ 377llm-leaderboard★ 360KADATH★ 273EvoTrace★ 196 · GitHub ↗ai-runtime-lab★ 124ai-eval-platform★ 117ai-reliability-copilot★ 99

Linhas conectam membros que estão mensuravelmente relacionados entre si. O tamanho do ponto reflete estrelas.

🧬 Membros
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground,…
★ 34k
mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to…
★ 27.8k
promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare…
★ 24.7k
opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive…
★ 21.7k
deepeval
The LLM Evaluation Framework
★ 18k
iFixAi
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in…
★ 12.1k
phoenix
AI Observability & Evaluation
★ 11.2k
garak
the LLM vulnerability scanner
★ 9.1k
chinese-llm-benchmark
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个…
★ 6.4k
helicone
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23…
★ 6.1k
AI-Infra-Guard
A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra…
★ 6.1k
giskard-oss
🐢 Open-Source Evaluation & Testing library for LLM Agents
★ 5.8k
ouroboros
Agent OS: Stop prompting. Start specifying. A Socratic interview gates the spec on an ambiguity score, then…
★ 5.4k
LLM-Engineers-Handbook
The LLM's practical guide: From the fundamentals to deploying advanced LLM and RAG apps to AWS using LLMOps…
★ 5.3k
AutoRAG
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
★ 5.1k
lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.4k
AI-Engineer-Headquarters
A collection of scientific methods, processes, algorithms, and systems to build stories & models.
★ 3.7k
trulens
Evaluation and Tracking for LLM Experiments and AI Agents
★ 3.5k
lmnr
Laminar - open-source observability platform purpose-built for AI agents. YC S24.
★ 3.2k
generative-ai
Comprehensive resources on Generative AI, including a detailed roadmap, projects, use cases, interview…
★ 2.6k
future-agi
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications.…
★ 1.9k
FuzzyAI
A powerful tool for automated LLM fuzzing. It is designed to help developers and security researchers…
★ 1.6k
prompty
Prompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty…
★ 1.3k
Tracely-ai
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR.…
★ 1.2k
judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and…
★ 1.1k
FinSight-AI
AI equity research agent with resilient workflows, evidence-grounded RAG, versioned reports, and automated…
★ 1k
awesome-evals
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs,…
★ 852
Tracely
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR.…
★ 833
agent-skills-eval
A test runner for agentskills.io-style AI agent skills
★ 718
Awesome-LLM-in-Social-Science
Awesome papers involving LLMs in Social Science.
★ 647
ClawBench
Open-source benchmark for browser AI agents on daily tasks.
★ 614
EnterpriseRAG-Bench
Dataset and benchmark for RAG on company internal documents.
★ 537
continuous-eval
Data-Driven Evaluation for LLM-Powered Applications
★ 516
rhesis
The testing platform for AI teams. Bring engineers, PMs, and domain experts together to generate tests,…
★ 381
aaabench
A long-horizon benchmark harness: give a coding agent a real game engine, professional conditions and time,…
★ 377
llm-leaderboard
A comprehensive set of LLM benchmark scores and provider prices. (deprecated, read more in README)
★ 360
KADATH
Evolutionary multi-agent runtime that breeds, evaluates, and improves autonomous agents across reproducible…
★ 273
EvoTrace
Compile real-world Claude Code and Codex trajectories into verified, tradable post-training assets.
★ 196 · GitHub ↗
ai-runtime-lab
Engineering deterministic, production-grade systems around non-deterministic LLMs — FSM, durable execution,…
★ 124
ai-eval-platform
★ 117
ai-reliability-copilot
Turn a production incident into a structured 9-section LLM response (severity, root cause, mitigation,…
★ 99
VLMForge
Evaluate visual models on your own images, JSON Schema, and production constraints with LangGraph, Pareto…
★ 89
🔗 Familias relacionadas

Medido a partir dos tópicos do GitHub compartilhados por ambos os projetos, ponderado pela raridade de cada tópico.