evals

17 Projekte teilen dieses GitHub-Topic

evals — mastra ★26.7kevalsphoenix — ★10.8kagentops — ★5.7kKiln — ★5klogfire — ★4.4ktrulens — ★3.5klmnr — ★3.1kai-system-design-guide — ★2.3kinspector — ★2.1ksimba — ★1.5kraglite — ★1.2kheym — ★825awesome-evals — ★765waku-agent — ★609agent-safety-eval-lab — ★341inferock-bench — ★123coder_eval — ★107phoenix★ 10.8kagentops★ 5.7kKiln★ 5klogfire★ 4.4ktrulens★ 3.5klmnr★ 3.1kai-system-design-guide★ 2.3kinspector★ 2.1ksimba★ 1.5kraglite★ 1.2kheym★ 825awesome-evals★ 765waku-agent★ 609agent-safety-eval-lab★ 341inferock-bench★ 123coder_eval★ 107

Linien verbinden Mitglieder, die messbar miteinander verwandt sind. Die Punktgröße spiegelt die Sterne wider.

🧬 Mitglieder
mastra
Mastra is the modern TypeScript framework for AI-powered applications and agents.
★ 26.7k
phoenix
AI Observability & Evaluation
★ 10.8k
agentops
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and…
★ 5.7k
Kiln
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data…
★ 5k
logfire
AI observability platform for production LLM and agent systems.
★ 4.4k
trulens
Evaluation and Tracking for LLM Experiments and AI Agents
★ 3.5k
lmnr
Laminar - open-source observability platform purpose-built for AI agents. YC S24.
★ 3.1k
ai-system-design-guide
AI system design guide for engineers building production AI systems and evals.
★ 2.3k
inspector
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
★ 2.1k
simba
OpenSource Production ready Customer service with built in Evals and monitoring
★ 1.5k
raglite
🥤 RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with DuckDB or PostgreSQL
★ 1.2k
heym
Build AI workflows by prompt or visual canvas. Heym is source-available and self-hosted, with agents, RAG,…
★ 825
awesome-evals
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs,…
★ 765
waku-agent
Waku Waku! Waku agent is your personal AI agent, on your own laptop, in code you can read in an afternoon —…
★ 609
agent-safety-eval-lab
Agent trace and tool-use safety evaluation lab.
★ 341
inferock-bench
Local LLM cost-tracking proxy for OpenAI, Anthropic, Gemini, and pinned OpenRouter calls with token usage,…
★ 123
coder_eval
Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for…
★ 107
🔗 Verwandte Familien

Gemessen anhand der von beiden Projekten geteilten GitHub-Themen, gewichtet nach der Seltenheit jedes Themas.