evals

19 Projekte teilen dieses GitHub-Topic

evals — mastra ★27.6kevalsphoenix — ★11.2kagentops — ★5.8klogfire — ★4.4ktrulens — ★3.5klmnr — ★3.2kai-system-design-guide — ★3kagent-skill-creator — ★2.4kfailproofai — ★1.6kwaku-agent — ★1.6ksimba — ★1.5kraglite — ★1.2kTracely-ai — ★1.2kheym — ★1.1kawesome-evals — ★852Tracely — ★833ai-engineer-notebooks — ★569agent-safety-eval-lab — ★323inferock-bench — ★140phoenix★ 11.2kagentops★ 5.8klogfire★ 4.4ktrulens★ 3.5klmnr★ 3.2kai-system-design-guide★ 3kagent-skill-creator★ 2.4kfailproofai★ 1.6kwaku-agent★ 1.6ksimba★ 1.5kraglite★ 1.2kTracely-ai★ 1.2kheym★ 1.1kawesome-evals★ 852Tracely★ 833ai-engineer-notebooks★ 569agent-safety-eval-lab★ 323inferock-bench★ 140

Linien verbinden Mitglieder, die messbar miteinander verwandt sind. Die Punktgröße spiegelt die Sterne wider.

🧬 Mitglieder
mastra
Mastra is the modern TypeScript framework for AI-powered applications and agents.
★ 27.6k
phoenix
AI Observability & Evaluation
★ 11.2k
agentops
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and…
★ 5.8k
logfire
AI observability platform for production LLM and agent systems.
★ 4.4k
trulens
Evaluation and Tracking for LLM Experiments and AI Agents
★ 3.5k
lmnr
Laminar - open-source observability platform purpose-built for AI agents. YC S24.
★ 3.2k
ai-system-design-guide
AI system design guide for engineers building production AI systems and evals.
★ 3k
agent-skill-creator
Build tested agent skills and govern their lifecycle through a user-defined marketplace: evidence, discovery,…
★ 2.4k
failproofai
Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy…
★ 1.6k
waku-agent
Waku Waku! Waku Agent is a local-first AI agent harness you actually own, including loop, memory, eval, all…
★ 1.6k
simba
OpenSource Production ready Customer service with built in Evals and monitoring
★ 1.5k
raglite
🥤 RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with DuckDB or PostgreSQL
★ 1.2k
Tracely-ai
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR.…
★ 1.2k
heym
Build agentic systems. Run them with confidence. Orchestrate agents, automate business processes, inspect…
★ 1.1k
awesome-evals
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs,…
★ 852
Tracely
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR.…
★ 833
ai-engineer-notebooks
Hands-on, framework-free Colab notebooks for the AI Engineer / Forward Deployed Engineer (FDE) skill set —…
★ 569
agent-safety-eval-lab
Agent trace and tool-use safety evaluation lab.
★ 323
inferock-bench
Local LLM cost-tracking proxy for OpenAI, Anthropic, Gemini, and pinned OpenRouter calls with token usage,…
★ 140
🔗 Verwandte Familien

Gemessen anhand der von beiden Projekten geteilten GitHub-Themen, gewichtet nach der Seltenheit jedes Themas.