Open source llm-evaluation projects
Every project in the registry tagged llm-evaluation, ranked by real GitHub adoption.
Open source LLM engineering platform for AI-powered applications
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simpl
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready d
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is sup
Open-source LLM tracing & evaluation for AI optimization
the LLM vulnerability scanner
A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.
Build reliable AI apps with comprehensive observability
🐢 Open-Source Evaluation & Testing library for LLM Agents
AI-powered platform for engineering LLM products
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Gu
Agentic LLM Vulnerability Scanner / AI red teaming kit 🧪
A powerful tool for automated LLM fuzzing. It is designed to help developers and security researchers identify and mitigate potential jailbreaks in their LLM AP
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI
Related tags
Frequently asked questions
How many open source llm-evaluation projects are there?
This registry tracks 14 projects tagged llm-evaluation, with 146,860 GitHub stars between them. The most-adopted is Langfuse at 34,729 stars.
Are these llm-evaluation projects free to use?
Yes — 12 of the 14 carry an explicit open-source licence across 2 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which llm-evaluation project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these llm-evaluation projects still maintained?
13 of the 14 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.