Open source evaluation projects
Every project in the registry tagged evaluation, ranked by real GitHub adoption.
Open source LLM engineering platform for AI-powered applications
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simpl
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready d
Build reliable AI apps with comprehensive observability
Simulation-based testing and evaluation for AI agents
AI-powered platform for engineering LLM products
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI
Related tags
Frequently asked questions
How many open source evaluation projects are there?
This registry tracks 7 projects tagged evaluation, with 97,700 GitHub stars between them. The most-adopted is Langfuse at 34,729 stars.
Are these evaluation projects free to use?
Yes — 6 of the 7 carry an explicit open-source licence across 2 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which evaluation project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these evaluation projects still maintained?
7 of the 7 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.