Open source evals projects
Every project in the registry tagged evals, ranked by real GitHub adoption.
Open-source LLM tracing & evaluation for AI optimization
Powerful observability made simple for developers
AI-powered platform for engineering LLM products
Waku Waku! Waku Agent is a local-first AI agent harness you actually own, including loop, memory, eval, all in code built to stay legible as it grows.
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI
Related tags
Frequently asked questions
How many open source evals projects are there?
This registry tracks 5 projects tagged evals, with 22,443 GitHub stars between them. The most-adopted is Arize Phoenix at 11,518 stars.
Are these evals projects free to use?
Yes — 4 of the 5 carry an explicit open-source licence across 2 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which evals project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these evals projects still maintained?
5 of the 5 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.