Open source evals projects

Every project in the registry tagged evals, ranked by real GitHub adoption.

projects 5 combined stars ★ 22K refresh nightly
01 Arize Phoenix ★ 12K

Open-source LLM tracing & evaluation for AI optimization

last push3 hours ago languagePython license
02 Logfire ★ 4.5K

Powerful observability made simple for developers

last push5 hours ago languagePython licenseMIT
03 Laminar ★ 3.3K

AI-powered platform for engineering LLM products

last push5 hours ago languageTypeScript licenseApache-2.0
04 waku-agent ★ 1.8K

Waku Waku! Waku Agent is a local-first AI agent harness you actually own, including loop, memory, eval, all in code built to stay legible as it grows.

last push5 hours ago languagePython licenseMIT
05 Tracely-ai ★ 1.4K

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI

last push34 hours ago languagePython licenseMIT

Related tags

← all tags

Frequently asked questions

How many open source evals projects are there?

This registry tracks 5 projects tagged evals, with 22,443 GitHub stars between them. The most-adopted is Arize Phoenix at 11,518 stars.

Are these evals projects free to use?

Yes — 4 of the 5 carry an explicit open-source licence across 2 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.

Which evals project should I choose?

The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.

Are these evals projects still maintained?

5 of the 5 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.