Open source llm-evaluation projects

Every project in the registry tagged llm-evaluation, ranked by real GitHub adoption.

projects 14 combined stars ★ 147K refresh nightly
01 Langfuse ★ 35K

Open source LLM engineering platform for AI-powered applications

last push3 hours ago languageTypeScript license
02 promptfoo ★ 25K

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simpl

last push4 hours ago languageTypeScript licenseMIT
03 opik ★ 22K

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready d

last push4 hours ago languagePython licenseApache-2.0
04 iFixAi ★ 15K

Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is sup

last push35 hours ago languagePython licenseApache-2.0
05 Arize Phoenix ★ 12K

Open-source LLM tracing & evaluation for AI optimization

last push3 hours ago languagePython license
06 garak ★ 9.3K

the LLM vulnerability scanner

last push25 hours ago languagePython licenseApache-2.0
07 AI-Infra-Guard ★ 6.4K

A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.

last push14 hours ago languagePython licenseApache-2.0
08 Helicone ★ 6.2K

Build reliable AI apps with comprehensive observability

last push28 hours ago languageTypeScript licenseApache-2.0
09 giskard-oss ★ 5.8K

🐢 Open-Source Evaluation & Testing library for LLM Agents

last push37 hours ago languagePython licenseApache-2.0
10 Laminar ★ 3.3K

AI-powered platform for engineering LLM products

last push5 hours ago languageTypeScript licenseApache-2.0
11 future-agi ★ 2.0K

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Gu

last push9 hours ago languagePython licenseApache-2.0
12 agentic_security ★ 2.0K

Agentic LLM Vulnerability Scanner / AI red teaming kit 🧪

last push7 days ago languagePython licenseApache-2.0
13 FuzzyAI ★ 1.6K

A powerful tool for automated LLM fuzzing. It is designed to help developers and security researchers identify and mitigate potential jailbreaks in their LLM AP

last push7 months ago languageJupyter Notebook licenseApache-2.0
14 Tracely-ai ★ 1.4K

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI

last push34 hours ago languagePython licenseMIT

Related tags

← all tags

Frequently asked questions

How many open source llm-evaluation projects are there?

This registry tracks 14 projects tagged llm-evaluation, with 146,860 GitHub stars between them. The most-adopted is Langfuse at 34,729 stars.

Are these llm-evaluation projects free to use?

Yes — 12 of the 14 carry an explicit open-source licence across 2 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.

Which llm-evaluation project should I choose?

The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.

Are these llm-evaluation projects still maintained?

13 of the 14 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.