
ContextGem: Effortless LLM extraction from documents

ContextGem is a free, open-source LLM framework that makes it radically easier to extract structured data and insights from documents — with minimal code.
💎 Why ContextGem?
Reliable structured extraction from documents typically involves writing extraction prompts, designing validation models, mapping outputs back to source references, orchestrating multi-step pipelines, and tracking usage across LLMs. ContextGem handles all of this through powerful abstractions — you describe what to extract in natural language, and the framework handles how.
The result: structured data with precise paragraph- and sentence-level references, automatic justifications, hierarchical multi-aspect extraction, and a unified, serializable document storage model — all from minimal code.
📖 Read more on the project motivation in the documentation.
⭐ Key features
| ✨ Automated dynamic prompts | 📐 Automated data modelling | 📍 Granular reference mapping |
| 💭 Built-in justifications | 🪆 Nested context extraction | 🔗 Unified declarative pipeline |
💡 What you can build
With minimal code, you can:
- Extract structured data from documents (text, images)
- Identify and analyze key aspects (topics, themes, categories) within documents (learn more)
- Extract specific concepts (entities, facts, conclusions, assessments) from documents (learn more)
- Build complex extraction workflows through a simple, intuitive API
- Create multi-level extraction pipelines (aspects containing concepts, hierarchical aspects)

📦 Installation
Using uv (recommended):
uv add contextgem
Or using pip:
pip install -U contextgem
🚀 Quick start
The following example demonstrates how to use ContextGem to extract anomalies from a legal document - a complex concept that requires contextual understanding. Unlike traditional RAG approaches that might miss subtle inconsistencies, ContextGem analyzes the entire document context to identify content that doesn't belong, complete with source references and justifications.
## Quick Start Example - Extracting anomalies from a document, with source references and justifications
import os
from contextgem import Document, DocumentLLM, StringConcept
## Sample document text (shortened for brevity)
doc = Document(
raw_text=(
"Consultancy Agreement\n"
"This agreement between Company A (Supplier) and Company B (Customer)...\n"
"The term of the agreement is 1 year from the Effective Date...\n"
"The Supplier shall provide consultancy services as described in Annex 2...\n"
"The Customer shall pay the Supplier within 30 calendar days of receiving an invoice...\n"
"The purple elephant danced gracefully on the moon while eating ice cream.\n" # 💎 anomaly
"Time-traveling dinosaurs will review all deliverables before acceptance.\n" # 💎 another anomaly
"This agreement is governed by the laws of Norway...\n"
),
)
## Attach a document-level concept
doc.concepts = [
StringConcept(
name="Anomalies", # in longer contexts, this concept is hard to capture with RAG
description="Anomalies in the document",
add_references=True,
reference_depth="sentences",
add_justifications=True,
justificati,
# see the docs for more configuration options
)
# add more concepts to the document, if needed
# see the docs for available concepts: StringConcept, JsonObjectConcept, etc.
]
## Or use `doc.add_concepts([...])`
## Define an LLM for extracting information from the document
llm = DocumentLLM(
model="openai/gpt-4o-mini", # or another provider/LLM
api_key=os.environ.get(
"CONTEXTGEM_OPENAI_API_KEY"
), # your API key for the LLM provider
# see the docs for mor