docetl is a free, open source data engineering & integration project written in Python and released under MIT. It has 4,097 GitHub stars, 443 forks and 45 open issues, and was last pushed 12 days ago. On this registry it ranks #16 of 39 tracked projects in Data Engineering & Integration, with 5 head-to-head comparisons available. It gained 4 stars over the last 3 tracked days.

What is docetl?

What it is

DocETL is a Python system for agentic LLM-powered data processing and ETL. It lives in the data engineering and document processing ecosystem, where pipelines must turn structured records and unstructured documents into queryable tables. Users describe each operation in natural language, and DocETL provides operators such as map, reduce, and filter, then orchestrates them across the data.

The concrete problem it addresses is manual LLM pipeline construction. Without DocETL, a developer writes each LLM call, wires calls together, and tunes accuracy, cost, and latency by hand. DocETL parallelizes work across records, optimizes pipelines by swapping models, rewriting prompts, decomposing operations, and replacing subtasks with code where possible, and returns tables that can be queried in a database.

Key capabilities

  • DocETL processes structured and unstructured data with LLMs, including document analysis, semantic data handling, and ETL or ELT workflows.
  • DocETL provides map, reduce, filter, resolve, split, gather, and extract operators for natural-language data processing tasks.
  • DocETL orchestrates and parallelizes pipeline work across records, so users do not manually wire each LLM call.
  • DocETL optimizes pipelines automatically for accuracy and cost by swapping models, rewriting prompts, decomposing operations, and replacing subtasks with code where possible.
  • DocETL supports a Python API with rate limits, model selection, schema inspection, preview runs, full collection, and cost reporting.
  • DocETL supports YAML low-code pipelines that declare datasets, operations, and output files, then run with docetl run pipeline.yaml.

Who uses it and how

  • Data engineers use it to build LLM-driven ETL or ELT pipelines over ticket, document, or file-based datasets, then write results to JSON or queryable tables.
  • Analysts and application developers use the Python API to classify records, reduce them by a key, summarize groups, inspect schemas, preview documents, and collect full results.
  • Low-code users define pipelines in YAML by naming datasets, default models, operations, and pipeline steps, then run the config from the command line.
  • Prompt developers use DocWrangler to edit prompts and see results in real time, hosted or locally, and users can get help from Claude Code or a copyable prompt.

Getting started

Install with pip install docetl, set an OpenAI or other LLM provider API key, then use the Python API, run a YAML pipeline with docetl run pipeline.yaml, or try DocWrangler at docetl.org/playground.

When to use it — and when not to

Use DocETL when the work is LLM processing over records or documents and users want declarative pipelines, automatic optimization, and tables that can be queried downstream. Do not use it as a database or file storage service; users still provide input files or datasets, LLM provider credentials, and a query destination. Because pipelines depend on model calls, rate limits, provider keys, and cost tracking, self-hosters must operate those dependencies and monitor spend, latency, and output quality.

project readme (upstream, from github) — read inline

DocETL: Declarative & Agentic Map-Reduce

Website Documentation Discord License: MIT

What is DocETL · Install · Python API · YAML · DocWrangler UI · Docs


What is DocETL

DocETL helps you process large collections of data (structured and unstructured) with LLMs. You write each operation in natural language, e.g., "pull out every complaint in this ticket," and DocETL

  • provides the operators you need (map, reduce, filter, and more) and orchestrates them, parallelizing work across your data,
  • optimizes your pipeline automatically, swapping models, rewriting prompts, decomposing operations, and replacing subtasks with code wherever possible, to raise accuracy and cut cost, and
  • returns tables, easy to query in your favorite database.

Without DocETL, you write each LLM call yourself, wire them together, and tune the result for accuracy, cost, and latency by hand.

CLI
DocWrangler UI

Install

pip install docetl
export OPENAI_API_KEY=your_key   # or any LLM provider key

Need Help Writing Your Pipeline?

Use Claude Code (recommended): run docetl install-skill and describe your task. See the quickstart.

If you'd rather use ChatGPT or the Claude app, copy the prompt at docetl.org/llms-full.txt into the chat before describing your task.


Python API (recommended)

Best for production code, notebooks, and scripting. Full guide

import docetl

docetl.default_model = "gpt-4o-mini"
docetl.rate_limits = {
    "llm_call": [{"count": 500, "per": 1, "unit": "minute"}],
    "llm_tokens": [{"count": 200_000, "per": 1, "unit": "minute"}],
}

# Classify support tickets, then summarize each category
pipeline = docetl.read_json("tickets.json")

pipeline = pipeline.map(
    prompt="Classify this support ticket: {{ input.text }}",
    output={"schema": {"category": "str", "priority": "str"}},
)

pipeline = pipeline.reduce(
    reduce_key="category",
    prompt="Summarize these tickets: {% for t in inputs %}{{ t.text }}{% endfor %}",
    output={"schema": {"summary": "str"}},
)

pipeline.schema()  # {'category': 'str', 'summary': 'str'}
pipeline.show()  # run on 5 docs and print results
rows = pipeline.collect()  # full run
print(f"Cost: ${pipeline.total_cost:.4f}")

YAML (low-code)

Declare your pipeline in a config file, no Python needed. Tutorial

datasets:
  tickets:
    type: file
    path: tickets.json

default_model: gpt-4o-mini

operations:
  - name: classify
    type: map
    prompt: "Classify this support ticket and assign a priority level."
    output:
      schema:
        category: str
        priority: str

pipeline:
  steps:
    - name: triage
      input: tickets
      operations: [classify]
  output:
    type: file
    path: output.json
docetl run pipeline.yaml

DocWrangler UI

Visual playground for interactive prompt development. Edit prompts, see results in real time. Try it at docetl.org/playground or run it locally.


Documentation

Python API Guide Frame API reference: operations, config, optimization
YAML Tutorial Step-by-step walkthrough of declarative pipelines
Operators Map, filter, reduce, resolve, split, gather, extract, and more
Optimization Automatic cost-accuracy optimization with MOAR
DocWrangler Setup Run the interactive UI locally or via Docker
Claude Code Quick Start Describe your task and let Claude build the pipeline

Community

Discord · Conversation Generator · Text-to-Speech · YouTube Transcript Topics


Development

git clone https://github.com/ucbepic/docetl.git && cd docetl
make install
make tests-basic  # < $0.01 with OpenAI

Papers

DocETL was created at the EPIC Data Lab and Data Systems and Foundations group at UC Berkeley.

DocETL, VLDB 2025 (paper)

@article{shankar2025docetl,
  title={DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing},
  author={Shankar, Shreya and Chambers, Tristan and Shah, Tarak and Parameswaran, Aditya G and Wu, Eugene},
  journal={Proceedings of the VLDB Endowment},
  volume={18}, number={9}, pages={3035--3048}, year={2025}
}

DocWrangler, UIST 2025, Best Paper Honorable Mention (paper)

@inproceedings{shankar2025docwrangler,
  title={Steering Semantic Data Processing With DocWrangler},
  author={Shankar*, Shreya and Chopra*, Bhavya and Hasan, Mawil and Lee, Stephen and Hartmann, Bj{\"o}rn and Hellerstein, Joseph M and Parameswaran, Aditya G and Wu, Eugene},
  booktitle={Proceedings of the ACM Symposium on User Interface Software and Technology (UIST)},
  year={2025}
}

MOAR, VLDB 2026 (paper)

@article{wei2026moar,
  title={Multi-Objective Agentic Rewrites for Unstructured Data Processing},
  author={Wei*, Lindsey Linxi and Shankar*, Shreya and Zeighami, Sepanta and Chung, Yeounoh and Ozcan, Fatma and Parameswaran, Aditya G},
  journal={Proceedings of the VLDB Endowment}, year={2026}
}

*Co-first authors

Frequently asked questions

Is docetl free to use?

docetl is open source under the MIT licence. There is no licence fee and no seat count — you can self-host it or, where the project offers one, pay a vendor for a managed version instead.

What does docetl do?

A system for agentic LLM-powered data processing and ETL

What is docetl written in?

docetl is primarily written in Python. Its source is publicly available at https://github.com/ucbepic/docetl, and it has 4,097 GitHub stars.