category
Open Source Data & Analytics Tools
Warehouses, pipelines and BI without the per-query invoice
All Data & Analytics tools
Text-based diagramming that renders in code and docs
LLM-ready web crawler built for AI data pipelines
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDce
Self-serve data exploration and dashboard builder
Apache ECharts is a powerful, interactive charting and data visualization library for browser
Scrapy, a fast high-level web crawling & scraping framework for Python.
Lightning-fast analytics for massive datasets
Unlock data insights with powerful, user-friendly analytics
The HTML5 Creation Engine: Create beautiful digital content with the fastest, most flexible 2D WebGL renderer.
A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。
38 editorial diagram types for Claude Code, Codex, and Pi. Self-contained HTML + SVG. No shadows. No Mermaid slop.
Unlock Product Insights with Open-Source Analytics
Empower your website with privacy-focused analytics
Best and simplest tool for website change detection, web page monitoring, and website change alerts. Perfect for tracking content changes, price drops, restock
👾 Fast and simple video download library and CLI tool written in Go
Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.
Python scraper based on AI
Privacy-focused, lightweight web analytics
Connect, query, visualize, and share your data
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or G
Elegant Scraper and Crawler Framework for Golang
🔥 人人可用的开源 BI 工具,数据可视化神器。An open-source BI tool alternative to Tableau.
Data Apps & Dashboards for Python. No JavaScript Required.
The SDK to extract data and interact with any site on the web. Get started with Claude Code, Codex, Eve, Mastra, and more.
Python ProxyPool for web spider
Open-source data integration for modern teams
Empower your digital strategy with actionable insights
Unify data models and metrics across your entire stack
Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.
🚀 Self-hosted TikTok & Douyin scraper and no-watermark video downloader — async REST API, MCP server, CLI and web console for posts, profiles, comments and pla
Open-source JavaScript charting library behind Plotly and Dash
A next-generation crawling and spidering framework.
No-code web scraping, crawling, and extraction platform
WebGL2 powered visualization framework
Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.
APIs for browser automation, testing, and bypassing bot-detection. Includes CDP Mode: A stealthy configuration for chromium that passes every bot detection test
Understand your audience with privacy-first analytics.
♾ A Graph Visualization Framework in JavaScript.
Distributed web crawler admin platform for spiders management regardless of languages and frameworks. 分布式爬虫管理平台,支持任何语言和框架
A JavaScript library aimed at visualizing graphs of thousands of nodes and edges
A scalable web crawler framework for Java.
Ultra-fast data transformation for AI with lineage
jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety.
Stealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.
High-performance browser automation bridge and multi-instance orchestrator with advanced stealth injection and real-time dashboard.
SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, P
Real-time Claude Code usage monitor with predictions and warnings
Privacy-focused analytics without the complexity
Zero-ETL, infinite possibilities. Live query APIs, code & more with SQL. No DB required.
Broadcast, Presence, and Postgres Changes via WebSockets
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
Python API for JMComic | 提供Python API访问禁漫天堂,同时支持网页端和移动端 | 禁漫天堂GitHub Actions下载器🚀
A Chrome DevTools Protocol driver for web automation and scraping.
Pydoll is a library for automating chromium-based browsers without a WebDriver, offering realistic interactions.
Transform SQL into beautiful data stories
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Web Crawler/Spider for NodeJS + server-side jQuery ;-)
Firebase SDK for Apple App Development
Sync and transform data from any source to any destination
Python binding for curl-impersonate fork via cffi. A http client that can impersonate browser tls/ja3/http2 fingerprints.
Flink CDC is a streaming data integration tool
Open-source BI for modern data teams
Distributed OLAP database for real-time analytics at scale
Declarative data automation language and Go runtime for structured extraction workflows.
Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, indepen
Comprehensive analytics for data-driven decisions
scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,emai
Allure Report is a flexible, lightweight multi-language test reporting tool. It provides clear graphical reports and allows everyone involved in the development
🦀 event stream processing for developers to collect and transform data in motion to power responsive data intensive applications.
Where data access meets operational intelligence
DotnetSpider, a .NET standard web crawling library. It is lightweight, efficient and fast high-level web crawling & scraping framework
pandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, Neptune, OpenSearch, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager,
A data integration framework
A system for agentic LLM-powered data processing and ETL
Headless Chrome .NET API
Scalable and efficient data transformation framework - backwards compatible with dbt.
Fast, Simple and a cost effective tool to replicate data from Postgres to Data Warehouses, Queues and Storage
Modern, privacy-friendly, and detailed web analytics that works without cookies or JS.
Apache DevLake is an open-source dev data platform to ingest, analyze, and visualize the fragmented data from DevOps tools, extracting insights for engineering
Tianji: Insight into everything, Website Analytics + Uptime Monitor + Server Status. not only another GA alternatives
Free Open Source Reporting tool for .NET6/.NET Core/.NET Framework that helps your application generate document-like reports
Python, SQL, and AI in a collaborative data notebook
The fastest business intelligence tool for humans and agents.
Powerful, customizable web analytics for your site
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
Convert wearable data into actionable health insights with open algorithms
API for Current cases and more stuff about COVID-19 and Influenza
Open-source, self-hosted revenue-first analytics for founders: web analytics, Session Replay, revenue attribution, and customer revenue integrations. datafast a
Data observability for modern data teams
LLM based data scientist, AI native data application. AI-driven infinite thinking redefines BI.
Python package for scraping recipes data
A lightweight stream processing library for Go
CLI task management & automation tool
CherryUSB is a tiny and beautiful, high performance and portable USB host and device stack for embedded system with USB IP
ARA Records Ansible and makes it easier to understand and troubleshoot.
The SpecterOps project management and reporting engine
To extract article from given URL
Privacy-first analytics for mobile and desktop apps
Cookie-less analytics for privacy-focused insights
A modern, scalable analytics system
🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
Database Reporting Tool and Tasks (.Net)
A Python stream processing engine modeled after Yahoo! Pipes
The open document intelligence platform for builders and hackers - DMS for the agentic world
Hop Orchestration Platform
OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestio
Actively maintained fork of Alibaba DataX — a fast, versatile ETL tool for RDBMS/NoSQL data transfer
Self-hosted SERP API for Google, Bing, Yandex, and more
JasperReports® - Free Java Reporting Library
Lightweight library for scraping web-sites with LLMs
javascript based business reporting platform :rocket:
Kibana Alert & Report App for Elasticsearch
Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality,
The Data Change Processing platform
Cookieless, privacy-first web analytics without the complexity
Privacy-first analytics without cookies or tracking
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scra
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry
Powerful open-source business intelligence for data-driven decisions
zerocode-tdd is a community-developed, free, open-source, outcome-driven automated testing framework for Data Pipelines, ETL, REST API, Kafka(Data Streams), Dat
Offen Fair Web Analytics
Power BI AI skills and Power BI agents for Claude Code and GitHub Copilot: a plugin marketplace of Power BI skills, subagents, and hooks for semantic models, DA
SeaTunnel is a distributed, high-performance data integration platform for the synchronization and transformation of massive data (offline & real-time).
Allure integrations for Python test frameworks
MasterParser is a powerful DFIR tool designed for analyzing and parsing Linux logs
This docker container allows you to see up to date reports simply mounting your "allure-results" directory in the container (for a Single Project) or your "proj
Web frontend for PuppetDB
Windows Event Log tooling for PowerShell and .NET: typed queries, reporting, export, WEC, automation, and the PSEventViewer module.
🤖 AI-powered web scraping editor with visual workflow builder. Build, test & deploy web scrapers using natural language. Powered by ScrapeGraphAI & LangGraph.
OpenChatBI is an intelligent chat-based BI tool powered by large language models, designed to help users query, analyze, and visualize data through natural lang
Countly Digital Analytics iOS SDK with macOS, watchOS and tvOS support.
Conduit streams data between data stores. Kafka Connect replacement. No JVM required.
Pen Test Report Generation and Assessment Collaboration
Undetected web-scraping & seamless HTML parsing in Python!
Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents.
Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for ever
Open-source, privacy-first and cookieless web analytics with revenue attribution and an MCP server.
Simple, powerful web & product analytics for user insights
Professional link management for modern teams
Seamless data import for your applications
Privacy-first, cookieless analytics without consent banners
Privacy-friendly web analytics for data-driven decisions
Privacy-first analytics you own completely
Frequently asked questions
How many open source Data & Analytics tools are there?
This registry tracks 145 open source Data & Analytics projects, with 1,716,734 combined GitHub stars. The list is ranked by stars and refreshed nightly.
What is the most popular open source Data & Analytics project?
Mermaid leads this category with 90,293 GitHub stars, followed by Crawl4AI.
Are these Data & Analytics tools free?
Yes — every project listed here is open source. Some also offer paid hosted versions alongside the free self-hosted option; the licence for each project is shown on its card.