matrixone is a free, open source machine learning infrastructure project written in Go and released under Apache-2.0. It has 1,889 GitHub stars, 312 forks and 745 open issues, and was last pushed 3 hours ago. On this registry it ranks #53 of 80 tracked projects in Machine Learning Infrastructure, with 5 head-to-head comparisons available.

MatrixOne All in One

license language macos linux codefactor release
mysql-compatible ai-native cloud-native
Docs || Official Website || Research Paper
English || 简体中文

Connect with us:

matrixone16 matrixone16

Contents

What is MatrixOne?

MatrixOne is the industry's first database to bring Git-style version control to data, combined with MySQL compatibility, AI-native capabilities, and cloud-native architecture.

At its core, MatrixOne is a HTAP (Hybrid Transactional/Analytical Processing) database with a hyper-converged HSTAP engine that seamlessly handles transactional (OLTP), analytical (OLAP), full-text search, and vector search workloads in a single unified system—no data movement, no ETL, no compromises.

🎬 Git for Data - The Game Changer

Just as Git revolutionized code management, MatrixOne brings Git-style workflows to data management. The design behind this capability is detailed in the arXiv paper Version Control System for Data with MatrixOne, and in practice it lets you manage your database like code:

  • 📸 Instant Snapshots - Zero-copy snapshots in milliseconds, no storage explosion
  • ⏰ Time Travel - Query data as it existed at any point in history
  • 🔀 Branch & Merge - Test migrations and transformations in isolated branches
  • ↩️ Instant Rollback - Restore to any previous state without full backups
  • 🔍 Complete Audit Trail - Track every data change with immutable history

Why it matters: Data mistakes are expensive. Git for Data gives you the safety net and flexibility developers have enjoyed with Git—now for your most critical asset: your data.


🎯 Built for the AI Era

🗄️ MySQL-Compatible

Drop-in replacement for MySQL. Use existing tools, ORMs, and applications without code changes. Seamless migration path.

🤖 AI-Native

Built-in vector search (IVF/HNSW) and full-text search. Build RAG apps and semantic search directly—no external vector databases needed.

☁️ Cloud-Native

Storage-compute separation. Deploy anywhere. Elastic scaling. Kubernetes-native. Zero-downtime operations.


🚀 One Database for Everything

The typical modern data stack:

🗄️ MySQL for transactions → 📊 ClickHouse for analytics → 🔍 Elasticsearch for search → 🤖 Pinecone for AI

The problem: 4 databases · Multiple ETL jobs · Hours of data lag · Sync nightmares

MatrixOne replaces all of them:

🎯 One database with native OLTP, OLAP, full-text search, and vector search. Real-time. ACID compliant. No ETL.

MatrixOne

⚡️ Get Started in 60 Seconds

1️⃣ Launch MatrixOne

docker run -d -p 6001:6001 --name matrixone matrixorigin/matrixone:latest

2️⃣ Create Database

mysql -h127.0.0.1 -P6001 -p111 -uroot -e "create database demo"

3️⃣ Connect & Query

Install Python SDK:

pip install matrixone-python-sdk

Vector search:

from matrixone import Client
from matrixone.orm import declarative_base
from sqlalchemy import Column, Integer, String, Text
from matrixone.sqlalchemy_ext import create_vector_column

# Create client and connect
client = Client()
client.connect(database='demo')

# Define model using MatrixOne ORM
Base = declarative_base()

class Article(Base):
    __tablename__ = 'articles'
    id = Column(Integer, primary_key=True, autoincrement=True)
    title = Column(String(200), nullable=False)
    content = Column(Text, nullable=False)
    embedding = create_vector_column(8, "f32")

# Create table using client API
client.create_table(Article)

# Insert some data using client API
articles = [
    {'title': 'Machine Learning Guide',
     'content': 'Comprehensive machine learning tutorial...',
     'embedding': [0.1, 0.2, 0.3, 0.15, 0.25, 0.35, 0.12, 0.22]},
    {'title': 'Python Programming',
     'content': 'Learn Python programming basics',
     'embedding': [0.2, 0.3, 0.4, 0.25, 0.35, 0.45, 0.22, 0.32]},
]
client.batch_insert(Article, articles)

client.vector_ops.create_ivf(
    Article,
    name='idx_embedding',
    column='embedding',
    lists=100,
    op_type='vector_l2_ops'
)

query_vector = [0.2, 0.3, 0.4, 0.25, 0.35, 0.45, 0.22, 0.32]
results = client.query(
    Article.title,
    Article.content,
    Article.embedding.l2_distance(query_vector).label("distance"),
).filter(Article.embedding.l2_distance(query_vector) < 0.1).execute()
for row in results.rows:
    print(f"Title: {row[0]}, Content: {row[1][:50]}...")

# Cleanup
client.drop_table(Article)  # Use client API
client.disconnect()

Fulltext Search:

...
from matrixone.sqlalchemy_ext import boolean_match

# Create fulltext index using SDK 
client.fulltext_index.create(
    Article,name='ftidx_content',columns=['title', 'content']
)

# Boolean search with must/should operators
results = client.query(
    Article.title,
    Article.content,
    boolean_match('title', 'content')
        .must('machine')
        .must('learning')
        .must_not('basics')
).execute()

# Results is a ResultSet object
for row in results.rows:
    print(f"Title: {row[0]}, Content: {row[1][:50]}...")
...

That's it! 🎉 You're now running a production-ready database with Git-like snapshots, vector search, and full ACID compliance.

💡 Want more control? Check out the Installation & Deployment section below for production-grade installation options.

📖 Python SDK Documentation →

📚 Tutorials & Demos

Ready to dive deeper? Explore our comprehensive collection of hands-on tutorials and real-world demos:

🎯 Getting Started Tutorials

Tutorial Language/Framework Description
Java CRUD Demo Java Java application development
SpringBoot and JPA CRUD Demo Java SpringBoot with Hibernate/JPA
[PyMySQL CRUD Dem

readme truncated — read the full docs on github

Frequently asked questions

Is matrixone free to use?

matrixone is open source under the Apache-2.0 licence. There is no licence fee and no seat count — you can self-host it or, where the project offers one, pay a vendor for a managed version instead.

What does matrixone do?

AI-native HTAP database with Git-for-Data and built-in vector search, serving as the data and memory backbone for intelligent agents and applications.

What is matrixone written in?

matrixone is primarily written in Go. Its source is publicly available at https://github.com/matrixorigin/matrixone, and it has 1,889 GitHub stars.