cassandra is a free, open source databases project written in Java and released under Apache-2.0. It has 10,096 GitHub stars, 4,118 forks and 529 open issues, and was last pushed 9 hours ago. On this registry it ranks #80 of 203 tracked projects in Databases, with 5 head-to-head comparisons available.

What is cassandra?

Apache Cassandra is an open source transactional distributed database written in Java and licensed under Apache-2.0, for users who need linear scalability and proven fault-tolerance on commodity hardware or cloud infrastructure without compromising performance.

What it is

Apache Cassandra is a highly-scalable partitioned row store. Rows are organized into tables with a required primary key, and the database organizes data by rows and columns in a manner similar to relational databases. It lives in the Java ecosystem: Java is required, Python is needed for cqlsh, and the project is listed under Infrastructure & Operations / Databases with the topics cassandra, database, and java. The Cassandra Query Language (CQL) is a close relative of SQL.

The concrete problem Cassandra solves is distribution of rows across multiple machines without forcing the application to manage placement. Partitioning means Cassandra can distribute data across multiple machines in an application-transparent manner, and it will automatically repartition as machines are added to or removed from the cluster. That replaces the application-managed sharding work that would otherwise be required to spread a row store over more than one machine.

Key capabilities

  • Partitioned row store: rows are organized into tables with a required primary key, and data can be distributed across multiple machines in an application-transparent manner.
  • Automatic repartitioning when machines are added to or removed from the cluster.
  • CQL (Cassandra Query Language), a close relative of SQL, with commands such as CREATE KEYSPACE, USE, CREATE TABLE, INSERT, and SELECT.
  • bin/cassandra -f starts the server in the foreground and logs to standard out; ctrl-C stops it.
  • bin/cqlsh provides an interactive command line client; its banner reports cqlsh 6.3.0, Cassandra 7.0-SNAPSHOT, CQL spec 3.4.8, and Native protocol v5.
  • Official Docker image cassandra on Docker Hub, alongside official downloads from the Apache Cassandra web site.
  • Apache-2.0 licence, with issues reported on the Apache Jira project CASSANDRA.

Who uses it and how

  • Developers evaluating the database stand up a one-node cluster by unpacking apache-cassandra-$VERSION.tar.gz, running bin/cassandra -f, and issuing CQL through bin/cqlsh.
  • Operators run clusters across multiple machines on commodity hardware or cloud infrastructure, relying on Cassandra to distribute rows and automatically repartition as machines join or leave.
  • Teams building row-store applications that need partitioning to remain application-transparent rather than manually sharded.
  • Users who want a CQL interface, close to SQL, for defining keyspaces, tables, and queries.
  • Projects that deploy through the official cassandra Docker image or through official Apache Cassandra downloads.

Getting started

Unpack apache-cassandra-$VERSION.tar.gz, then run bin/cassandra -f to start the server in the foreground and bin/cqlsh to open the interactive CQL client. Java must be available, and Python is required for cqlsh.

How it compares

The supplied facts do not name any similar tools or paid products for comparison. On the evidence provided, Apache Cassandra stands alone in this registry.

When to use it — and when not to

A self-hoster must operate the Java runtime and Python for cqlsh, and must run a cluster whose data is partitioned across machines. Cassandra is not a fit for teams that require standard SQL semantics rather than CQL, or that do not want to operate a distributed row store. The README excerpt is sparse and getting-started-focused, and the repository shows 529 open issues, so operators should expect to consult the full documentation and Jira rather than rely on the README alone.

project readme (upstream, from github) — read inline

image:https://img.shields.io/badge/License-Apache%202.0-blue.svg[License, link=https://github.com/apache/cassandra/blob/trunk/LICENSE.txt] image:https://ci-cassandra.apache.org/job/Cassandra-trunk/badge/icon[Build Status, link=https://ci-cassandra.apache.org/job/Cassandra-trunk/]     image:https://img.shields.io/badge/Official-Downloads-brightgreen[Official Downloads, link=https://cassandra.apache.org/$$_$$/download.html] image:https://img.shields.io/docker/pulls/$$_$$/cassandra[Docker Pulls, link=https://hub.docker.com/r/$$_$$/cassandra]     image:https://img.shields.io/badge/Slack-4A154B?style=flat&logo=slack&logoColor=white[Slack, link=https://infra.apache.org/slack.html] image:https://img.shields.io/badge/Bluesky-0285FF?logo=bluesky&logoColor=fff&color=0285FF[Bluesky, link=https://bsky.app/profile/cassandra.apache.org] image:https://img.shields.io/badge/-LinkedIn-blue?style=flat-square&logo=Linkedin&logoColor=white&link=https://www.linkedin.com/company/apache-cassandra/[LinkedIn, link=https://www.linkedin.com/company/apache-cassandra/] image:https://img.shields.io/badge/YouTube-FF0000?style=flat&logo=youtube&logoColor=white[Youtube, link=https://www.youtube.com/c/PlanetCassandra]

Apache Cassandra

Apache Cassandra is a highly-scalable partitioned row store. Rows are organized into tables with a required primary key.

https://cwiki.apache.org/confluence/display/CASSANDRA2/Partitioners[Partitioning] means that Cassandra can distribute your data across multiple machines in an application-transparent matter. Cassandra will automatically repartition as machines are added and removed from the cluster.

https://cwiki.apache.org/confluence/display/CASSANDRA2/DataModel[Row store] means that like relational databases, Cassandra organizes data by rows and columns. The Cassandra Query Language (CQL) is a close relative of SQL.

For more information, see https://cassandra.apache.org/[the Apache Cassandra web site].

Issues should be reported on https://issues.apache.org/jira/projects/CASSANDRA/issues/[The Cassandra Jira].

Requirements

  • Java: see supported versions in build.xml (search for property "java.supported").
  • Python: for cqlsh, see bin/cqlsh (search for function "is_supported_version").

Getting started

This short guide will walk you through getting a basic one node cluster up and running, and demonstrate some simple reads and writes. For a more-complete guide, please see the Apache Cassandra website's https://cassandra.apache.org/doc/latest/cassandra/getting-started/index.html[Getting Started Guide].

First, we'll unpack our archive:

$ tar -zxvf apache-cassandra-$VERSION.tar.gz $ cd apache-cassandra-$VERSION

After that we start the server. Running the startup script with the -f argument will cause Cassandra to remain in the foreground and log to standard out; it can be stopped with ctrl-C.

$ bin/cassandra -f

Now let's try to read and write some data using the Cassandra Query Language:

$ bin/cqlsh

The command line client is interactive so if everything worked you should be sitting in front of a prompt:


Connected to Test Cluster at localhost:9160. [cqlsh 6.3.0 | Cassandra 7.0-SNAPSHOT | CQL spec 3.4.8 | Native protocol v5] Use HELP for help. cqlsh>

As the banner says, you can use 'help;' or '?' to see what CQL has to offer, and 'quit;' or 'exit;' when you've had enough fun. But lets try something slightly more interesting:


cqlsh> CREATE KEYSPACE schema1 WITH replication = { 'class' : 'SimpleStrategy', 'replication_factor' : 1 }; cqlsh> USE schema1; cqlsh:Schema1> CREATE TABLE users ( user_id varchar PRIMARY KEY, first varchar, last varchar, age int ); cqlsh:Schema1> INSERT INTO users (user_id, first, last, age) VALUES ('jsmith', 'John', 'Smith', 42); cqlsh:Schema1> SELECT * FROM users; user_id | age | first | last ---------+-----+-------+------- jsmith | 42 | john | smith

cqlsh:Schema1>

If your session looks similar to what's above, congrats, your single node cluster is operational!

For more on what commands are supported by CQL, see https://cassandra.apache.org/doc/trunk/cassandra/developing/cql/index.html[the CQL reference]. A reasonable way to think of it is as, "SQL minus joins and subqueries, plus collections."

Wondering where to go from here?

Frequently asked questions

Is cassandra free to use?

cassandra is open source under the Apache-2.0 licence. There is no licence fee and no seat count — you can self-host it or, where the project offers one, pay a vendor for a managed version instead.

What does cassandra do?

Open source transactional distributed database. Linear scalability and proven fault-tolerance on commodity hardware or cloud infrastructure without compromising

What is cassandra written in?

cassandra is primarily written in Java. Its source is publicly available at https://github.com/apache/cassandra, and it has 10,096 GitHub stars.