Ceph is a free, open source storage solutions project written in C++ and released under a custom open-source licence. It has 17,050 GitHub stars, 6,504 forks and 1,353 open issues, and was last pushed 6 hours ago. On this registry it ranks #3 of 8 tracked projects in Storage Solutions, with 5 head-to-head comparisons available. It gained 21 stars over the last 6 tracked days.

Ceph — Scalable open-source distributed storage system

What is Ceph?

What it is

Ceph is a scalable open-source distributed storage system written in C++ and described as a platform for object, block, and file storage. It lives in infrastructure and operations storage solutions, where it provides a shared storage layer across distributed nodes rather than a single local disk.

The concrete problem it solves is the need for one storage foundation that can serve different workload types while remaining available and scalable across a cluster. The project topics point to distributed storage, cloud storage, distributed file system, erasure coding, high availability, high performance, and integrations with Kubernetes, NFS, iSCSI, FUSE, and HDFS.

Key capabilities

  • Ceph provides object, block, and file storage from one distributed platform, so operators can use it for multiple storage interfaces instead of running separate systems.
  • It supports erasure coding, a storage technique listed among its topics for distributing data across nodes.
  • It exposes block storage through iSCSI, allowing clients to use Ceph-backed volumes over a network.
  • It supports file access through NFS, FUSE, and HDFS topics, covering common file-system integration paths.
  • It is associated with Kubernetes and cloud storage topics, indicating use in container and cloud-oriented storage deployments.
  • It is described as high-performance and highly-available, reflecting goals for scalable distributed storage clusters.

Who uses it and how

  • Storage teams that need a shared distributed backend for object, block, and file workloads use Ceph as the central platform.
  • Kubernetes operators can use Ceph in environments where container workloads need persistent storage, as indicated by the Kubernetes topic.
  • Hosts can mount Ceph-backed file data through FUSE, which is listed among the project topics.
  • Clients that need block volumes over a network can use the iSCSI integration listed in the topics.
  • File workloads that require NFS or HDFS-compatible access can use Ceph according to the NFS and HDFS topics.

Getting started

Clone the ceph/ceph repository, run ./install-deps.sh on Debian or Ubuntu after installing curl, then use ./do_cmake.sh and ninja to build the development environment. The README also says to install the vstart cluster with ninja install, and for production binaries it recommends building .deb or .rpm packages or consulting ceph.spec.in and debian/rules.

When to use it — and when not to

Use Ceph when you need a self-hosted distributed storage platform that can cover object, block, and file interfaces and you are willing to operate a C++ project with source-based build steps. Avoid it if you need a turnkey hosted service or a simple install path, because the README provides git clone, dependency installation, CMake, ninja, and package build instructions rather than a hosted option. Also review the mixed license inventory and the NOASSERTION repository license metadata before redistribution, and note that large parallel ninja builds can exhaust memory because each job needs about 2.5 GiB of RAM.

project readme (upstream, from github) — read inline

Ceph - a scalable distributed storage system

See https://ceph.com/ for current information about Ceph.

Status

OpenSSF Best Practices Issue Backporting

Contributing Code

Most of Ceph is dual-licensed under the LGPL version 2.1 or 3.0. Some miscellaneous code is either public domain or licensed under a BSD-style license.

The Ceph documentation is licensed under Creative Commons Attribution Share Alike 3.0 (CC-BY-SA-3.0).

Some headers included in the ceph/ceph repository are licensed under the GPL. See the file COPYING for a full inventory of licenses by file.

All code contributions must include a valid "Signed-off-by" line. See the file SubmittingPatches.rst for details on this and instructions on how to generate and submit patches.

Assignment of copyright is not required to contribute code. Code is contributed under the terms of the applicable license.

Checking out the source

Clone the ceph/ceph repository from github by running the following command on a system that has git installed:

git clone [email protected]:ceph/ceph

Alternatively, if you are not a github user, you should run the following command on a system that has git installed:

git clone https://github.com/ceph/ceph.git

When the ceph/ceph repository has been cloned to your system, run the following commands to move into the cloned ceph/ceph repository and to check out the git submodules associated with it:

cd ceph
git submodule update --init --recursive --progress

Build Prerequisites

section last updated 06 Sep 2024

We provide the Debian and Ubuntu apt commands in this procedure. If you use a system with a different package manager, then you will have to use different commands.

#. Install curl:

apt install curl

#. Install package dependencies by running the install-deps.sh script:

./install-deps.sh

#. Install the python3-routes package:

apt install python3-routes

Building Ceph

These instructions are meant for developers who are compiling the code for development and testing. To build binaries that are suitable for installation we recommend that you build .deb or .rpm packages, or refer to ceph.spec.in or debian/rules to see which configuration options are specified for production builds.

To build Ceph, follow this procedure:

  1. Make sure that you are in the top-level ceph directory that contains do_cmake.sh and CONTRIBUTING.rst.

  2. Run the do_cmake.sh script:

    ./do_cmake.sh
    

    See build types.

  3. Move into the build directory:

    cd build
    
  4. Use the ninja buildsystem to build the development environment:

    ninja -j3
    

    [!IMPORTANT]

    Ninja is the build system used by the Ceph project to build test builds. The number of jobs used by ninja is derived from the number of CPU cores of the building host if unspecified. Use the -j option to limit the job number if build jobs are running out of memory. If you attempt to run ninja and receive a message that reads g++: fatal error: Killed signal terminated program cc1plus, then you have run out of memory.

    Using the -j option with an argument appropriate to the hardware on which the ninja command is run is expected to result in a successful build. For example, to limit the job number to 3, run the command ninja -j3. On average, each ninja job run in parallel needs approximately 2.5 GiB of RAM.

    This documentation assumes that your build directory is a subdirectory of the ceph.git checkout. If the build directory is located elsewhere, point CEPH_GIT_DIR to the correct path of the checkout. Additional CMake args can be specified by setting ARGS before invoking do_cmake.sh. See cmake options for more details. For example:

    ARGS="-DCMAKE_C_COMPILER=gcc-7" ./do_cmake.sh
    

    To build only certain targets, run a command of the following form:

    ninja [target name]
    
  5. Install the vstart cluster:

    ninja install
    

Build Types

do_cmake.sh by default creates a "debug build" of Ceph (assuming .git exists). A Debug build runtime performance may be as little as 20% of that of a non-debug build. Pass -DCMAKE_BUILD_TYPE=RelWithDebInfo to do_cmake.sh to create a non-debug build. The default build type is RelWithDebInfo once .git does not exist.

CMake mode Debug info Optimizations Sanitizers Checks Use for
Debug Yes -Og None ceph_assert, assert gdb, development
RelWithDebInfo Yes -O2, -DNDEBUG None ceph_assert only production

CMake Options

The -D flag can be used with cmake to speed up the process of building Ceph and to customize the build.

Building without RADOS Gateway

The RADOS Gateway is built by default. To build Ceph without the RADOS Gateway, run a command of the following form:

cmake -DWITH_RADOSGW=OFF [path to top-level ceph directory]
Building with debugging and arbitrary dependency locations

Run a command of the following form to build Ceph with debugging and alternate locations for some external dependencies:

cmake -DCMAKE_INSTALL_PREFIX=/opt/ceph -DCMAKE_C_FLAGS="-Og -g3 -gdwarf-4" \
..

Ceph has several bundled dependencies such as Boost, RocksDB and Arrow. By default, cmake builds these bundled dependencies from source instead of using libraries that are already installed on the system. You can opt to use these system libraries, as long as they meet Ceph's version requirements. To use system libraries, use cmake options like WITH_SYSTEM_BOOST, as in the following example:

cmake -DWITH_SYSTEM_BOOST=ON [...]

To view an exhaustive list of -D options, invoke cmake -LH:

cmake -LH
Preserving diagnostic colors

If you pipe ninja to less and would like to preserve the diagnostic colors in the output in order to make errors and warnings more legible, run the following command:

cmake -DDIAGNOSTICS_COLOR=always ...

The above command works only with supported compilers.

The diagnostic colors will be visible when the following command is run:

ninja | less -R

Other available values for DIAGNOSTICS_COLOR are auto (default) and never.

Tips and Tricks

  • Use "debug builds" only when needed. Debugging builds are helpful for development, but they can slow down performance. Use -DCMAKE_BUILD_TYPE=Release when debugging isn't necessary.
  • Enable Selective Daemons when testing specific components. Don't start unnecessary daemons.
  • Preserve Existing Data skip cluster reinitialization between tests by using the -n flag.
  • To manage a vstart cluster, stop daemons using ./stop.sh and start them with ./vstart.sh --daemon osd.${ID} [--nodaemonize].
  • Restart the sockets by stopping and restarting the daemons associated with them. This ensures that there are no stale sockets in the cluster.
  • To track RocksDB performance, set export ROCKSDB_PERF=true and start the cluster by using the command ./vstart.sh -n -d -x --bluestore.
  • Build with vstart-base using debug flags in cmake, compile, and deploy via ./vstart.sh -d -n --bluestore.
  • To containerize, generate configurations with vstart.sh, and deploy with Docker, mapping directories and configuring the network.
  • Manage containers using docker run, stop, and rm. For detailed setups, consult the Ceph-Container repository.

Troubleshooting

  • Cluster Fails to Start: Look for errors in the logs under the out/ directory.
  • OSD Crashes: Check the OSD logs for errors.
  • Cluster in a Health Error State: Run the ceph status command to identify the issue.
  • RocksDB Errors: Look for RocksDB-related errors in the OSD logs.

Building a source tarball

To build a complete source tarball with everything needed to build from source and/or build a (deb or rpm) package, run

./make-dist

This will create a tarball like ceph-$version.tar.bz2 from git. (Ensure that any changes you want to include in your working directory are committed to git.)

Running a test cluster

From the ceph/ directory, run the following commands to launch a test Ceph cluster:

cd build
ninja vstart        # builds just enough to run vstart
../src/vstart.sh --debug --new -x --localhost 

readme truncated — read the full docs on github

Frequently asked questions

Is Ceph free to use?

Ceph is open source. There is no licence fee and no seat count — you can self-host it or, where the project offers one, pay a vendor for a managed version instead.

What does Ceph do?

Scalable open-source distributed storage system

What is Ceph written in?

Ceph is primarily written in C++. Its source is publicly available at https://github.com/ceph/ceph, and it has 17,050 GitHub stars.