Open source model-serving projects
Every project in the registry tagged model-serving, ranked by real GitHub adoption.
A framework for efficient model inference with omni-modality models
Open-Source Personal Cloud OS for Always-On Agents
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
OpenLake is a high performance storage engine for efficient LLM inference and GPU Training
🏕️ Reproducible development environment for humans and agents
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
Serverless LLM Serving for Everyone.
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Related tags
Frequently asked questions
How many open source model-serving projects are there?
This registry tracks 9 projects tagged model-serving, with 26,296 GitHub stars between them. The most-adopted is vllm-omni at 6,855 stars.
Are these model-serving projects free to use?
Yes — 9 of the 9 carry an explicit open-source licence across 2 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which model-serving project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these model-serving projects still maintained?
8 of the 9 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.