Open source llm-serving projects
Every project in the registry tagged llm-serving, ranked by real GitHub adoption.
A high-throughput and memory-efficient inference and serving engine for LLMs
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Related tags
Frequently asked questions
How many open source llm-serving projects are there?
This registry tracks 8 projects tagged llm-serving, with 123,241 GitHub stars between them. The most-adopted is vllm at 92,028 stars.
Are these llm-serving projects free to use?
Yes — 8 of the 8 carry an explicit open-source licence across 1 distinct licence, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which llm-serving project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these llm-serving projects still maintained?
7 of the 8 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.