# Open-source LLM inference runtimes in 2026: llama.cpp vs Ollama vs vLLM vs SGLang
Direct answer
Choose llama.cpp for maximum control over GGUF models and mixed CPU/GPU hardware, Ollama for the simplest local developer experience, MLX-LM for Apple-silicon experimentation, MLC LLM for browser or mobile deployment, and vLLM or SGLang for concurrent GPU serving. TensorRT-LLM fits NVIDIA-specific optimization. Treat TGI as an existing-installation option because its repository is now in maintenance mode.
The useful comparison is not one benchmark table
People searching for “llama.cpp vs Ollama vs vLLM” usually want a quick winner. Current results commonly answer with a feature table, hardware matrix, model-format comparison and a decision tree. That structure is useful, but a single speed ranking is misleading. These projects occupy different layers, and their performance changes with model architecture, quantization, prompt length, batch size, concurrency, accelerator, memory bandwidth and request shape.
The first decision is therefore not “which runtime has the highest tokens per second?” It is “what job must the system perform?” A private assistant on a laptop, an API for twenty concurrent users, an iPhone feature and a multi-GPU agent platform are four different systems. The same model weights can behave very differently because the surrounding engine schedules memory and requests differently.
This guide compares the projects using current upstream documentation reviewed on 3 October 2026. It does not reproduce their marketing benchmarks or pretend that results measured on one machine transfer to another. Instead, it maps each stack to the workload it is designed to serve and gives you a repeatable way to validate the final choice.
Quick decision table
| Runtime | Best starting point for | Main strength | Main trade-off |
|---|---|---|---|
| llama.cpp | GGUF on laptops, desktops and mixed CPU/GPU systems | Broad hardware backends, quantization and fine control | More knobs and operational assembly |
| Ollama | Developers who want local models working quickly | Model lifecycle, simple CLI and HTTP API | Less direct control than configuring the underlying engine yourself |
| MLX-LM | Apple-silicon research and application prototypes | Native MLX workflow, conversion, quantization and fine-tuning | Apple-specific rather than a general deployment layer |
| MLC LLM | Browser, mobile and heterogeneous edge targets | One engine across WebGPU, native mobile and several GPU APIs | Compilation and target packaging add complexity |
| vLLM | High-throughput OpenAI-compatible GPU APIs | Continuous batching, PagedAttention and broad serving features | Best value appears with suitable accelerators and concurrent traffic |
| SGLang | Agentic, structured and large-scale serving | Serving optimizations plus agent-oriented execution features | Fast-moving stack with a larger operational surface |
| TensorRT-LLM | NVIDIA-focused deployments needing deep optimization | NVIDIA-specific kernels, quantization and distributed runtime | Hardware specialization and a steeper build/tuning path |
| TGI | Maintaining an existing Hugging Face deployment | Mature integrations and observability | Upstream is in maintenance mode and recommends alternatives |
| TensorSharp | .NET teams evaluating a native GGUF and agent runtime | Managed .NET surface with compatible APIs | Smaller ecosystem; benchmark claims need local reproduction |
There is no universal winner. The safest shortlist comes from deployment shape, then a benchmark using your actual model and prompts.
First separate engine, model manager and serving platform
Comparison pages often put every project in one column as if each were a drop-in replacement. They are not.
An inference engine loads weights, allocates memory, evaluates tokens and applies kernels. A model manager adds downloading, naming, configuration and a friendly lifecycle. A serving platform adds queues, batching, routing, metrics, distributed execution and an API contract. Some projects span more than one layer, but their center of gravity still matters.
llama.cpp began as a compact C/C++ engine and now also ships `llama-server`, an OpenAI-compatible server with continuous batching, parallel decoding, embeddings, reranking, constrained JSON, tool use and monitoring endpoints. Ollama presents a higher-level experience for obtaining and running models with a CLI and REST API. vLLM and SGLang are designed around serving many requests efficiently. MLC LLM treats portability to browser and mobile runtimes as a first-class problem.
That distinction explains why “Ollama versus llama.cpp” is partly a user-experience decision, while “vLLM versus SGLang” is usually a scheduler, cache and production-serving decision.
llama.cpp: broad local hardware control
llama.cpp is the most flexible starting point when you have GGUF weights, constrained memory or non-datacenter hardware. Its official repository lists CPU execution plus CUDA, HIP, Metal, Vulkan, SYCL and other backends, with CPU/GPU hybrid inference when a model does not fit entirely on an accelerator. Integer quantization is central to the project, which makes it practical to trade model size and quality against available RAM or VRAM.
Use llama.cpp when you need to control context size, GPU layer offload, quantization, batch parameters and cache behavior directly. It is also a sensible portability baseline: the same GGUF model can often move among a workstation, Apple silicon, AMD hardware and an NVIDIA host, even though the best flags and performance will change.
The project is no longer only a command-line generator. `llama-server` exposes OpenAI-compatible chat, responses and embeddings routes and supports concurrent users. Still, assembling authentication, TLS, multi-instance routing, autoscaling and production monitoring remains the operator's responsibility. Its strength is transparent control, not a fully managed platform.
Choose it for a single-user assistant, offline batch job, embedded desktop feature or careful experimentation on heterogeneous hardware. For more local options, see the free local AI tools directory.
Ollama: the shortest path to a local developer API
Ollama is usually the easiest answer when the requirement is “install a model locally and call it from an application.” Its upstream project supports macOS, Windows and Linux, offers Docker deployment, and exposes a REST API for model execution and management. Its OpenAI-compatibility documentation covers familiar client shapes, reducing integration work for applications already written against that interface.
The main benefit is not a claim that Ollama always evaluates tokens faster than llama.cpp. It is that model acquisition, manifests, local storage, process lifecycle and API access arrive as one coherent product. That can save hours for prototypes, teaching, private chat and development tools.
Choose Ollama when operational simplicity matters more than tuning every backend parameter. Choose direct llama.cpp when you need a specific build, experimental flag, unusual offload layout or a minimal component inside another product. Because the projects and their embedded engine versions evolve, verify the exact release rather than assuming every llama.cpp feature is immediately exposed through Ollama.
MLX-LM: the focused Apple-silicon path
MLX-LM is a Python package for generation and fine-tuning on Apple silicon using MLX. Its official README documents Hugging Face integration, quantization, conversion, low-rank and full fine-tuning, distributed inference and both CLI and Python APIs. That makes it especially attractive to Mac developers who want a native research workflow rather than a cross-platform server first.
Use MLX-LM for interactive experiments, local evaluation, conversion and fine-tuning on Apple hardware. It can serve applications too, but its defining advantage is tight alignment with the Apple-silicon software stack. If you must deploy the same runtime to Linux, Windows, AMD and NVIDIA machines, llama.cpp or another portable serving layer is usually a more direct baseline.
The model format also matters. MLX-compatible weights and GGUF are different packaging ecosystems. Do not select a runtime before confirming that the exact model architecture, tokenizer, quantization and chat template you need are supported.
MLC LLM: browser, mobile and heterogeneous edge deployment
MLC LLM is the most distinctive option in this comparison because it targets WebGPU and WebAssembly in browsers as well as iOS, Android and native GPU backends. Its project describes a common `MLCEngine` with OpenAI-compatible APIs across REST, Python, JavaScript and mobile environments.
That makes MLC compelling when local inference is a product feature delivered to end-user devices, not merely a server on your LAN. A browser demo, privacy-sensitive mobile assistant or offline field application has constraints that vLLM was not designed to solve.
The trade-off is build and packaging complexity. Web and mobile targets require careful model compilation, download-size planning, device testing and memory controls. Treat “supports the platform” as the beginning of validation, not proof that the desired model will fit or perform well on every device.
vLLM: throughput-oriented GPU serving
vLLM is a strong default candidate for an OpenAI-compatible service with concurrent GPU traffic. Its official project highlights PagedAttention, continuous batching, chunked prefill, prefix caching, speculative decoding, distributed parallelism, streaming and several quantization methods. Its server supports chat, completions, responses, embeddings and other routes.
The important concept is utilization. A local interactive user often sends one request and waits. A service receives requests with different prompt and output lengths at overlapping times. Continuous batching and paged key/value-cache management aim to keep accelerators productive while controlling memory fragmentation.
Use vLLM when you have supported accelerator hardware, multiple users or jobs, and a need for a familiar API. Benchmark both throughput and latency: maximizing aggregate tokens per second can hurt tail latency, and a configuration that excels at short prompts may struggle with long-context prefill. Also harden the deployment at a reverse proxy; vLLM's documentation warns that its API-key option does not protect every endpoint.
If you are evaluating hosted capacity before buying hardware, the free GPU credits guide and free AI API guide can help identify zero-cost test paths, subject to current provider terms.
SGLang: agentic and structured serving at scale
SGLang describes itself as an inference framework optimized for agentic workloads, reinforcement-learning rollouts and large-scale serving. Its current hardware matrix spans NVIDIA and AMD GPUs, Google TPUs, Intel hardware, Apple silicon and additional accelerators. The surrounding ecosystem includes routing and cache integrations as well as model-serving components.
SGLang belongs on the shortlist when the workload repeatedly reuses prompt prefixes, produces structured output, calls tools or fans out agent requests. Those patterns can benefit from scheduling and cache techniques that go beyond a basic completion server. It is also relevant when a team wants one fast-moving framework for both research-scale experiments and production serving.
That breadth creates operational cost. Supported features vary by model and hardware backend, and a fast release cadence means upgrades need regression tests. Compare SGLang with vLLM using the same model revision, precision, request distribution and concurrency. A third-party benchmark with different flags is not a procurement decision.
TensorRT-LLM: optimize for NVIDIA when specialization is acceptable
TensorRT-LLM is NVIDIA's toolkit for optimizing and deploying LLM inference on NVIDIA GPUs. It combines a Python model-definition layer with specialized kernels, quantization paths, runtime components and distributed execution.
It deserves evaluation when the deployment is committed to NVIDIA hardware and the team can absorb engine-building and tuning complexity. That specialization can expose optimizations unavailable to a more portable stack. It is less attractive as a first prototype or when portability across vendors is a firm requirement.
Do not assume the most specialized runtime automatically wins. Model support, conversion effort, sequence lengths and traffic shape determine whether the additional engineering pays back. Measure end-to-end service latency and operational effort, not kernel throughput alone.
TGI: stable existing deployments, but not the default for a new build
Hugging Face Text Generation Inference helped establish production features such as continuous batching, tensor parallelism, streaming, quantization, metrics and tracing. Its repository now states that TGI is in maintenance mode and recommends engines including vLLM and SGLang for new deployments, plus llama.cpp or MLX for local use.
That does not mean a healthy TGI installation must be removed immediately. Existing integrations, known operating procedures and stable model coverage have value. It does mean a greenfield selection should account for reduced feature development and compare migration paths before increasing dependence on TGI-specific behavior.
TensorSharp: a .NET-native candidate with disclosure
The source item behind this guide was submitted by the TensorSharp maintainer, who disclosed that affiliation. TensorSharp's official repository presents a native .NET engine and agent runtime for GGUF models with CPU and accelerator backends, a browser UI, and Ollama/OpenAI-compatible APIs. It also publishes comparisons against llama.cpp.
Those comparisons are project-authored evidence, not independent proof of general superiority. TensorSharp may be worth a trial for teams that need a managed .NET integration or its agent runtime, but reproduce results using the same model file, context, prompt set, warm-up, thread count and hardware. A smaller ecosystem also means checking release cadence, issue response, model coverage and upgrade behavior more carefully.
The appropriate conclusion is inclusion, not endorsement: it solves a legitimate .NET integration problem, while performance and compatibility must be verified locally.
Formats and compatibility can decide before performance
Model names alone are insufficient. A runtime needs the exact architecture implementation, tokenizer, chat template, quantization and multimodal projector used by the artifact.
llama.cpp and TensorSharp emphasize GGUF. MLX-LM uses MLX-compatible weights. vLLM and SGLang commonly consume Hugging Face-style model repositories, while TensorRT-LLM may require an engine-building workflow. MLC compiles models for its target runtime. Conversion is possible in many cases, but it adds storage, time and a new artifact to validate.
Before benchmarking, confirm:
- The exact model revision is supported, not only the model family name.
- The intended quantization works on the target backend.
- The chat template produces the expected control tokens.
- Tool calls, JSON constraints, embeddings or multimodal inputs work if required.
- Context length and key/value-cache memory fit under realistic concurrency.
- The license permits the intended model and product use.
A workload-based selection path
One person, one laptop
Start with Ollama if you value convenience. Start with llama.cpp if you already have GGUF files or need to tune memory and offload. On Apple silicon, test MLX-LM alongside them when Python experimentation or fine-tuning matters.
A local API for a small team
Ollama or `llama-server` may be enough when concurrency is low. Add authentication and network boundaries yourself; “local” software exposed on a shared network is still a service. Record queue time, time to first token and cache behavior with several simultaneous users.
A public or internal GPU service
Benchmark vLLM and SGLang first. Include TensorRT-LLM when the fleet is NVIDIA-only and deeper optimization is justified. Test failure recovery, metrics, rolling upgrades, overload behavior and model reloads in addition to generation speed.
A browser or mobile feature
Evaluate MLC LLM. On Apple-only applications, MLX may also be relevant depending on the product layer. Device distribution, package size, thermal throttling and memory pressure matter more than a desktop benchmark.
An existing TGI estate
Keep it stable while building a compatibility test against vLLM or SGLang. The upstream maintenance notice is a planning signal, not an emergency migration order.
A .NET-native product
Compare TensorSharp with a separate llama.cpp server or bindings. Favor the option that passes your model-coverage, security, observability and maintenance tests; do not decide from the submitting maintainer's benchmark alone.
How to benchmark without fooling yourself
Use the model revision and quantization you will actually deploy. Freeze sampling settings and chat template. Separate cold model load, prompt prefill and output decode. Then test at concurrency 1 and at realistic concurrency rather than reporting one best-case tokens-per-second number.
Measure at least:
- model load time and peak RAM/VRAM;
- time to first token at short and long prompt lengths;
- output tokens per second per request;
- aggregate throughput at target concurrency;
- median, 95th and 99th percentile latency;
- cache-hit behavior for repeated prefixes;
- correctness of structured output and tool calls;
- stability under cancellation, overload and model reload;
- energy or cloud cost per completed useful task.
Warm-up runs should be separated from steady state. Repeat tests and publish the full command, hardware, software revision and prompt distribution. If one runtime requires a different quantization, say so—the result then compares complete deployment choices rather than engines alone.
Practical verdict
For most local users, the first comparison should be Ollama versus llama.cpp: convenience against direct control. Apple developers should add MLX-LM, and teams shipping to browsers or phones should add MLC LLM. For a shared accelerator service, start with vLLM and SGLang, then add TensorRT-LLM when an NVIDIA-specific path is acceptable.
TGI remains relevant for installed systems but its maintenance status changes the greenfield calculation. TensorSharp is a credible specialist candidate for .NET teams, with the important caveat that its source opportunity and benchmarks came from its maintainer.
The runtime is not the model, and the benchmark is not the workload. Shortlist by deployment layer, validate compatibility, then measure your own prompts at your own concurrency. That process is slower than choosing the top row in a comparison table—and far more likely to produce a reliable system.
For continuing evaluations, browse the FreeAI Tokens editorial hub and check current access conditions in the free AI tiers directory.
Sources and evidence
- llama.cpp repository and llama-server documentation: hardware backends, quantization and server features.
- Ollama repository and OpenAI compatibility documentation: installation, lifecycle and APIs.
- MLX-LM repository: Apple-silicon generation, conversion, quantization and fine-tuning.
- MLC LLM repository: WebGPU, mobile, native backends and common engine APIs.
- vLLM repository and OpenAI-compatible server documentation: scheduling, cache and serving capabilities.
- SGLang repository: workload focus, installation and hardware support.
- TensorRT-LLM repository: NVIDIA optimization and runtime architecture.
- Text Generation Inference repository: features and current maintenance notice.
- TensorSharp repository: .NET feature claims and project-authored benchmarks.
Sources were reviewed on 3 October 2026. Performance claims without independent reproduction are identified as project-authored; this article is a workload guide, not a benchmark ranking.