From llama.cpp to vLLM: Inside the Architecture of a Production LLM MaaS Platform
When we first start working with open-source LLMs, the problem looks deceptively simple:
How do I run this model?
Download a model, load it into an inference runtime, send a prompt, and get a response.
Projects such as llama.cpp have made this experience remarkably accessible. Its goal is to provide efficient LLM and VLM inference with minimal setup across a wide range of hardware, including CPUs, Apple Silicon, NVIDIA and AMD GPUs, and other accelerators. It also supports low-bit quantization and CPU+GPU hybrid inference. (GitHub)
But the problem changes dramatically when we move from running an LLM for ourselves to providing LLM inference as a service.
Now we need to answer questions such as:
- How many requests can one GPU serve?
- How do we efficiently handle hundreds or thousands of concurrent requests?
- How do we maximize GPU utilization?
- How do we manage the KV cache?
- How do we serve models across multiple GPUs?
- How do we route requests between model replicas?
- How do we scale capacity when traffic changes?
- How do we implement authentication, quotas and billing?
- How do we isolate different customers?
- How do we operate thousands of GPUs reliably?
This is where the architecture evolves from llama.cpp → vLLM → LLM Model-as-a-Service (MaaS).
1. First Principle: An LLM Is a Compute Workload
Before discussing any particular inference engine, it is useful to step back.
An LLM is essentially a very large computation graph backed by a large set of model weights.
At the simplest level:
Application
│
│ Prompt
▼
┌──────────────┐
│ LLM Runtime │
└──────┬───────┘
│
▼
Model Weights
│
▼
CPU / GPU
The first engineering challenge is therefore:
How can we execute the model efficiently on available hardware?
This is the problem space where llama.cpp is particularly interesting.
2. llama.cpp: Making LLM Inference Portable
llama.cpp is implemented in C/C++ and is designed to run LLMs and VLMs with minimal setup across a wide variety of hardware.
Its supported environments include:
- x86 CPUs
- ARM CPUs
- Apple Silicon
- NVIDIA GPUs
- AMD GPUs
- Vulkan
- SYCL
- other accelerator backends
It also supports CPU-specific instruction sets such as AVX, AVX2, AVX512 and AMX, while Apple Silicon receives dedicated optimization through ARM NEON, Accelerate and Metal. (GitHub)
Conceptually:
llama.cpp
│
┌───────────────┼───────────────┐
│ │ │
▼ ▼ ▼
CPU Apple Silicon GPU
│ │ │
x86/ARM Metal CUDA/HIP
This portability is one of its greatest strengths.
3. Quantization Changes the Economics of Local Inference
Another important part of the llama.cpp ecosystem is aggressive quantization.
Instead of representing every model weight using FP16 or BF16, a model can be represented at lower precision:
FP16
│
├── Q8
├── Q6
├── Q5
├── Q4
├── Q3
└── Q2
The basic trade-off is:
Lower precision
│
├── Smaller model
├── Lower memory requirement
└── Potentially faster inference
│
└── Some loss of numerical precision
This can make a huge difference on consumer hardware.
Imagine:
Large Model
│
┌───────┴───────┐
│ │
FP16 Q4
│ │
Very large Much smaller
│ │
▼ ▼
Server-class Workstation
memory hardware
llama.cpp supports integer quantization from very low bit widths through 8-bit, making it particularly attractive for local and resource-constrained inference. (GitHub)
4. CPU + GPU Hybrid Inference
One particularly useful feature is hybrid inference.
Suppose we have:
GPU VRAM = 24 GB
System RAM = 64 GB
but our model requires more than 24 GB.
Instead of simply failing:
Model
│
▼
24 GB VRAM
│
X
llama.cpp can divide the workload between GPU and system memory:
Model
│
┌───────┴───────┐
│ │
▼ ▼
GPU VRAM System RAM
│ │
└───────┬───────┘
▼
llama.cpp
This makes it possible to experiment with models that would otherwise exceed the available VRAM. (GitHub)
For a local AI enthusiast, developer, researcher or edge deployment, this is extremely powerful.
5. llama.cpp Can Also Run an API Server
There is another misconception worth clearing up.
llama.cpp isn’t merely a command-line program.
It provides a server interface that can expose an OpenAI-compatible API.
For example:
Application
│
│ HTTP API
▼
┌──────────────┐
│ llama-server │
└──────┬───────┘
│
GGUF Model
│
CPU / GPU
So you can build applications or agents on top of llama.cpp without embedding the inference runtime directly into your application. (GitHub)
This raises an important question:
If llama.cpp can serve multiple requests through an API, why do we need something like vLLM?
The answer is scale and workload characteristics.
6. The Problem Changes at High Concurrency
Consider a local application:
User
│
▼
llama.cpp
│
▼
GPU
Now compare that with an LLM API service:
User 1 ────┐
User 2 ────┤
User 3 ────┤
User 4 ────┤
User 5 ────┤
... │
User 999 ──┤
User 1000 ─┘
│
▼
LLM Server
│
▼
GPU Pool
Each request can have a different:
- prompt length
- output length
- arrival time
- context length
- generation speed
- priority
The GPU therefore needs to be treated as a shared, dynamically scheduled compute resource.
This is where modern LLM serving systems become interesting.
7. Continuous Batching
Traditional batch processing might look like:
Batch 1
──────────────────────────────
Request A █████████████████
Request B ███████
Request C █████████████
↓
Wait for all
↓
Batch 2
──────────────────────────────
Request D █████
Request E ███████████
The problem is obvious.
Request B may finish long before Request A, but the traditional batch may still be waiting for the slowest request.
LLM generation is especially well suited to a different approach:
Continuous batching.
Conceptually:
Time ─────────────────────────────────────────>
A ███████████████████
B ███████
C █████████████
D █████
E ███████████
F ██████
↑ ↑
requests enter/leave
the active batch
The scheduler continuously adjusts the active set of sequences.
Instead of thinking:
batch → execute → finish → next batch
we can think:
arrive → schedule → generate
↑ ↓
dynamically
update
This keeps the accelerator busy with useful work.
vLLM lists continuous batching as one of its core serving capabilities. (vLLM)
8. The Real Memory Problem: KV Cache
There is another major obstacle to high-concurrency LLM serving.
It is not just the model weights.
It is the KV cache.
During autoregressive generation, the model maintains key/value states associated with previously processed tokens.
For one request:
Prompt
│
▼
Token 1
│
▼
Token 2
│
▼
Token 3
│
▼
...
The corresponding KV data needs to remain available during generation.
With one request:
GPU Memory
┌───────────────────────────────┐
│ Model Weights │
├───────────────────────────────┤
│ Request A KV Cache │
└───────────────────────────────┘
With hundreds of requests:
GPU Memory
┌───────────────────────────────┐
│ Model Weights │
├───────────────────────────────┤
│ Request A KV │
│ Request B KV │
│ Request C KV │
│ Request D KV │
│ ... │
│ Request N KV │
└───────────────────────────────┘
The KV cache can become a significant memory consumer.
And because requests have different sequence lengths and lifetimes, managing that memory efficiently becomes a central serving problem.
9. PagedAttention
This is where vLLM’s original research contribution becomes particularly interesting.
The PagedAttention paper describes a KV-cache management approach inspired by virtual memory and paging in operating systems. Instead of requiring each request’s KV cache to occupy one large contiguous memory region, KV data can be managed using smaller blocks. (DOI)
Conceptually:
GPU Memory
┌──────┬──────┬──────┬──────┬──────┬──────┐
│ A1 │ C1 │ B1 │ A2 │ D1 │ C2 │
├──────┼──────┼──────┼──────┼──────┼──────┤
│ B2 │ A3 │ E1 │ C3 │ D2 │ ... │
└──────┴──────┴──────┴──────┴──────┴──────┘
A logical request can then reference its blocks:
Request A
│
├── A1
├── A2
└── A3
Request C
│
├── C1
├── C2
└── C3
This allows the inference engine to use GPU memory more flexibly and enables sharing of KV-cache blocks in appropriate scenarios.
The original vLLM research reported substantial throughput improvements over contemporary serving systems, particularly as sequences and models became larger. (DOI)
10. vLLM: From Runtime to Serving Engine
At this point we can see the conceptual difference.
llama.cpp primarily asks:
How do I execute an LLM efficiently on this machine?
vLLM asks:
How do I efficiently execute many LLM requests on this GPU?
A simplified vLLM worker looks like:
vLLM Worker
│
┌──────────┴──────────┐
│ │
▼ ▼
Request Scheduler Model Executor
│ │
│ ┌──────┴──────┐
│ │ │
│ Model Weights KV Cache
│
└──────────────┬──────────────
▼
GPU
The actual vLLM architecture is significantly more sophisticated, but this mental model captures the important idea:
The inference engine is actively managing requests, memory and GPU execution as one serving system.
vLLM currently combines continuous batching with optimized attention memory management, optimized kernels, quantization, prefix caching and multiple forms of distributed parallelism. (vLLM)
11. One GPU Quickly Becomes Many GPUs
Large models often cannot fit onto a single GPU.
Suppose we have:
Large Model
│
┌──────────┼──────────┐
▼ ▼ ▼
GPU 0 GPU 1 GPU 2
Now the inference engine must coordinate computation across GPUs.
vLLM supports multiple distributed inference strategies, including tensor, pipeline, data and expert parallelism. (vLLM)
A simplified example:
vLLM
│
┌────────┼────────┐
▼ ▼ ▼
GPU 0 GPU 1 GPU 2
│ │ │
└────────┼────────┘
│
GPU Interconnect
The challenge is no longer just model execution.
It becomes:
How do we efficiently coordinate a distributed GPU computation?
12. From One vLLM Process to a GPU Fleet
Now imagine that we’re not running one model.
We’re operating an LLM service.
We might have:
GPU Fleet
┌────────┬────────┬────────┬────────┐
│ │ │ │ │
Node 1 Node 2 Node 3 Node 4 ...
│ │ │ │
GPU GPU GPU GPU
And perhaps:
Qwen Cluster
└── vLLM replicas
Llama Cluster
└── vLLM replicas
DeepSeek Cluster
└── vLLM replicas
At this point, vLLM alone is no longer enough.
We need orchestration.
13. Kubernetes: Managing the Infrastructure
This is where Kubernetes commonly enters the architecture.
It is important to understand the division of responsibilities:
Kubernetes manages infrastructure and workloads.
vLLM manages LLM inference.
For example:
Kubernetes Cluster
│
┌──────────────┼──────────────┐
▼ ▼ ▼
GPU Node GPU Node GPU Node
│ │ │
vLLM Pod vLLM Pod vLLM Pod
│ │ │
Model A Model A Model B
Kubernetes can handle things such as:
- workload scheduling
- service discovery
- health management
- deployment
- rolling upgrades
- failure recovery
- autoscaling
But LLM workloads make GPU scheduling considerably more complicated than ordinary CPU workloads.
14. GPU Scheduling Is Not Just CPU Scheduling
Imagine two workloads.
Workload A
Model A
1 GPU
Workload B
Model B
8 GPUs
Tensor Parallel
The scheduler cannot treat them as equivalent workloads.
For Workload A:
Node 1
┌─────┐
│ GPU │ ← enough
└─────┘
For Workload B:
Node
┌─────┬─────┬─────┬─────┐
│ GPU │ GPU │ GPU │ GPU │
├─────┼─────┼─────┼─────┤
│ GPU │ GPU │ GPU │ GPU │
└─────┴─────┴─────┴─────┘
The placement of those GPUs can matter because of:
- GPU memory
- GPU type
- NVLink topology
- PCIe topology
- NUMA topology
- inter-GPU bandwidth
- model parallelism requirements
Therefore, at large scale, the platform needs to become increasingly GPU- and workload-aware.
15. The MaaS Layer
Now we finally reach the actual Model-as-a-Service platform.
A customer doesn’t want to know:
Which Kubernetes node?
Which GPU?
Which vLLM pod?
Which model replica?
They simply want:
POST /v1/chat/completions
model = qwen3-32b
The platform hides the infrastructure.
Conceptually:
Customer
│
▼
API Gateway
│
▼
Model Router
│
▼
Inference
│
▼
GPU
This abstraction is the essence of MaaS.
16. API Gateway
The first platform layer usually deals with the customer-facing API.
For example:
Customer
│
▼
┌──────────────────────┐
│ API Gateway │
├──────────────────────┤
│ Authentication │
│ API Keys │
│ Rate Limiting │
│ Quotas │
│ Request Validation │
│ TLS │
└──────────┬───────────┘
│
▼
The gateway answers:
Who is this customer?
Are they allowed to use this model?
Are they within quota?
Are they exceeding their rate limit?
This is already outside the scope of an inference engine.
17. Model Routing
The next layer is the model router.
Suppose the customer asks for:
model = qwen3-32b
There might be dozens or hundreds of replicas behind that logical model name.
Model Router
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Qwen Pod Qwen Pod Qwen Pod
│ │ │
vLLM vLLM vLLM
│ │ │
GPU GPU GPU
The router may consider:
Requested model
│
├── Model availability
├── Replica health
├── Queue depth
├── GPU utilization
├── Latency
├── Region
├── Tenant priority
└── Capacity
This is an important architectural boundary.
The customer sees a logical model.
The platform decides which physical inference worker should execute the request.
18. Control Plane vs Data Plane
At this point, a production MaaS architecture naturally separates into two major domains.
Control Plane
The control plane decides:
What should be running, where should it run, and how should it be configured?
Control Plane
│
┌───────────────┼────────────────┐
▼ ▼ ▼
Model Registry Deployment Scheduler
│ Manager │
▼ ▼ ▼
Model Versions Autoscaling GPU Placement
│ │ │
└───────────────┼────────────────┘
│
▼
GPU Fleet
Data Plane
The data plane handles the actual customer inference traffic.
Data Plane
Customer
│
▼
API Gateway
│
▼
Model Router
│
▼
Inference Worker
│
▼
vLLM / SGLang / TRT-LLM
│
▼
GPU
│
▼
Generated Tokens
This separation is fundamental to scalable distributed systems.
The control plane manages the state of the platform.
The data plane handles the traffic flowing through the platform.
19. A Complete LLM MaaS Architecture
We can now put all the pieces together.
Internet
│
▼
┌────────────────────┐
│ API Gateway │
│ │
│ Auth / TLS │
│ Rate Limit │
│ Quota │
└─────────┬──────────┘
│
▼
┌────────────────────┐
│ Model Router │
│ │
│ Model Selection │
│ Load Balancing │
│ Routing Policy │
└─────────┬──────────┘
│
┌────────────────┼────────────────┐
│ │ │
▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────┐
│ Qwen │ │ Llama │ │ DeepSeek │
│ Pool │ │ Pool │ │ Pool │
└────┬─────┘ └────┬─────┘ └────┬─────┘
│ │ │
vLLM × N vLLM × N vLLM × N
│ │ │
└────────────────┼────────────────┘
│
┌──────┴──────┐
│ GPU Fleet │
└──────┬──────┘
│
Kubernetes
│
▼
GPU Infrastructure
And alongside it:
┌────────────────────────┐
│ Control Plane │
├────────────────────────┤
│ Model Registry │
│ Deployment Manager │
│ Scheduler │
│ Autoscaling │
│ Capacity Management │
│ Health Management │
└───────────┬────────────┘
│
▼
GPU Fleet
This is much closer to what we should think of as a production LLM MaaS architecture.
20. Billing and Usage Accounting
There is another layer that is easy to overlook.
MaaS isn’t just an inference system.
It is also a commercial service.
The platform needs to know:
Customer
│
▼
API Key
│
▼
Request
│
├── Model
├── Input tokens
├── Output tokens
├── Latency
├── Region
└── Status
│
▼
Usage Metering
│
▼
Billing
For example:
Customer A
├── 10M input tokens
├── 2M output tokens
└── Qwen3-32B
Customer B
├── 50M input tokens
├── 8M output tokens
└── Llama model
The inference layer produces the raw execution metrics.
The MaaS platform turns those metrics into:
- usage
- quotas
- invoices
- dashboards
- cost attribution
Again, this functionality sits above the inference engine.
21. Why a MaaS Provider May Use More Than vLLM
It is tempting to say:
“Production LLM MaaS = Kubernetes + vLLM.”
That’s a useful starting point, but it is not quite correct.
A real platform might have:
Model Router
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Qwen Llama DeepSeek
│ │ │
vLLM SGLang TensorRT-LLM
│ │ │
GPU Pool GPU Pool GPU Pool
Different inference engines can have different advantages depending on:
- model architecture
- accelerator
- workload
- quantization
- latency requirements
- throughput requirements
- parallelism strategy
The important architectural principle is:
The MaaS platform should abstract the inference engine from the customer.
The customer asks for:
model = qwen3-32b
They shouldn’t need to care whether the backend happens to be:
vLLM
SGLang
TensorRT-LLM
22. Agentic AI Changes the Workload
This architecture becomes even more interesting when we move from traditional chat applications to AI agents.
A traditional application might look like:
User
│
▼
LLM
│
▼
Response
An agentic application looks more like:
Agent
│
┌─────────────┼─────────────┐
▼ ▼ ▼
LLM Call Tool Call LLM Call
│
▼
RAG
│
▼
LLM Call
│
▼
Tool Call
│
▼
LLM Call
│
▼
Response
One user request can generate many LLM requests.
Now imagine thousands or millions of agents operating concurrently.
The workload becomes:
Agent 1 ─┐
Agent 2 ─┤
Agent 3 ─┤
Agent 4 ─┤
Agent 5 ─┤
... │
Agent N ─┘
│
▼
Model Router
│
▼
Inference Fleet
│
▼
GPUs
This makes efficient scheduling increasingly important.
The optimization target is no longer simply:
How many tokens per second can one request generate?
It becomes:
How efficiently can the platform schedule millions of dynamically arriving inference workloads across a finite GPU fleet?
23. The Architecture Is Really About Resource Management
This leads to a broader way of thinking about LLM infrastructure.
At the bottom:
GPU
│
▼
Model Execution
Then:
vLLM
│
▼
Request Scheduling
KV Cache Management
Batch Management
Then:
MaaS Platform
│
▼
Model Routing
Tenant Management
Usage / Billing
Capacity Management
Then:
Infrastructure
│
▼
Kubernetes / Scheduler
│
▼
GPU Fleet
The problem therefore evolves from:
Compute the next token
to:
Schedule the next token
and ultimately to:
Schedule the right computation on the right GPU for the right customer at the right time.
That is a much more interesting distributed-systems problem.
24. Three Levels of LLM Infrastructure
The entire evolution can be summarized in three stages.
Level 1 — Run the Model
Model
│
▼
llama.cpp
│
▼
CPU / GPU
The question:
Can I run this model efficiently?
Typical scenarios:
- local AI
- development
- experimentation
- personal assistants
- edge inference
- CPU inference
- Apple Silicon
- consumer GPUs
Level 2 — Serve the Model
Multiple Users
│
▼
vLLM
│
▼
GPU(s)
The question:
Can I efficiently serve concurrent requests?
Typical concerns:
- continuous batching
- KV-cache efficiency
- GPU utilization
- latency
- throughput
- multi-GPU inference
Level 3 — Operate an LLM Platform
Customers
│
▼
API Gateway
│
▼
Model Router
│
┌──────────┼──────────┐
▼ ▼ ▼
vLLM SGLang TRT-LLM
│ │ │
└──────────┼──────────┘
▼
GPU Fleet
│
Kubernetes
│
▼
Infrastructure
The question becomes:
Can I turn a fleet of expensive accelerators into a reliable, scalable, multi-tenant inference service?
25. The Mental Model
A useful way to visualize the entire stack is:
┌─────────────────────────────────────────────────────┐
│ AI Applications │
│ Chat / RAG / Coding / Agents / VLM │
├─────────────────────────────────────────────────────┤
│ MaaS Platform │
│ API / Auth / Quota / Billing / Routing / Metering │
├─────────────────────────────────────────────────────┤
│ Inference Serving │
│ vLLM / SGLang / TensorRT-LLM / ... │
├─────────────────────────────────────────────────────┤
│ GPU Orchestration │
│ Kubernetes / Scheduler / Autoscaling │
├─────────────────────────────────────────────────────┤
│ GPU Infrastructure │
│ NVIDIA / AMD / Other Accelerators │
└─────────────────────────────────────────────────────┘
At the other end of the spectrum:
┌─────────────────────────────────────────────────────┐
│ Local AI │
├─────────────────────────────────────────────────────┤
│ llama.cpp │
├─────────────────────────────────────────────────────┤
│ GGUF / Quantization │
├─────────────────────────────────────────────────────┤
│ CPU / GPU / Apple Silicon │
└─────────────────────────────────────────────────────┘
These aren’t mutually exclusive technologies.
They represent different points in the inference stack and different optimization goals.
Conclusion: From Running Models to Operating Intelligence
The evolution from llama.cpp → vLLM → LLM MaaS is ultimately an evolution in the problem we are trying to solve.
llama.cpp focuses on making powerful models executable across a remarkably broad range of hardware.
vLLM focuses on turning GPU compute into highly efficient, concurrent LLM serving.
An LLM MaaS platform goes one level further: it turns a fleet of accelerators and inference engines into a reliable, scalable, multi-tenant service.
The architecture eventually looks something like:
LLM MaaS
│
┌─────────────────┼─────────────────┐
│ │ │
Product Plane Control Plane Data Plane
│ │ │
▼ ▼ ▼
API / Billing Model Registry Model Router
Auth / Quota Deployment Inference
Metering Scheduling vLLM
Observability Autoscaling SGLang
Capacity TRT-LLM
│
▼
GPU Fleet
│
Kubernetes
And perhaps the most important lesson is this:
The hard part of LLM infrastructure is gradually moving from “how do we execute the model?” to “how do we efficiently orchestrate computation, memory, networking and accelerators around the model?”
For traditional chat applications, this already matters.
For agentic AI, it becomes even more important because a single user interaction can generate a cascade of inference, tool-calling and retrieval operations.
The future of LLM infrastructure therefore isn’t just about building faster models.
It is about building better systems for turning GPU compute into intelligence on demand.
References
-
llama.cpp — GGML/llama.cpp project and documentation. GitHub — llama.cpp
-
vLLM Documentation — serving architecture, continuous batching, PagedAttention, quantization and distributed inference. vLLM Documentation
-
Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention”, SOSP 2023. ACM Digital Library — PagedAttention Paper
-
PagedAttention preprint — arXiv:2309.06180. (arXiv)
-
vLLM Project — official project site. vLLM :::
One editorial change I’d strongly recommend for the Medium version: make the six architecture diagrams actual visual figures rather than code-block diagrams. In particular, the PagedAttention, continuous batching, model router, and control-plane/data-plane diagrams deserve proper visual treatment; those are the parts readers are most likely to save/share.