← All articles

How to run vLLM locally

vLLM is a high-throughput inference server for NVIDIA GPUs. Running it locally puts the model on your hardware, your latency, and your cost structure. VIMS integrates it as the default local inference engine.

vLLM is a high-throughput inference server for NVIDIA GPUs. Running it locally puts the model on your hardware, your latency, and your cost structure. VIMS integrates it as the default local inference engine.

What vLLM is

vLLM is an open-source inference engine built around PagedAttention, a technique for managing the key-value cache the way an operating system manages memory pages. The practical effect: it serves many concurrent requests from one model without running out of VRAM, with throughput well above naive serving.

It exposes an OpenAI-compatible API. Anything that speaks to the OpenAI chat completions endpoint can speak to vLLM.

Hardware requirements

The model you can serve is bounded by VRAM:

Modelfp16Quantized (AWQ/GPTQ 4-bit)
7B~16 GB~6 GB
13B~28 GB~10 GB
70B~140 GB (multi-GPU)~42 GB

A 7B or 13B quantized model runs on a single consumer GPU. A 70B model wants multiple datacenter cards or heavy quantization.

The manual path

Standalone:

pip install vllm
vllm serve meta-llama/Llama-3.1-8B-Instruct

The server listens on port 8000 with an OpenAI-compatible API:

curl http://localhost:8000/v1/chat/completions \
  -d '{"model":"meta-llama/Llama-3.1-8B-Instruct","messages":[{"role":"user","content":"hello"}]}'

That is the whole server. What it does not give you: model download management, quantization selection, multiple agents sharing one instance, or anything that decides which model answers which request. You wire that yourself.

The VIMS path

VIMS wraps vLLM in a provider manager that handles the lifecycle (Blog 01: Local-First AI):

  • Model download from Hugging Face, with quantization picks for smaller GPUs
  • Single-instance management, so you do not end up running three copies of a 13B on one card
  • Chat, tool calls, and embeddings routed through the same instance, so agents in different shells share one backend
  • Per-agent model choice: one instance on local vLLM, another on a cloud API, another on Ollama, all from one fleet view (Blog 02: Runtimes and Swarms)
  • Federated inference: peers in a team can query your local model over the P2P swarm

When the local model is not enough, agents pay for cloud access from their own wallets, not your credit card. The default stays local: no per-token cost, no audio or data leaving the machine, no dependency on an endpoint staying up.

When local is the wrong call

Local serving wins on cost at volume, privacy, and latency stability. It loses on frontier capability: the strongest models still live behind APIs. The practical split most teams land on: local for the high-volume work (classification, summarization, tool routing, voice loops), cloud API for the hard reasoning, and the payment gate in between (Blog 09: Agent Wallets).

To run it inside VIMS with agent shells and swarms on top, see Desktop and Compute.