← All articles

Local-First AI: vLLM and Ollama

Part 2 of the [VIMS Blog Network](00-vims-convergence-point). Agents should think on your hardware, not someone else's.

Part 2 of the VIMS Blog Network. Agents should think on your hardware, not someone else's.

The Cloud Trap

Most agent platforms have a hidden dependency: they only work when someone else's GPU is available. You send a prompt to an API endpoint, wait for a response, and pay per token. If the endpoint is down, your agent is brain-dead. If the pricing changes, your agent costs more. If the provider decides your use case violates terms of service, your agent stops existing.

That is a leash.

VIMS takes a different position: local-first. When you ask an agent to do something, the inference should happen on your hardware when possible. Your GPU, your models, your latency, your cost structure. When local models are not sufficient, agents use their own wallets to pay for cloud API access, not on your credit card. (Blog 09: Agent Wallets)

vLLM: NVIDIA-Native Serving

VIMS integrates vLLM as its primary local inference engine for NVIDIA GPUs. vLLM is a high-throughput, memory-efficient inference server that uses PagedAttention to manage the KV cache, so it can serve multiple concurrent requests without running out of VRAM.

The integration handles the full lifecycle: model download from Hugging Face, quantization support for fitting larger models into smaller GPUs, and single-instance management so you are not accidentally running three copies of a 13B model on one card. The VIMS provider manager routes chat requests, tool calls, and embeddings through the same vLLM instance, so agents in different shells share one inference backend.

When an agent needs to reason through a complex task (breaking down a user request, planning a multi-step workflow, generating code, summarizing a document), that reasoning happens on your GPU. No round-trip to a cloud API. No per-token cost. No dependency on an external service being up.

Ollama: Bundled Model Management

Not every machine has an NVIDIA GPU. VIMS bundles an Ollama supervisor that manages a curated catalog of models, from small quantized models that run on CPU to larger models that leverage whatever acceleration is available. Ollama provides an HTTP API for model listing, pulling, and inference, and VIMS wraps it with the same provider manager interface.

The result is that VIMS works on a wide range of hardware: a workstation with multiple NVIDIA GPUs running vLLM, a laptop with an AMD card running Ollama, or a server with only CPU running quantized models. The agent does not care which backend is serving; it sends a chat request and gets a response.

Agent Shells and Parallel Runtimes

An agent in VIMS is a process: a shell instance with its own filesystem workspace, security policy, tool access, and lifecycle. The vLLM provider enables parallel agent instances, multiple shells running simultaneously, each consuming from the same inference backend. This is the foundation for swarms: coordinated groups of agent instances that work together on a problem.

VIMS ships nine agent shells (Rust, Python, TypeScript, and Go runtimes with different strengths) unified under one fleet manager. Swarms compose agents across these runtimes, with eight topologies for different coordination patterns. Each instance can use a different inference provider, local or cloud, giving you fine-grained control over the cost and capability of every agent. The full runtime catalog, cross-runtime coordination, per-instance model control, and swarm architecture are covered in detail in the next article. (Blog 02: Runtimes and Swarms)

When an agent joins a meeting or a team room, it does so as its shell instance, carrying its identity, its tools, and its security policy into the collaboration. (Blog 08: Agent Identity, Blog 11: Flow and Teams)

Federated Inference

Local-first does not mean isolated. VIMS supports federated inference across P2P team rooms: a peer in a team can query another peer's local model, or broadcast a query to all online peers. Your local vLLM instance becomes a resource that other team members can access, and their models become resources for you.

This is peer-to-peer model serving. Your model runs on your hardware, but its output is available to your collaborators. The team surface handles the routing: directed queries go to a specific peer, and broadcasts go to all online peers and aggregate responses.

Federated inference is the bridge between local-first and collaborative AI. You get the privacy and cost benefits of local inference, plus the diversity and capacity of multiple models across the team. (Blog 06: P2P Surfaces, Blog 11: Flow and Teams)

Voice: The Local Loop

When an agent joins a meeting as a synthetic participant, the entire voice loop runs locally: whisper.cpp handles speech-to-text with the large-v3 model, the configured LLM provider handles reasoning, and Piper handles text-to-speech. The audio is captured from the meeting's WebRTC stream, transcribed, reasoned about, and spoken back, all on your hardware, with no cloud STT or TTS API in the loop.

This means a meeting agent can participate in a voice conversation with sub-second latency, fully on-device, while the meeting itself is connected peer-to-peer over Hyperswarm. The only thing that leaves your machine is the synthesized speech, streamed over WebRTC to other participants. (Blog 07: Voice and Meeting Agents)

The Local-First Contract

Local-first is a contract:

  1. Your hardware, your models. Inference runs on your GPU when possible. No mandatory cloud dependency.
  2. Your data stays local. Documents, databases, knowledge bases, and conversation logs live on your machine. P2P replication shares them with peers you choose, not with a central server.
  3. Agents are processes. Shell instances with filesystems, security policies, and lifecycles. They persist between interactions.
  4. Collaboration extends local. Federated inference and P2P surfaces let you share your local capabilities with collaborators without surrendering control.

This contract is what makes VIMS agents genuinely autonomous: they think on your hardware, act within your security policies, and collaborate on your terms, with no dependency on a cloud endpoint and no data going to a server you do not control.


Previous: VIMS: The Convergence Point Next: Runtimes and Swarms: The Multi-Runtime Architecture