Voice and Meeting Agents: Whisper to Room
Part 8 of the [VIMS Blog Network](00-vims-convergence-point). VIMS agents listen and speak.
Part 8 of the VIMS Blog Network. VIMS agents listen and speak.
The Voice Gap
Text-based agents are useful, but they exclude an entire modality of human-computer interaction: voice. In meetings, on calls, in conversation, the natural interface is speech. An agent that can only text is an agent that can only participate in async workflows. An agent that can listen, reason, and speak is an agent that can participate in real-time collaboration.
Building a voice agent is hard. You need speech-to-text (STT) that runs fast enough for real-time conversation, a language model that can reason over transcribed speech, and text-to-speech (TTS) that sounds natural enough to not break the flow of conversation. You need to handle interruptions, overlapping speech, and the messy reality of human conversation. And if you want it to work locally (without sending audio to a cloud API), you need all three components to run on your hardware.
VIMS solves this with a fully local voice pipeline and meeting agents that join rooms as synthetic participants.
The Voice Pipeline
The VIMS voice loop has three stages, all running locally:
Stage 1: Speech-to-Text via whisper.cpp
whisper.cpp is the C++ port of OpenAI's Whisper model, optimized for local CPU/GPU inference. VIMS uses the large-v3 model (the most accurate Whisper variant) running locally on your hardware.
When an agent is in a meeting, it subscribes to the audio stream from the WebRTC data channel. Incoming audio is buffered, fed to whisper.cpp, and transcribed in near-real-time. The transcription includes timestamps, so the agent knows when each phrase was spoken and by whom (when the meeting multiplexer provides speaker diarization).
There is no per-minute cost, no audio leaving your machine, no latency from network round-trips. The transcription happens on your CPU or GPU, with latency determined by your hardware, typically a few hundred milliseconds for the large-v3 model on a modern machine. (Blog 01: Local-First AI)
Stage 2: Reasoning via the Configured LLM
The transcribed text becomes the agent's input. The agent reasons over it using whatever LLM provider is configured: vLLM serving a local model, Ollama with a quantized model, or a cloud API if the agent's wallet pays for it.
The reasoning step is where the agent decides what to say. It considers the conversation context, the meeting topic, any instructions it was given (e.g., "take notes," "summarize action items," "answer questions about the product roadmap"), and produces a response.
This is the same reasoning engine that powers text-based agent interactions. The voice pipeline uses the same LLM the agent always uses, so the agent's voice responses are as capable as its text responses.
Stage 3: Text-to-Speech via Piper
Piper is a fast, local neural TTS system. VIMS uses it to convert the agent's text response into speech, which is then streamed back over the WebRTC data channel to other meeting participants.
The TTS voice is configurable. VIMS uses the en_US-ryan-high voice by default: a natural-sounding English voice. The audio is streamed, not batched; the agent starts speaking as soon as the first sentence is synthesized, not after the entire response is generated. This reduces perceived latency and makes the conversation feel more natural.
Meeting Agent Peers
An agent in a VIMS meeting is a synthetic peer: a full participant in the meeting with its own audio stream.
When an agent joins a meeting, it:
- Connects via WebRTC: the same protocol human participants use. The agent's audio is a WebRTC media stream, just like any other participant.
- Subscribes to other participants' audio: the agent receives the mixed audio stream from the meeting, transcribes it, and reasons over it.
- Responds via TTS: when the agent has something to say, it synthesizes speech and streams it back. Other participants hear it as a voice in the meeting, not a text message in a sidebar.
- Renders an avatar: the agent's avatar is displayed in the meeting tile. If the agent has an on-chain NFT, the avatar is rendered from the NFT's visual data. (Blog 08: Agent Identity)
The meeting agent lifecycle is managed by the meeting room, not by the agent instance. The room tracks agent peers: their IDs, display names, instance IDs, and join timestamps. Agents can be added and removed dynamically during a meeting.
The Schedule → Meeting Auto-Attach
VIMS integrates voice agents with the scheduling system. When a meeting is booked with an agent attached, the agent automatically joins the meeting room at the scheduled start time and leaves shortly after the end time.
This means you can schedule a meeting with a synthetic agent the same way you schedule a meeting with a human: pick a time, add the agent as an attendee, and the agent shows up. The agent's calendar entry includes the meeting room ID, so it knows where to connect.
The auto-attach flow is:
- Booking created: a schedule booking includes an agent ID and a meeting room ID.
- Start time reached: the agent joins the meeting room as a synthetic peer.
- Meeting proceeds: the agent participates in the conversation.
- End time reached: the agent leaves the room.
The agent joins the call, listens, speaks, and participates. You can schedule a weekly standup with an agent that takes notes, a customer call with an agent that answers technical questions, or a brainstorming session with an agent that contributes ideas. (Blog 11: Flow and Teams)
Inject Audio: Headless Verification
How do you test a voice agent without a microphone? VIMS provides an inject-audio endpoint that feeds a synthetic audio payload directly into the agent's voice loop. You provide a WAV file (mono, any sample rate); the agent processes it through the full STT → LLM → TTS pipeline and returns the synthesized speech.
This is used for:
- Automated testing: verify that an agent responds correctly to specific audio inputs without setting up a WebRTC session.
- Pre-rendered prompts: make an agent say a specific thing in a meeting without waiting for the conversation to reach a point where it would naturally speak.
- Accessibility: users who cannot speak can type a message, have it synthesized to audio, and injected into the meeting on their behalf.
The Full Local Loop
The critical architectural point: the entire voice loop runs locally. whisper.cpp runs on your CPU/GPU. The LLM runs on your vLLM/Ollama instance. Piper runs on your CPU. The only thing that leaves your machine is the synthesized speech, streamed over WebRTC to other participants.
This means:
- No cloud STT API: no per-minute cost, no audio sent to a third party
- No cloud LLM dependency: the same local model that handles text chat handles voice reasoning
- No cloud TTS API: Piper synthesizes speech locally
- Sub-second latency: the loop is local, so latency is determined by your hardware, not by network round-trips
- Privacy: the audio never leaves your machine until it is synthesized speech ready for the meeting
A meeting agent can participate in a voice conversation with sub-second latency, fully on-device, while the meeting itself is connected peer-to-peer over Hyperswarm. (Blog 06: P2P Surfaces)
Security and Audit
Every voice interaction is recorded in the security audit log. The transcription, the agent's reasoning, the synthesized response are all logged with correlation IDs linking to the meeting, the agent instance, and the timestamp. If an agent says something unexpected in a meeting, you can trace exactly what it heard, what it reasoned, and what it said. (Blog 03: Security First)
The SDK Surface
Meeting agents are not limited to the VIMS desktop app. The SDK exposes MeetingAddAgent via ExtendedHooks, so any SDK app can add an agent to a meeting room. An app that manages customer support calls can add a VIMS agent to transcribe, summarize, and suggest responses, all within the app's own UI, powered by the VIMS voice pipeline. (Blog 12: SDK, Coder, and Terminal)
Previous: P2P Surfaces Next: Agent Identity: Nostr Keys and On-Chain NFTs
blog