← All articles

Data and Knowledge: Vector Search and SQL

Part 5 of the [VIMS Blog Network](00-vims-convergence-point). Agents need knowledge: structured data, searchable documents, and encoded media, beyond prompts.

Part 5 of the VIMS Blog Network. Agents need knowledge: structured data, searchable documents, and encoded media, beyond prompts.

The Fragmentation Problem

Ask an agent to "find the customer who churned last quarter and summarize their support tickets." What does it need?

It needs a SQL database (customer records, churn status, ticket history). It needs a document store (email threads, meeting notes). It needs a vector index (semantic search over unstructured support transcripts). It needs a knowledge base (product documentation, cancellation policies). And it needs all of these to be connected: the customer ID in the SQL database links to the email thread in the document store, which links to the relevant policy in the knowledge base.

Most agent frameworks bolt on a vector database and call it done. The agent gets semantic search over chunked text, and that is the entire data story. SQL is "external": you write a custom tool integration. Documents are "external": you upload them to a separate service. The knowledge base is a prompt prefix.

VIMS takes a different approach: a vertically integrated data and knowledge stack where SQL, vector search, document indexing, and media encoding are first-class subsystems of the operating system, accessible to agents as native tools, and shareable peer-to-peer.

The Database App

VIMS includes a full-featured Database App that connects to PostgreSQL, MySQL, SQLite, and SQL Server. It is a complete database workbench designed for both humans and agents.

Schema Discovery

When you connect a database, VIMS reads the full schema: tables, views, indexes, foreign keys, and triggers. The schema is presented as an ER diagram: nodes are tables with their columns and primary keys, edges are foreign-key relationships. The ER graph doubles as an agent tool, so an agent can reason about the database structure before writing queries.

Schema search lets you find tables, columns, views, indexes, and triggers by name, with substring match ranked by exactness. When an agent needs to find "the table that stores customer churn status," it searches the schema, finds the relevant table, and reads its columns.

The SQL Editor

The SQL editor is schema-aware. Autocomplete suggests keywords, dialect-specific functions, table names, and column names based on the connected database. Queries are classified before execution: VIMS distinguishes SELECT (read-only), INSERT/UPDATE/DELETE (write), and DROP/TRUNCATE/ALTER (dangerous), and applies appropriate safety gates.

For write operations, the editor enforces a stage → review → approve → merge workflow. Changes are staged, reviewed via diff, approved through HITL, and then executed. Before any write, VIMS automatically creates a point-in-time backup, so if the query goes wrong, you can restore. (Blog 03: Security First)

Natural-Language to SQL

The database chat feature lets you ask questions in plain English. "How much did we spend on cloud last quarter?" becomes a SELECT query with an explanation of what it does. If the query is read-only, it executes automatically and returns results. If it involves writes, it goes through the write confirmation flow.

The assistant is grounded in the actual schema: it reads the table names, column types, and foreign key relationships before generating SQL. It knows the difference between a VARCHAR and a JSONB column. It knows which tables are related and how to join them.

For complex scenarios, the assistant has specialized modes: explain (describe what a query does), optimize (propose a more efficient rewrite), and fix (correct a failing query given the error message). (Blog 05: MCP Connectors)

Cross-Database Operations

VIMS supports operations across multiple connections. Schema diff compares two databases, showing added/removed tables and added/removed/changed columns. Data compare finds rows that exist in one database but not the other, or that have changed. Table transfer copies data between connections, with optional table creation and truncation.

These are agent-accessible tools. An agent can diff a staging database against production, identify drift, and propose a migration script, all within the governed security model.

SQL Source Mirroring

For offline-first scenarios, VIMS creates read-only local replicas of remote databases. A mirror is a SQLite snapshot of a remote database: same schema, same data, available without a network connection. The mirror can be refreshed on demand or on a schedule.

This means an agent can query a production database even when the network is down, even when the remote server is under maintenance, even when you are on a plane. The data is local. Queries run against the mirror, not the remote server. When connectivity returns, the mirror is refreshed.

Mirrors are P2P-shareable. A team member can share their database mirror with the team, and every peer gets a local copy, with no VPN, no connection-string sharing, no firewall rules. (Blog 06: P2P Surfaces)

Vector Search

VIMS includes a vector search engine using HNSW (Hierarchical Navigable Small World) indices. Documents are chunked, embedded, and indexed. Retrieval is hybrid: it balances recency (newer documents are weighted higher) with relevance (semantic similarity to the query).

The vector index is part of the knowledge base, which is part of the operating system. Agents search the KB via a natural-language query and get ranked snippets with source citations. The KB is the same surface that indexes uploaded files, chat conversations, agent reasoning transcripts, and swarm outputs. Everything an agent has ever read, written, or reasoned about is searchable.

The KB is also P2P-shareable. A team can share a knowledge base, and every peer can search it. The index is replicated over the swarm, raw documents and vector embeddings both. This means semantic search works offline, on your local copy of the KB, with the same relevance ranking as the online version. (Blog 11: Flow and Teams)

Pixe Encoding

VIMS introduces a unique media encoding pipeline called Pixe. Documents are converted to QR-encoded video: a visual representation that can be stored, transmitted, and decoded losslessly. Each frame of the video contains a QR code encoding a chunk of the document; the full video is the complete document.

Why? Because video is the most resilient storage medium. You can store a Pixe-encoded video on a hard drive, stream it over a network, project it on a wall, or even print it on film. The encoding is self-describing, with no external metadata needed. The decoder reads the QR frames, reassembles the chunks, and reconstructs the original document.

The Pixe pipeline includes vector indexing: each encoded document is chunked and embedded, so the encoded library is searchable via the same vector search engine as the KB. You can search for a concept, find the matching Pixe video, and decode it, all programmatically.

P2P Key-Gated Sharing

Both the KB and Database mirrors are P2P primitives. Sharing is key-gated: the owner generates an access key with a specific scope (read, write, admin, or limited to specific sources). The key is a plaintext secret shown once at generation time; the system stores only the hash.

Recipients use the key to access the shared resource over the P2P swarm. The key can be revoked at any time, and revocation is immediate: the recipient's access is cut off on the next swarm sync. Keys can have expiry dates, so temporary access is automatic.

This is how teams share knowledge bases and database mirrors without a central server. The owner controls access; the swarm handles replication; the keys enforce permissions. (Blog 06: P2P Surfaces)

The Integrated Picture

An agent in VIMS does not need to be told where data lives. The database, the knowledge base, the vector index, and the Pixe library are all native tools. When a flow includes a database node, it queries the connected database. When it includes an MCP node, it can search the KB. When a team shares a database mirror, every peer's agent can query it. When an agent reasons over a document, the reasoning is indexed in the KB for future search.

The data stack is part of the operating system the agent runs on.


Previous: Security First Next: MCP Connectors: Client and Server