Vector databases became the default infrastructure choice for AI agent context in 2023 and 2024. The pattern is everywhere: chunk your documents, embed them, store them in Pinecone or Weaviate or pgvector, retrieve by semantic similarity, inject into the prompt. It works well for a specific class of problems. It is the wrong tool for a different class that enterprise teams frequently conflate with the first one.
The distinction matters because building on the wrong abstraction creates problems that are hard to fix later. If your AI agent needs to query live transactional data, the vector database retrieval pattern is not just suboptimal: it is architecturally incompatible with what you need.
What vector databases are actually good at
Vector databases solve the problem of semantic retrieval over a static or slowly-changing document corpus. You have a large set of text documents, you want to retrieve the ones most semantically relevant to a query, and you cannot enumerate the relevant documents by structured key. Knowledge bases, documentation corpora, FAQ databases, support ticket archives. The retrieval is approximate but useful, the latency is acceptable for synchronous use, and the data can tolerate a staleness window of hours or days.
These characteristics describe a retrieval-augmented generation (RAG) pattern. RAG is genuinely useful for a large fraction of AI assistant use cases. It is not the architecture for enterprise operational data access.
Where vector retrieval breaks down for operational data
Operational enterprise data has properties that are incompatible with vector retrieval. The data changes continuously. A customer's account balance changes with every transaction. An open service ticket changes status several times a day. An inventory record changes with every shipment. If you embed this data and store it in a vector database, it is stale by the time the agent queries it. The staleness window that is acceptable for documentation retrieval, measured in hours, is not acceptable for data that drives operational decisions.
The second problem is precision. Vector retrieval is approximate: it returns documents that are semantically similar to a query. For operational data, the agent does not need semantically similar records. It needs the specific customer record with a given account number, the specific invoice with a given ID, the specific contract associated with a specific opportunity. These are structured key lookups, not semantic searches. Embedding structured records and retrieving by vector similarity adds noise that structured queries do not have.
The third problem is access control. Vector databases do not have row-level or field-level access controls as a first-class feature. Adding access control requires retrieving records and filtering them in application code, which means the data was already retrieved and potentially logged before the access decision was made. For regulated enterprise data, access control at the retrieval boundary is a hard requirement, not an application-layer concern.
The context bus model
A context bus is a different architecture. Instead of pre-indexing data into a separate store, the context bus maintains live connections to authoritative data sources and executes structured queries at request time. The agent describes what context it needs, using a typed interface, and the bus routes that request to the appropriate connector, enforces access policy, executes the query against the live source, and returns the result.
The key properties of this model are: data freshness (queries hit live sources, not cached embeddings), structural precision (queries are structured and typed, not approximate), access control at the query boundary (the policy is evaluated before data is retrieved and returned), and full auditability (every query event is captured with its authorization context).
This architecture is more complex to build than a vector retrieval pipeline. It requires maintained connector infrastructure, a policy engine, and a query routing layer. The tradeoff is that it is the only architecture that is actually correct for operational enterprise data.
When to use which
These architectures are not mutually exclusive. Many enterprise AI systems need both. The right way to think about it is by data type and query type.
Use vector retrieval for: policy documents, product documentation, historical ticket archives, training materials, FAQ corpora. The data is text-heavy, updates infrequently, and retrieval is by semantic intent rather than structured key. Staleness of hours or days is acceptable.
Use a context bus for: customer records, financial transactions, HR data, inventory, CRM opportunity data, service tickets. The data is structured, changes frequently, has precise key-based access patterns, and requires access control and audit at the retrieval layer. Staleness of more than seconds or minutes is not acceptable.
The confusion arises because both patterns are described as "giving the agent access to data." At that level of abstraction, they look the same. The implementation differences only become visible when you ask specific questions: Is this data live or document-style? Do access policies apply per-field? Does a compliance team need to see exactly what was retrieved?
What this means for teams evaluating AI agent infrastructure
If your AI agent needs to work with enterprise operational data, vector database infrastructure alone will not serve you. You can use it for the document retrieval portions of your context strategy. For live operational data, you need connector infrastructure with live query execution, policy enforcement, and audit capture.
This is not a criticism of vector databases. They are excellent tools for the problems they were designed to solve. The mistake is using them as the default answer to "how does the agent access data" without asking whether the data in question fits the vector retrieval model. For most enterprise operational data, it does not.