Enterprise data was not designed with AI agents in mind. Your SAP instance, your Salesforce org, your data warehouse, and your customer platform all have access control models built around the assumption that a human or a well-defined service account is making requests. The requests are predictable. The scope of what gets accessed is bounded by a UI or by an application that was engineered to access specific tables or endpoints.
AI agents break both of those assumptions. An agent's data requests are generated dynamically based on its reasoning state. The scope of what it might access in a single session is wider than any UI-driven interaction. And the agent does not have the contextual judgment that a human employee has about what they should and should not be looking at.
These are not problems that agent frameworks solve. They are solved by the layer between the agent and the data. This article makes the case for why that layer needs to exist, and what happens when it does not.
What direct agent-to-data access actually looks like
The most common implementation pattern for AI agent data access is direct: the agent is given credentials for the data source, and a set of tools or functions that let it query the source directly. The agent decides which query to run, the query is executed, and the result is returned to the agent's context. The agent's next action is based on what it found.
This works at the proof-of-concept stage. The agent gets the data it needs. Developers can move fast. The integration is simple. The problems emerge when you ask: what exactly did the agent access? What policy was applied to that access? If the agent queried a customer record and returned information in a response, who authorized that retrieval and is there a record of it?
The answers are: we do not know precisely, no policy was applied, and no structured record exists. The data source logs may show that a service account made a query. The application logs may show that the agent produced an output. The connection between the two, what the agent accessed and why in order to produce that output, is not captured anywhere.
Failure mode one: access without policy
In direct agent-to-data integrations, the de facto access policy is the credential's permission scope. Whatever the service account can read, the agent can read. This is a coarse-grained control. In most enterprise environments, service accounts are provisioned for operational use cases that have broad read access to support multiple functions. The access scope that is appropriate for the service account's original use case is not necessarily appropriate for every AI agent use case.
In a regulated environment, this is not just an engineering problem. Healthcare organizations operating under HIPAA have a minimum necessary standard: access to Protected Health Information should be limited to the minimum information necessary for the purpose. Financial services organizations have similar principles in their data governance frameworks. An AI agent with broad service-account-level access to a regulated data source is structurally out of compliance with these principles, even if no one has explicitly made a wrong access decision. The policy gap is architectural.
A middleware layer with a policy engine enforces the correct access scope at the query boundary. The agent requests context. The middleware evaluates the request against a declared policy for that agent's identity. The policy specifies which fields, which record types, and which query patterns are permitted. The agent receives only what the policy allows, regardless of what the underlying credential permits.
Failure mode two: no audit trail
When an AI agent accesses data directly, audit reconstruction is possible only through the logs of the underlying data source. SAP RFC call logs, Salesforce API access logs, database query logs. These logs record that a query was made by a service account. They do not record which agent made the request, what the agent's stated purpose was, what policy rule was evaluated, or what was returned in the context.
This matters when you need to answer questions like: during the data access incident last month, which customer records did the AI agent touch? Was the agent accessing data within its authorized scope at the time? If a customer exercises their right of access under GDPR and asks what automated processing accessed their data, what can you tell them?
These are not edge cases. They are the normal operational and compliance questions that arise when AI agents become part of production data workflows. Without a middleware layer that captures structured audit events at the context request level, the answers are unavailable or require forensic reconstruction that may not be possible.
The middleware captures every context request as a structured event: agent identity, timestamp, connector, query scope, policy rule evaluated, resolution (permitted or denied), and a summary of what was returned. This event stream is the audit trail. It is produced automatically, at the correct layer of abstraction, without requiring individual agents to implement logging.
Failure mode three: access surface expansion as agents multiply
In most organizations that have been building AI agents for more than six months, there are multiple agents in production or near-production. Each was built by a team that needed access to data sources for their specific use case. Each integration was built independently. The collective access surface, which data sources any agent in the organization can reach, was never explicitly reviewed or bounded. It accumulated.
The access surface of a direct-integration agent fleet tends to expand monotonically. Adding a new agent means adding new credentials or reusing existing ones with broader scope. Access is never narrowed because narrowing requires coordinated effort across teams that do not have a shared governance mechanism. The question "what can our AI agents collectively access in our data environment" has no clean answer because the answer is distributed across N independent credential configurations.
A middleware layer with a connector registry and policy engine is the governance mechanism. Every agent that needs data access registers a connector configuration. Every connector configuration specifies which data source it connects to and which policy governs access. The total access surface of the agent fleet is visible in one place: the connector registry. Access scope for individual agents is visible in their policy configurations. When a team wants to give a new agent access to a new data source, they do so through the governed path, with an access decision that is recorded.
The objection: is this not just adding overhead?
Yes, a middleware layer adds engineering overhead relative to a direct integration. Building a connection from an agent to a data source is faster than building the same connection through a policy-enforced connector registry. The question is what the overhead is for, and whether the alternative is actually cheaper in total.
The overhead pays for three things: access control that is correct and documented, an audit trail that exists when it is needed, and an access surface that is bounded and visible. If your agents do not need those things, the overhead is not justified. If they do need them, the alternative is not "no overhead." It is a larger overhead paid later, often under adverse conditions: a compliance audit, a data incident, or a regulatory inquiry. That retroactive overhead is higher than the proactive overhead of building the middleware layer before deploying agents into regulated data environments.
We built Strattum because we kept seeing teams discover the three failure modes above after deployment rather than before. The middleware layer we are describing is not a theoretical architecture. It is the solution to a set of problems we observed repeatedly, in environments where the stakes of getting it wrong were not limited to engineering inconvenience.