Skip to main content

Command Palette

Search for a command to run...

An AI-ready data layer for enterprise agents

What grounding is, why chained agents make it the first thing to get right, and what one shared data layer is worth when every business unit builds agents on it

Updated
•13 min read•View as Markdown

The goal: one AI-ready data layer that every business unit can build agents on, with each unit keeping control of its own definitions, access and audit. Squid generates a grounded AI ontology from the systems you already run, has it confirmed by the people who own the definitions, and puts every source under one governed, real-time context layer that employees and agents share.

Days, not months to a new production agent on a layer that already exists · 95% less build effort for every agent after the first · Any harness, your agent framework or ours — open, extensible, model-agnostic


01. The goal: one data layer, controlled by each business unit

Most large companies now have agent projects running in several business units at once. Finance wants an agent that reconciles billing. The customer team wants one that spots accounts about to leave. Support wants help triaging tickets.

Each team tends to start the same way. It finds the data, works out which of several tables is the real one, writes the joins, re-implements the access rules, and builds a test set. A few months later the next team does the same work on much of the same data, and arrives at a slightly different definition of "active customer."

Gartner expects organizations to abandon 60% of AI projects that lack AI-ready data through 2026.[^1] In our experience the model is rarely what stalls these projects. The data work underneath each agent is, and it gets repeated for every project.

So the goal we build toward is a layer that does that work once and lets every agent reuse it. An AI-ready data layer has four properties:

  1. Agents query data where it lives. No new copy, no migration project on the critical path.
  2. Business terms have one agreed definition. "Active subscriber" or "net revenue" means the same thing to every agent, and a named owner approved it.
  3. Access is checked on every request, as the person asking. The agent can reach what that user can reach, and nothing more.
  4. Every answer can be traced and scored. You can see which rows and passages it came from, and measure how often answers are supported by them.

The ontology is for employees too

The same layer serves people as well as agents. Squid generates a grounded AI ontology of the business: its customers, products, contracts, sites and the relationships between them, with every entity and definition linked back to the records it came from. Employees can browse it or ask questions of it in plain language. An analyst asking "which enterprise accounts have open billing disputes and a renewal this quarter?" gets an answer assembled from billing, CRM and ticketing, with the sources cited and filtered to what that analyst is allowed to see. Because the ontology is generated from live data and confirmed by owners, it stays current, where a hand-drawn model tends to drift out of date.

That matters for adoption. Employees who use the ontology every day are also the people who notice when a definition is wrong, and their corrections flow back to the owners for approval. The layer gets better through use.

The second half of the goal matters as much as the first. A shared layer only works in a large company if it doesn't take control away from the business units. So the layer is shared and the control is federated:

Shared, built once Controlled by each business unit
Semantic layer, grounded AI ontology and knowledge graph, identity and policy engine, in-place query execution, evaluation harness Its definitions, its access policy, its audit view, its choice of model, its usage limits

Key terms

  • Grounding: tying each step of an agent's work to your own data, with the source attached, so the answer can be checked against the rows or passages it cites.
  • Permissions (access control): deciding whether the person asking may see those sources at all. A separate question from grounding: an answer can be perfectly grounded and still show someone a salary they should not see.
  • Semantic layer: the agreed business meaning of your data: what "net revenue" is, which table is authoritative, how two systems join.
  • Ontology and knowledge graph: the entities in your business, how they relate, and the records that belong to each. A grounded AI ontology is generated from your data, with every entity linked to its source records.
  • Control plane: borrowed from networking: the part of a system that decides what may happen, kept separate from the part that does the work. Here, the layer that decides what an agent may read, run and write, and records what it did.
  • Grounding rate: the share of claims in an agent's answers that an independent evaluator finds supported by the sources actually retrieved, reported per agent and per business unit.

02. Two questions to ask before the next agent ships

1. Which systems can this agent read today, and who approved that? A common answer is a service account with wide read access, plus a line in the system prompt asking the model to be careful with sensitive data. An auditor will treat the first part as the real access level and the second as a hope.

2. When the agent gets something wrong, how do we find out? Plenty of teams learn about errors from the user who received them.

The two questions point at two different problems. The first is about permissions; the second is about grounding. A system can solve one and fail the other. We build them in the same layer because both have to act on the same request path, at the moment the agent reaches for data.

03. Accuracy: errors compound across steps, and what to do about it

A typical agent request has it pick a source, write a query, check the query, read a document, reconcile the two, and then act. If each of five steps is right 95% of the time and the errors are independent, the chain is right 77% of the time. At 99% per step it finishes at 95%. Over eight steps, four points of per-step accuracy become a 26-point gap (92.3% vs 66.3%).

Steps 95% per step 99% per step Grounded + checks and retry (≈99.8%)
1 95.0% 99.0% 99.8%
5 77.4% 95.1% 99.0%
8 66.3% 92.3% 98.3%

Illustrative. The first two columns assume independent failures and no recovery. The third assumes checks catch 80% of step errors and the retry succeeds at the grounded rate; the 80% is an example, not a measurement.

The model is too simple in two directions: errors are often correlated (a wrong definition at step one spoils every later step), and real systems catch errors before they propagate. What we do about the decay, in order of impact:

  1. Fewer steps. The semantic layer and knowledge graph already hold which table, which join and which ID matches which account, so the agent makes one resolved lookup instead of three guesses.
  2. Check each step against something. Queries are validated against the schema and approved definitions before they run; claims are checked against retrieved rows and passages. The grounding layer supplies the reference to check against.
  3. Recover inside the turn. On a failed check, the agent corrects and retries before answering.
  4. Stop before anything irreversible. Writes go through human approval, so a surviving error shows up as a rejected proposal before anything changes.

Stronger models help as well: they write better queries and reason more carefully over results. What a model cannot do on its own is know your definitions, authoritative sources or users' entitlements.

End-to-end accuracy. We measure the whole chain. Each deployment gets test questions with known answers, written with the business owners; a separate evaluator agent scores every response. In a benchmark run with a global investment bank, DataMind answered all 120 of the bank's benchmark questions correctly. That is evidence, not a guarantee, so in production we keep scoring live traffic, track the grounding rate per agent, and collect user ratings.

04. Governance: where the controls sit

Centralizing data before governing it made sense for analytics. For agents it usually means a long migration and a second copy of sensitive data whose permissions drift from the original's. The alternative is to govern at the point of access.

On a structured question, the model receives the schema and writes a query. The query runs in the source system under the asking user's permissions. The model then sees the result that query returns for that user, usually small, filtered or aggregated. It does not receive whole tables, and it never sees rows the user couldn't see directly.

Six control points:

  1. Identity inherited from Okta, Entra ID, Auth0, Cognito or Keycloak (read path)
  2. Authorization per request: RBAC and ABAC, respecting source-system permissions (read path)
  3. Model field of view: schema plus the permitted result set, configured per business unit (read path)
  4. Data stays in place, in-region where residency rules apply (read path)
  5. Writes gated by a person before they execute (write path)
  6. Every turn logged and replayable (both)

A prompt instruction is advice to a model. A query-time authorization check is a control.

05. What a grounded data layer gives you

Grounding is usually discussed as an accuracy technique. In practice it is infrastructure. Once the semantic layer, the knowledge graph and the ontology exist and are governed, they stop being part of any one project and become something every agent after the first inherits. Six things change.

Pillar Anchor Detail
Ready Agentic data governance All data, AI-ready One governed path to every source, structured and unstructured alike. Identity, policy, lineage and audit applied on the data path, not written into prompts. No migration required.
Grounding context Real time Agents read the current state of your systems at the moment of the request, not a nightly snapshot or a drifted index.
Fast Time to market Days, not months A new agent ships in days on a layer that already exists. Built one at a time, with the data work repeated each round, the same agent takes months.
Build effort 95% less for every agent after the first. Connectors, entity resolution, approved definitions, access policies and test sets are built once and reused by every team.
Open Any harness Yours or ours Bring your own agent framework or use Squid's, through SDKs, MCP and A2A. Model-agnostic, including models in your own tenant.
The ontology Keeps learning Generated in depth from your systems, confirmed by the people who own the definitions, continuously updated as new sources and approvals arrive. Versioned, and yours to keep.

The first agent pays for the layer. Every agent after it inherits one. That single fact is what moves time to market from months to days and takes most of the build effort out of every project that follows.

Two second-order effects follow from the same design. Inference costs fall, because an agent that has context does not push rows, documents and retries into the prompt to compensate for what it is missing; the model receives the schema and the database does the retrieval. Review costs fall, because an answer that cites the database row and the contract clause behind it takes a glance to check rather than an investigation.

06. How we build it at Squid

DataMind is our implementation of this layer. Grounding comes first because evaluation, governance and reuse all depend on a shared, approved model of what the data means.

  1. Semantic layer: business vocabulary inferred from schemas, joins and documents, approved by definition owners, versioned.
  2. Grounded AI ontology and knowledge graph: generated from records, tickets, contracts and code, each entity linked to its sources; employees can browse and query it directly, and it keeps learning.
  3. Query execution: the model writes a validated query that runs in the source system's own dialect, under the user's permissions; the model sees only the permitted result set.
  4. Governed agents: multi-step work with checks and retries, human approval for writes, and a log of every turn.

Ontology-first platforms ask your team to model the enterprise and then add AI, which gives a rigorous result at a high up-front cost. We infer a first model from your data and ask your owners to confirm it, which is faster to start and depends on the quality of your existing schemas and documents.

07. How you know it is working

  • An evaluator agent scores each response against the sources retrieved; the share of supported claims is the grounding rate.
  • Continuous evaluation runs test questions and scores live traffic, so regressions surface before users find them.
  • Reflection and retry lets the agent correct its own failures within the turn.

Six questions for your next vendor meeting

  1. On a structured question, what does the model actually see: the schema, the result set, or copied tables?
  2. Whose identity does the agent act as, the user's or a service account's?
  3. Where do the business definitions come from, and who approves changes?
  4. When two systems disagree, does the agent pick one or show the conflict?
  5. Can each business unit own its definitions, access policy and audit view on one shared layer?
  6. What was the grounding rate on last week's production traffic, and how was it measured?

If you're planning agents in more than one business unit, the most useful early decision is whether they will share a data layer. It changes the cost of every agent that follows, and it's much harder to retrofit after five teams have built their own.


About DataMind: Squid AI's AI-ready data layer for enterprise agents and the employees who work with them: a generated, grounded AI ontology and knowledge graph, semantic layer, in-place query execution and governed agents, on your data and your identity provider, with controls for each business unit. SOC 2 Type II · ISO 27001. getsquid.ai/solutions/datamind

[1]: Gartner, "Lack of AI-Ready Data Puts AI Projects at Risk," February 2025.