Enterprise AI

Why AI agents give inconsistent answers — and the semantic layer they're missing

A semantic layer for AI agents defines business terms once. Only 27% of companies have one, and Microsoft and Bloomberg just shipped products that assume you do.

Updated 8 min read

Two vendors shipped the same admission in the same week.

On 29 September 2026, Bloomberg launched Enterprise MCP, a way for clients’ AI agents to pull its licensed data. The pitch wasn’t access. It was meaning. The next day, Microsoft opened the preview of Work IQ’s business context, built to ground Copilot and agents in what Dynamics 365 records mean and which rules apply to them.

Neither one sells a better model. Both sell the layer underneath.

TL;DR: A semantic layer for AI agents is the governed, machine-readable place where terms like “active customer,” “on-hand,” and “complete” are defined once, so every agent resolves them the same way. TDWI’s 2026 benchmark found only 27% of organizations have one. That’s why agents give confident, inconsistent answers about a company’s own data: two systems carry two definitions, and the agent inherits whichever it queried. Microsoft and Bloomberg now ship this layer as a product, but only for meaning that has already been written down. Define your terms and procedures first, make drift fail loud, then turn the agent on.

The model got better. The definitions didn’t.

On June 26, 2026, OpenAI previewed GPT-5.6 Sol, built for long-horizon agentic work. It set a new state of the art on the Terminal-Bench 2.1 multi-step benchmark. It’s a real jump in what a single agent can plan and carry out.

Here’s what none of that touches. Ask that model, running as an agent inside your company, “how many active customers do we have,” and the answer depends on which system it reads. Your CRM says active means logged in within 90 days. Billing says active means holding an unexpired contract. Both are internally correct. Neither is the company’s answer, because the company never wrote one down. The smarter model just returns the wrong-for-your-purposes number with more composure.

Bloomberg’s head of enterprise data, Tony McManus, put it plainly in the launch release. The bottleneck in financial services, he said, has moved from model capability to data readiness. His example: a last price, on its own, doesn’t say what currency it’s in or what kind of price it is. And his summary: “Connections were never the hard part. Readiness is.”

Reasoning was never the broken layer. Meaning was.

The named object under every “data contract”

I’ve written before that integration and data contracts decide whether AI works, and that agentic planning breaks when your systems don’t agree on a number. The semantic layer is the thing those pieces were circling. It’s one governed, machine-consumable place where each business term has a single definition. It sits between the raw systems and every agent that reads them.

Without it, the definition doesn’t disappear. It scatters. It lives in a SQL view one analyst wrote in 2023, in a field label that means one thing in Salesforce and another in NetSuite, in the head of the ops manager who knows “complete” means dispatched in the field app but invoiced in finance. An agent querying across them doesn’t get a contradiction it can flag. It gets one answer, confidently, from whichever table it happened to hit.

A human analyst catches this. They ask “active by which definition?” The agent doesn’t ask. It answers.

September 2026: the platforms start selling the layer

The last week of September is worth reading closely, because two very different vendors described the same missing piece in almost the same words.

Bloomberg Enterprise MCP (29 September). It exposes Data License Plus content, more than 100 million securities and over 50,000 fields, through the Model Context Protocol. The part that matters is the metadata. Per Bloomberg’s release, it describes what each field means, how it was calculated and when it applies, plus the qualifiers that make a number interpretable: exchange, currency, price type, period, as-of date. On top of that sit packaged “Skills”: finding outliers in a set of tickers, reviewing trades that breach price deviation thresholds, screening securities for sanctions exposure.

Microsoft Work IQ business context (preview from 30 September). Per Microsoft’s Dynamics 365 blog of 25 September, Work IQ grounds Copilot and agents in Dynamics 365 and Power Platform data, with Dataverse as the storage layer. Its semantic model derives context automatically from your existing Dynamics 365 and Power Apps configuration. Alongside it come business skills, which Microsoft defines as instructions for a business task combined with the context and resources an agent needs to follow them. The example is a renewal: check unresolved critical cases, confirm a service improvement plan owner, document the customer commitment, get approval before offering a concession.

Put the two next to each other and the pattern is obvious. Metadata that says what a field means. Procedures that say what to do with it. That’s a semantic layer plus a written process, sold as a feature.

Now look at why each one works.

Bloomberg can ship 50,000 defined fields because defining fields is Bloomberg’s product. Somebody there has been writing down “how was this calculated, and when does it apply” for decades. The MCP server is the last step on top of that.

Microsoft’s version derives context from your configuration. So it inherits your configuration. If your opportunity stages were set up in 2019 and half the sales team uses “Closed Lost – Other” to mean “handed to a partner,” Work IQ will learn exactly that, faithfully. If the real renewal rule is “never discount a customer with an open escalation, unless the VP says so,” and that rule lives in one sales manager’s head, it isn’t in Dataverse. There’s nothing to derive.

That four-step renewal skill in Microsoft’s example looks simple. It’s also the output of a process-mapping exercise most companies have never done. Somebody had to decide that an open critical case blocks a renewal, that a service plan needs a named owner, and that concessions need approval. Work IQ makes that decision cheap to store and easy for an agent to follow. It doesn’t make it any cheaper to discover.

Six sources, one finding

Here are the independent data-readiness studies and vendor launches across 2026, side by side. None of them is about model quality.

Source (2026)Sample / typeThe finding
TDWI Agentic AI Readiness Benchmark (May)161 organizations27% have a governed, enterprise-wide, machine-consumable semantic layer; 47% report broadly trusted structured data; fewer than 10% run multi-agent systems in production
Fivetran Agentic AI Readiness Index (May)400 data professionals15% fully prepared for agentic AI in production, while 41% already run it there; 42% name data quality and lineage as a top barrier
Modern Data Report540+ data leaders80% rank a semantic layer with standardized definitions as the most important enabler of AI; 68% say their data isn’t trustworthy enough for AI; 93% encounter conflicting metrics
Google Cloud Knowledge Catalog (April)Vendor productBuilt to map and infer business meaning across the data estate, because “AI is only as smart as its context”
Bloomberg Enterprise MCP (29 Sep)Vendor productField-level metadata (meaning, calculation, applicability, currency, as-of date) across 50,000+ fields, plus packaged workflow Skills
Microsoft Work IQ business context (preview 30 Sep)Vendor productSemantic model derived from existing Dynamics 365 / Power Apps configuration, plus “business skills” that package a procedure for agents

Read down the right-hand column. The surveys say most companies are pushing agents toward production faster than their data can define itself. The vendors, one after another, are building the definition layer and selling it. Google said in April that an agent that doesn’t understand your definition of “margin” is forced to guess. Five months later Bloomberg and Microsoft shipped products whose whole job is to stop the guessing.

When the vendors selling you the agents start selling you the semantic layer, the argument is over.

Why “forced to guess” is the expensive part

Guessing sounds harmless until you follow it downstream. A planning agent doesn’t return “I’m not sure which definition you mean.” It returns a number, a routing decision, a drafted approval, as if the definition were settled. The error doesn’t announce itself. It rides quietly into a board deck, a customer promise, a replanning loop that runs every fifteen minutes on a figure two systems compute differently.

And it drifts. Someone renames a CRM stage from “Closed Won” to “Committed.” The connector still fires. The agent still runs. But the field it was counting now means something subtly different, and nothing in the stack notices, because a schema that still validates can carry a meaning that quietly changed. Three vendors in the pipeline, three definitions of the same event, and no one owns the reconciliation.

Derived context has the same exposure. A semantic model built from your configuration is a snapshot of that configuration. Change the configuration and the meaning moves with it, whether or not anyone decided it should.

The working version: define, then deploy

The fix isn’t a bigger model, and it isn’t a platform checkbox. It’s unglamorous, and it’s the entire job.

Define the terms the agent acts on, in a place a machine can read. Not a wiki page. Pick the decisions the agent will actually make and list the handful of terms those decisions turn on: active customer, on-hand, complete, margin, past due. Resolve each to one governed definition. You don’t need to model the whole enterprise. You need the terms in the agent’s path.

Decide which system wins. When the CRM and billing disagree on “active,” the semantic layer names the authoritative source and the rule. That’s a business call, not a data-engineering one. Which is exactly why it keeps not getting made.

Write the procedure before you package it. A business skill, a Bloomberg Skill, a Copilot Studio instruction set: they’re all containers. Fill one the way you’d brief a new hire. Every step, every exception, and who approves what. If you can’t write the renewal rule down in plain language, the agent can’t follow it either.

Make drift fail loud. When an upstream field is renamed or a value changes meaning, the pipeline should stop and flag, not keep answering on a definition that shifted underneath it. A loud failure is a ticket. A silent one is a number nobody questions until it’s expensive.

Start narrow, with a human on the output. One workflow, one set of defined terms, someone checking the agent before it acts. Earn the autonomy on data you’ve actually defined.

Per TDWI, fewer than one in ten organizations are running multi-agent systems in production. The ones that are didn’t win by buying the strongest model. They wrote down what their own words mean, and what their own process is, before they let an agent act on either.

The platforms will now store that for you. Writing it is still the work, and it’s the conversation worth having before the next agent goes live.

FAQ

What is a semantic layer for AI agents?
A semantic layer is the governed, machine-readable place where your business terms are defined once: what counts as an 'active customer,' what 'on-hand' inventory means, when an order is 'complete,' how 'margin' is calculated. It sits between the raw systems and the agents that read them, so every agent resolves a term the same way. Without it, each system carries its own quiet definition and an agent inherits whichever one it happens to query. TDWI's May 2026 Agentic AI Readiness Benchmark found just 27% of organizations have a governed, enterprise-wide semantic layer that's machine-consumable. That 27% is the real gate on putting agents into production.
What is a business skill in Microsoft Work IQ?
In Microsoft's words (Dynamics 365 blog, 25 September 2026), a business skill 'combines instructions for a business task with the context and supporting resources an agent needs to follow them.' Microsoft's example is a renewal: check unresolved critical cases, confirm a service improvement plan owner, document the customer commitment, and get approval before offering a concession. Business skills are part of Work IQ's business context, which grounds Copilot and agents in Dynamics 365 and Power Platform data stored in Dataverse. The preview starts 30 September 2026, with rollout continuing through October. A business skill is only as good as the procedure written into it, and at most companies nobody has written that procedure down yet.
Semantic layer vs business skill: what's the difference?
A semantic layer defines what your data means. A business skill defines what to do with it. The semantic layer answers 'what counts as an active customer, and which system is authoritative.' The business skill answers 'when a renewal comes up, check these cases, confirm this owner, get this approval before discounting.' An agent needs both. Microsoft's Work IQ and Bloomberg's Enterprise MCP, both released in the last week of September 2026, ship them together: context about the data plus packaged procedures. Neither half can be derived from software configuration alone if the definition or the procedure only exists in someone's head.
Why do AI agents give inconsistent or wrong answers about company data?
Because two systems hand the agent two different meanings of the same word and nothing reconciles them. Your CRM might define 'active customer' as anyone who logged in within 90 days, while billing defines it as anyone with an unexpired contract. Ask an agent 'how many active customers do we have' and the number depends on which table it read. Both answers are internally correct and neither is the company's answer. The agent isn't hallucinating. It's faithfully reporting an ambiguity you never resolved.
How many enterprises actually have a governed semantic layer?
Few. TDWI's May 2026 benchmark of 161 organizations found 27% have a governed enterprise-wide semantic layer and only 47% report broadly trusted structured data. Fivetran's 2026 Agentic AI Readiness Index, surveying 400 data professionals, found just 15% fully prepared to support agentic AI in production while 41% are already running it there. The Modern Data Report 2026 found 80% of data practitioners rank a semantic layer with standardized definitions as the most important enabler of AI, and 93% say they run into conflicting metrics.
How do you build a semantic layer before deploying AI agents?
Start from the decisions, not the data. Pick the decisions the agent will make, list the handful of terms those decisions depend on, and resolve each one to a single governed meaning with a named authoritative system. Write the procedure the agent should follow the way you'd write it for a new hire, including the approval step. Make the pipeline fail loud when an upstream system renames a field or changes what a value means. Then start narrow, one workflow, with a human checking the output. The model is the last step. The shared definition is the project.