← All posts

Structured Data Markup for AI Agent Discoverability

Schema markup has become essential infrastructure for AI agents evaluating B2B software vendors.

September 28, 2026

Cover illustration for “Structured Data Markup for AI Agent Discoverability”

A procurement lead at a mid-size logistics firm types a single line into an AI agent: compare enterprise accounting platforms that integrate with our ERP and hold SOC 2 Type II. Within seconds, a comparison table comes back, three columns, feature rows, pricing tiers. One vendor everybody in the industry would expect to see just isn't there. Not because the product is worse. It's missing because the agent couldn't parse what the product does.

That's the new default in B2B software buying. Instead of a person clicking through search results and landing on a homepage, an agent takes the request, filters by what the buyer actually means, builds a feature matrix on the fly, and in some deployments now, completes the purchase once the buyer's conditions are met. No ad gets clicked. No hero banner gets glanced at. No call-to-action button gets pressed. The agent reads structured data and nothing else.

The adoption numbers back up how fast this happened. Gartner has projected a 25 percent drop in traditional search volume during 2026 as agents take over the work Google used to do for product research. Forrester expects at least one in five B2B sellers to be responding to AI-driven buyer agents with automated counteroffers in 2026. Gartner's own forecast for agentic AI in procurement points to explosive growth by 2030, though it's worth being clear that's a projection, not a number anyone has already banked.

None of this is subtle for vendors. If an AI system can't read your product data, you don't lose the deal, you never make the list in the first place. Merit stops being the variable. Parseability is.

What AI agents read during a vendor evaluation

Discoverability in this world has nothing to do with rank. It's a binary: does your product show up inside the generated answer, or does it not.

Four behaviors, taken together, explain what makes a vendor land in that table. First, intent filtering: the agent works from an operational sentence, not a keyword. "Find enterprise workflow platforms that integrate with this ERP and meet SOC 2 Type II under a given ACV" It is a request an agent has to decompose into checkable facts. Second, it builds a feature matrix straight from public technical documentation, review sites, and independent write-ups, never from a narrative tour of your homepage. Third, in the more advanced setups already running in enterprise procurement, the agent watches pricing and contract terms and fires off a purchase once its criteria clear, using catalog pricing and API-based ordering rather than a human signing anything. Fourth, some agents track consumption and schedule reorders on their own, cutting people out of routine restocking decisions.

Citation behavior has already come apart from ranking in a way that should worry anyone still optimizing for position ten and above the fold. Ranking well and getting cited are now two separate games.

Codewave's case work makes the failure mode concrete. An engineering IT company with genuinely strong documentation still lost comparisons because a large chunk of its critical product information was unstructured or contradicted itself across different pages. Agents could find the company. They just couldn't answer a direct architectural question about it, and moved on to a competitor whose data held together. That's the whole problem in miniature: findable is not the same as answerable.

Agents also seem to discount claims they can't check. When ChatGPT, Perplexity, or Google's AI answer "best tools for X," they pull from pages carrying real numbers, named sources, and clear structure. A page that just says "we're the leading platform" gives the engine nothing to cite, so it doesn't. The agent-era version of domain authority isn't design polish or ad budget. It's three things: structured data markup, how complete the semantics are, and whether independent third parties back the claim up.

Schema markup as the foundational layer for AI search visibility

Schema markup used to be a way to earn a rich snippet, a star rating under a search result, a little extra real estate on the page. That job description no longer covers what it does. By 2026, schema had become the layer AI engines actually depend on to verify facts, map entities, and decide who gets cited, with ChatGPT, Gemini, Claude, and Perplexity all parsing structured data as part of how they build an answer.

This isn't inference from behavior alone. The platforms said so directly. Microsoft confirmed at SMX Munich in March 2025 that it uses schema markup for its generative AI features, Google made comparable statements around the same window at its Search Central Live events, and ChatGPT later confirmed it uses structured data to decide which products even show up in its results. When three separate platforms independently confirm the same mechanism, that constitutes infrastructure. That's infrastructure.

The practical shift for a B2B brand is that discoverability now runs through data architecture more than through web design or how much gets spent on ads. Structured data has effectively become the front door.

Which creates a specific, avoidable failure. If your specs only live in a PDF white paper, or sit behind a form fill before anyone can read them, an agent likely can't index that content at all, and it will default to whichever competitor made the same information easy to reach. This is a solvable problem, not a mysterious one, and that's exactly why it's inexcusable to leave unsolved.

A handful of schema types carry outsized weight for B2B software vendors specifically. Organization schema fixes brand identity across sources, linking a company's name, logo, and web presence into one entity an agent can recognize consistently. Service schema is the one that answers "what does this thing actually do," which is the exact question the Codewave example shows agents failing to get answered when the data's messy. Article schema on blog posts and research content is what earns citation from independent publishers rather than just the vendor's own site. FAQPage schema, when it's built around real questions buyers actually ask rather than manufactured ones, appears directly inside AI answer formats. And Person schema, attached to founders, engineers, or named specialists, anchors credibility to an actual human being an agent can treat as a verifiable source rather than an anonymous corporate voice.

Its product pages used to be feature-heavy and visually rich, the kind of page a design team would be proud of, but gave an agent nothing quick to grab onto. The fix was almost embarrassingly simple: lead with one plain sentence defining what the product does, then three bullets on who it's for and what problem it solves. Agents lean on meta titles and descriptions as a first pass before deciding whether to read further, and that reliance is what makes that opening sentence do more work than the entire visual design produces.

A business without clean, error-free structured data is functionally invisible on the fastest-growing discovery channel in software procurement today.

The emerging protocol layer, MCP, UCP, and agent-readable endpoints beyond schema

Schema markup is the floor, not the ceiling. Beyond it, a separate layer of protocols is emerging that lets an agent query a vendor's systems directly, in real time, rather than just reading static tags on a page.

It's not settled yet. WebMCP manifests, mcp.json, and agent.json are all competing signal formats right now, a genuinely early-stage environment, with most observers expecting real consolidation to take another 12 to 18 months.

UCP lets a business advertise what it can do at a dedicated discovery endpoint, and covers catalog search, cart building, identity linking, checkout, and order management, running on REST and JSON-RPC with support for AP2, A2A, and MCP built in rather than competing against them. It was published January 11, 2026 with Shopify, Etsy, Wayfair, Target, Walmart, and more than twenty other partners.

The advice from people actually watching this space closely is to move now rather than wait for the standard to finish settling. Publishing and maintaining agent contracts on a primary domain, even in draft form, puts a vendor in a better position once a final standard does emerge, and the concrete first step is auditing the domain for what agent capabilities already exist versus what's actually announced, then assigning someone to own closing that gap. Platforms built "AI-native," with intelligence baked into the data model and workflows from the start, behave differently than "AI-added" tools that bolt AI features onto older systems, and that same distinction applies to how a vendor should think about its own agent-facing infrastructure. It has emerged as the leading candidate for agent discovery standardization, with major platforms coalescing around it through 2025 and into 2026, as reported by efficientlyconnected.com. Google's Universal Commerce Protocol (UCP).

The go-to-market implication is straightforward once you see it. A vendor that publishes a clean, queryable endpoint gives a procurement agent something concrete to fetch, confirm, and cite. A vendor that doesn't is gambling that the agent will reconstruct its capabilities correctly from marketing copy, which is a bad bet. But being queryable only answers whether an agent can find you and read your shape. It says nothing about whether the agent believes what it reads, which is a separate problem. Model Context Protocol (MCP). It originated at Anthropic before being donated in December 2025 to the Agentic AI Foundation, a directed fund under the Linux Foundation co-founded by Anthropic, Block, and OpenAI, placing it under neutral governance rather than one vendor's roadmap, as reported by anglera.com.

Why discoverability alone does not win the evaluation

Confidence in individual reviews fell from 79 percent to 42 percent between 2020 and 2025. It marks a sizable wobble. It's a sign that the credibility scaffolding built for human buyers, testimonials, star ratings, review aggregates, was already weakening before agents entered the picture, and it appears to be eroding faster now that agents are the ones doing the reading.

Getting found and getting picked are two separate problems, and structured data only solves the first one. Gartner's research shows B2B buyers spend a genuinely small slice of their time actually meeting with potential suppliers; almost everything else is self-directed research. When an agent is doing that research instead of a person, it can't respond to a case study's narrative arc or feel anything about a customer's story. It can only pull out facts it can verify. A striking majority of B2B buyers, per UserEvidence's Evidence Gap report, say they've ruled out a vendor specifically because the evidence offered didn't hold up.

What survives that scrutiny is narrow and specific: concrete numbers, named sources, structure an agent can parse cleanly. A customer story with a real name attached, a real metric, and an explanation of the mechanism behind the result, that's citation material. "We're the leading platform" is not, no matter how confidently it's written.

The very investments that make a vendor discoverable to an agent also put every unverifiable claim directly in that agent's path. Structured data gets you read. It doesn't get you believed. Social proof now has to work for two audiences at once, the human buyer looking at the site and the AI model that buyer consulted before ever clicking through. Forrester's State of Business Buying research found generative AI tools were the single most-cited meaningful interaction type buyers used while researching purchases, yet a real share of those buyers reported feeling less confident afterward, having run into unreliable or flatly inaccurate information along the way. Some meaningful portion of that bad experience traces straight back to vendor pages and catalogs that answer questions poorly, which puts the responsibility back on the vendor side of the equation. Structured data is the entry ticket. What it carries has to hold up on its own.

Traditional social proof as seen by a diligent AI agent

Walking through what most B2B SaaS vendors are actually bringing into an agent evaluation right now makes the gap obvious fast.

Logo walls are a row of brand marks with no claim attached to any of them. An agent can't verify that any specific company listed there uses any specific feature, because the logo asserts nothing checkable. It's decoration standing in for evidence.

Testimonial quotes tend to be unattributed, or attributed loosely enough that an agent has no way to confirm the person exists, that their company is actually a paying customer, or that the outcome described happened the way the quote implies.

Static case study PDFs are written for a person reading start to finish, in order, absorbing a narrative. They're frequently locked behind a form fill, which makes them invisible to a crawler outright, and even when they're openly accessible, they usually carry no machine-readable structure at all.

Star ratings and aggregate review scores collapse everything into one number with no resolution at the feature or use-case level.

An AI agent evaluating social proof encounters an unverified string it cannot confirm independently, and a string it can't confirm is a string it generally won't cite. The proof B2B software vendors have spent years perfecting was built for exactly the audience that's no longer making the first cut, which is the mismatch driving this shift. This is the asset inventory most B2B SaaS vendors carry into an agent evaluation.

Positioning to AI Agents

Sources