Schema markup won't make AI cite you. Wrong schema can make it doubt you.
Schema gets sold as the switch for AI search: add the markup, get the citation. It isn't one, and brands that treat it that way end up with pages full of markup that does nothing, or markup that quietly contradicts what the page says. Structured data does real work for AI visibility, just not that work. Here's what it actually does, the three ways it goes wrong, and how to tell whether yours is helping.
Yes, but not the way it's usually sold. Schema markup doesn't earn a citation by itself. No engine treats a JSON-LD block as a reason to recommend you. What it does is remove ambiguity: it pins down which organisation you are, states your key facts in a form that doesn't depend on interpretation, and confirms what your page already says in plain text. That makes you easier to identify and harder to misdescribe. The flip side matters just as much: markup that contradicts your page, uses the vocabulary wrongly, or claims things you can't back gives engines a reason to doubt you. A passing validator won't catch most of it. The free audit shows whether your structured data is helping or hurting.
Every GEO checklist has a line that says "add schema." It's usually the most concrete item on the list, which is why it gets done first and trusted most. Somebody installs a plugin, a validator turns green, and the box is ticked.
Months later the brand still isn't in the answers, and nobody can say whether the markup did anything. That's the normal outcome when schema is treated as a ranking lever. It isn't one. It's closer to a notarised copy of what your page says: useful when it matches, worthless when nobody reads it, and harmful when it disagrees with the original.
What schema markup is, in one paragraph
Schema markup is structured data written in the schema.org vocabulary, a shared dictionary of types (Organization, Service, Article, FAQPage) and the properties each type can carry. On most sites it lives in a <script type="application/ld+json"> block in the page head, in a format called JSON-LD. Visitors never see it. Machines read it to learn, without guessing, what a page is about and what the thing it describes actually is.
The words "without guessing" matter most in that definition. Prose has to be interpreted. Markup mostly doesn't. Everything schema does for AI visibility follows from that, and so do all of its limits.
The myth: markup in, citation out
The pitch goes like this: AI engines love structured data, so add more of it and you'll get cited more. Three facts undercut it.
- The search platforms don't describe it as a switch. Google's guidance for its AI features says you don't need any special markup to appear in them, and that structured data should match what's visible on the page. Microsoft's Bing team has said publicly that schema helps its models understand content. Understanding is a real benefit, but it isn't a promise of being chosen.
- Rich results and AI citations are different things. Much of the "schema works" evidence comes from rich results: stars, FAQ dropdowns, how-to steps in classic search. Google cut FAQ rich results back to a narrow set of authoritative government and health sites in 2023 and retired how-to rich results the same year. Plenty of sites kept the markup and lost the display, and neither change says anything about whether an AI answer quotes you.
- Not every pipeline reads it. When an assistant fetches a page live to answer a question, many retrieval setups reduce the page to its visible text first, and script blocks don't survive that. The search indexes behind many engines do parse structured data. The model writing the answer may never see your JSON-LD at all. It sees whatever your page says in words.
The three jobs schema actually does
Drop the myth and there's still a strong case for structured data. It does three jobs well, and all three are about removing doubt:
Identity
It says which organisation you are, and ties your site to the profiles and knowledge-base entries that describe the same entity. Engines resolving "which Acme do they mean?" lean on exactly this kind of signal. It's the single biggest reason to have schema, and it's the job most sites do worst.
Liftable facts
Prices, hours, locations, service areas, founding dates. These are the facts an assistant most often gets wrong, and markup states them unambiguously. Paired with the same fact in plain text, it gives an engine two agreeing sources on one page.
Agreement
When the markup and the prose say the same thing, each confirms the other. The page reads as settled rather than contested. That consistency is what makes a fact look established, on a page and across the web.
Notice what's not on that list: persuasion. Schema can't make you sound more authoritative or more relevant than you are. Its job is to make sure what you actually are gets read correctly.
Three ways schema goes wrong, and why validators miss them
Here's the part checklists skip. Bad markup doesn't just fail to help. It can work against you, because the one job structured data has is to be trustworthy, and wrong markup is untrustworthy in a form built to be taken at face value.
- It contradicts the page. The markup says one price and the pricing table says another. The address block is from before the move. The
Organizationdescription is from two rebrands ago. A plugin set it once and nobody has looked since. An engine now has two versions of the same fact from the same source, which is worse than having one. - It's valid JSON with the wrong vocabulary. Every schema.org property belongs to specific types. Put a property on a type that doesn't define it (opening hours on a plain
Organizationinstead of a local-business type, say) and the block still parses and still looks fine. Consumers that follow the vocabulary just ignore it. The fact you meant to state never gets stated. - It claims what you can't back. Profile links to accounts you don't own, ratings the page doesn't show, services you no longer offer. Search platforms have long treated markup that isn't reflected in visible content as a quality problem, and self-serving review markup is the classic example. For identity, the damage is worse: link your entity to the wrong profile and you've told engines that someone else's record is yours.
position on each Service. But position is only defined on the list-item wrapper, not on the thing inside it. The JSON was valid and the page rendered fine, and the schema.org validator warned once per service. We fixed the structure, then taught our own page audit to catch that class of mistake on every site it scans, so nobody else's markup can fail that way without being told.Most of the free tools people rely on stop at syntax, which is why this sticks around. Here's the gap:
| Question | A syntax validator | What an engine needs |
|---|---|---|
| Does it parse? | Yes, this is the whole check | Table stakes. Unparseable markup is simply absent. |
| Is the vocabulary right? | Sometimes, as a warning that's easy to miss | Required. A property on the wrong type states nothing. |
| Does it match the page? | Never. It can't read your prose. | The point. Disagreement is worse than silence. |
| Is it true? | Never | Assumed, which is why a false claim costs you. |
| A green validator means your markup can be read. It doesn't tell you whether it's right, and "right" is the only kind of markup that helps. | ||
Which markup is worth having
The instinct after reading the schema.org type list is to add everything that applies. Resist it. More types means more surface to fall out of date, and none of them is worth anything if it's wrong. Useful markup comes in three layers:
- One identity block, in one home. A single, complete description of your organisation: the name you actually use, what you are, and the profiles that are verifiably yours. It should be defined once and referenced everywhere else. Don't restate it slightly differently on every page, because five near-identical organisations are exactly the ambiguity you were trying to remove.
- The thing you sell. Your services or products, and the facts buyers ask assistants about: what it is, who it's for, where it's offered, what it costs if you publish that. These are the facts engines paraphrase, so these are the ones most worth stating twice.
- The page itself. What kind of page this is, what it's about, when it last changed, and, when the page genuinely has them, the questions it answers. Page-level markup is where most drift happens, so it's where "matches the visible content" matters most.
We don't publish the exact property set we build, or how we decide which profiles count as verifiably yours. That's the part a look-alike generator would copy and fill with guesses, and a confident guess in markup is how brands end up linked to the wrong company. The principle holds without it: state less, but only what you can prove. Every property should be something you'd defend if a buyer read it aloud.
How to tell if it's working
This is the uncomfortable truth about structured data: you can't watch it work directly. No engine reports "we used your JSON-LD." What you can observe is the outcome it's meant to produce, which is engines identifying you correctly and stating your facts accurately. That's what to measure.
So the real test isn't whether your markup validates. It's four questions, and the last one has to be asked of more than one engine:
- Does it describe what's on the page? Every fact in the markup should appear in the visible content, with the same value.
- Is the vocabulary right? Every property should sit on a type that defines it, or it's decoration.
- Is the identity unambiguous? One organisation, described once, linked only to profiles you can prove are yours.
- Do engines get the facts right? Ask them. If an assistant still quotes last year's price, the markup isn't winning, whatever the validator says.
The first three are checks on your site. The fourth is a check on the world, and it's the only one that tells you anything changed. Keep unknown separate from clean, too: an engine that didn't answer hasn't confirmed your facts. It hasn't told you anything.
Markup is a promise about your page. Engines reward the promise only when the page keeps it.
That loop is what AI Access runs. It scans your pages the way an AI crawler does, flags markup that's missing, mistyped, or out of step with the page, and builds a corrected identity block from facts you've verified rather than facts a model guessed. Hallucination Watch then asks the engines whether they're getting you right, which is the check a validator can't do.
Key takeaways
- Schema isn't a citation switch. No engine recommends you because of a JSON-LD block. Rich results and AI citations are different things, and some retrieval pipelines never see your markup at all.
- It removes doubt. Structured data pins down who you are, states your key facts without ambiguity, and confirms what the page says. That makes you easier to identify and harder to misdescribe.
- Wrong markup is worse than none. Markup that contradicts the page, uses a property on the wrong type, or claims what you can't back gives engines a reason to doubt you, and a syntax validator catches almost none of it.
- State less, but only what you can prove. One identity block in one home, the thing you sell, and the page itself. Every property should be something you'd defend out loud.
- Measure the outcome, not the markup. The real test is whether engines get your facts right. The free audit checks your structured data, and AI Pulse keeps checking it.
Questions about schema and AI search.
Does schema markup help you show up in ChatGPT and other AI answers?
Indirectly, yes. Directly, no. Schema markup isn't a ranking signal that earns a citation on its own. What it does is remove ambiguity: it pins down which organisation you are, states key facts like prices, hours and locations in a form that doesn't need interpreting, and confirms what your page already says. That makes you easier for an engine to identify and harder to misdescribe, which is a real advantage when an assistant is deciding what to say about you. It works only when the markup is accurate and matches the visible page. Markup that disagrees with the page does more harm than good.
Do large language models actually read JSON-LD?
It depends on the pipeline, which is why you shouldn't rely on markup alone. The search indexes many engines draw on do parse structured data, and Microsoft's Bing team has said publicly that schema helps its models understand content. But when an assistant fetches a page live to answer a question, many setups reduce the page to its visible text first, and script blocks don't survive that. The safe assumption is that some engines see your markup and some only see your words, so every fact in the markup should also be stated plainly on the page.
Is FAQ schema still worth adding?
Only where the page really is a set of questions and answers, and not for the display. Google cut FAQ rich results back to a narrow set of authoritative government and health sites in 2023, so most sites won't get the dropdowns in classic search whatever they mark up. The markup can still accurately describe a page that answers common questions, and that's a legitimate use. The rule is the same as for any type: mark up what's visibly there, with the same wording, and don't add FAQ markup to a page that doesn't have an FAQ.
Can bad schema markup hurt your AI visibility?
Yes, in three ways. Markup that contradicts the page gives engines two versions of the same fact from the same source. Markup that puts a property on a type that doesn't define it is valid JSON that consumers ignore, so the fact you meant to state never gets stated. And markup that claims what you can't back, like profile links to accounts you don't own or ratings the page doesn't show, is treated as a quality problem, and for identity it can tie you to someone else's record. None of these are caught by a validator that only checks syntax.
Which schema types matter most for AI visibility?
Start with identity: one complete description of your organisation, defined once and referenced everywhere, linked only to profiles you can prove are yours. Most wrong AI answers about a brand trace back to engines being unsure who the brand is, and that block does more to fix it than anything else. After that, mark up the thing you sell (services or products and the facts buyers ask about), then the page itself. Resist adding every type that could apply. More markup means more surface to go stale, and wrong markup is worse than none.
How do you know if your schema markup is working?
You can't watch it work directly, because no engine reports which markup it used. What you can check is whether the markup is right and whether it's having the effect it's for. Every fact in it should appear on the page with the same value, every property should sit on a type that defines it, and your identity should be unambiguous. Then ask the engines: if an assistant still states last year's price or confuses you with another company, the markup isn't winning, whatever a validator says. Check on a cadence, because pages change and markup drifts out of step with them.
Related posts & guides.
Find out what your markup is really saying.
Enter your domain and the free audit shows how AI-ready your site is, structured data included, with no call, no setup and nothing to install. Then let AI Pulse keep it checked: the complete platform across all 10 AI engines, a single GEO score, what changed, and a ranked list of fixes written in plain language. $99/month, and a stale or broken block turns up in your report instead of in a wrong answer.