BlogPlaybook

Structured data for AI search: what schema.org actually does

Benjamin Libor10 min read

Structured data is a block of JSON-LD on a page that says what the page is: an article by this author on this date, a product with this price, an organisation with these profiles. It reliably helps search engines understand and display pages, and assistants that search inherit that. No assistant has documented that schema alone earns a citation, and it does nothing for a page with thin content. Add the handful of types that describe real things on your pages, validate them, and spend the rest of the time on the content.

In short

  • Structured data (schema.org vocabulary in JSON-LD) describes the page to machines; it reliably helps Google's rich results and entity understanding, which assistants that search inherit.
  • No assistant has documented that structured data by itself earns citations, and it does not fix thin content or missing facts.
  • For a B2B tech company, six types cover nearly everything: Organization, SoftwareApplication or Product with offers, Article, FAQPage, BreadcrumbList and HowTo where there is a real procedure.
  • Only mark up what is visible on the page; FAQ spam and invented ratings do harm.
  • Validate with the schema.org validator and Google's Rich Results Test, and put this after access, content and links in the order of priorities.

What structured data is

Structured data is a machine-readable description of a page, placed in the page's HTML. The vocabulary almost everyone uses is schema.org, a shared list of types (Organization, Article, Product, FAQPage) and their properties (name, author, datePublished, price). The format almost everyone uses is JSON-LD: a small block of JSON inside a <script type="application/ld+json"> tag, usually in the page head.

The block does not change what a visitor sees. It restates, in a fixed shape, what the page already says in prose: this is an article, written by this person, published on this date, about this product, which costs this much. A crawler that reads it does not have to guess which of the dates on the page is the publication date or which string is the author's name.

Everything structured data can do follows from "a crawler can read facts about the page without parsing the prose", and everything it cannot do follows from "it only restates what is already there".

What it reliably does

Two effects are documented and dependable.

Rich results in Google. Google's search documentation lists the structured data types that can change how a result looks: article dates and images, product prices and availability, breadcrumbs, organisation logos, and so on. When the markup is valid and the content qualifies, the result on the page shows more, and a result that shows more tends to get more clicks. This is the original reason structured data exists, and it still works.

Clearer entities. Structured data tells search engines that "Fernwick" on this page is a software product made by a company called Fernwick with these social profiles, not a surname or a place. Organization with sameAs links to your LinkedIn page, Crunchbase entry and GitHub organisation connects the name on your site to the same name elsewhere. The better a search engine's picture of the entity, the more confidently it can answer "what is Fernwick" and "who makes Fernwick".

Both effects reach assistants indirectly. ChatGPT, Perplexity, Gemini and Google's AI Overviews search before they answer, and they draw candidate pages from search indexes that were built with structured data in them. A page that a search engine understands well is a page that is found for the right queries, which is the first condition for being cited. How ChatGPT decides which pages to cite explains the search step.

What it doesn't do

Three things people hope for and will not get.

It does not earn citations by itself. At the time of writing, none of OpenAI, Perplexity, Google or Anthropic documents structured data as a factor in choosing which pages an assistant cites. The crawlers fetch HTML and read text; the JSON-LD block is there in the HTML, and may well be read, but there is no published evidence that it changes the choice. Treat any claim that "schema is the key to AI search" as a sales pitch until someone shows the documentation.

It does not fix thin content. A page with no facts, marked up as Article with a date and an author, is still a page with no facts. Assistants quote passages; the markup contains no passages. If the problem is that the page does not say anything quotable, the fix is writing content that AI answers quote, not markup.

It does not override what the page says. A price in the JSON-LD that disagrees with the price in the pricing table is a contradiction, not a correction. Search engines' guidelines say the markup must match the visible content, and markup that does not can get rich results removed. At best the wrong number is ignored; at worst the page loses trust.

The honest framing: structured data is hygiene. It is worth an afternoon, not a quarter.

Which types are worth adding for a B2B tech company

Most sites need six types. Add them in this order; each one describes something that is really on the page.

TypeWhereWhat to includeWhy
OrganizationHome page (and the site template)name, url, logo, sameAs to LinkedIn, Crunchbase, GitHub, XTies the brand name to one entity across the web
SoftwareApplication or ProductProduct and pricing pagesname, description, applicationCategory, operatingSystem, offers with price and priceCurrencyStates what the product is and what it costs in a form crawlers read
ArticleBlog posts, guides, docsheadline, author (a Person with a name and a url), datePublished, dateModified, publisherDates and authorship, unambiguous
FAQPagePages that show a real FAQEach visible question and its visible answerGives search engines the question-answer pairs you already show
BreadcrumbListEvery page below the home pageThe path from the home page to this oneShows where the page sits in the site; renders as breadcrumbs in Google
HowToPages that are a real procedureThe steps, in order, each with a name and textMarks a procedure as a procedure; only where the page is one

Two notes on the table. SoftwareApplication is the right type for a SaaS product; Product is for things with SKUs, though it is also accepted for software and some teams use both. Pick one and use it consistently. And FAQPage and HowTo have lost most of their visible effect in Google: in 2023 Google announced it would show FAQ rich results only for a small set of authoritative government and health sites, and would stop showing HowTo rich results. The markup is still valid, still describes the page accurately, and still costs nothing when the content is real. It is no longer a reason to add an FAQ to a page that does not need one.

Skip Review and AggregateRating unless you show real, collected ratings on the page, and Event and VideoObject unless the page is one.

What to avoid

Three mistakes do active harm.

  • Marking up content that is not on the page. An FAQ in the JSON-LD with no FAQ in the HTML, a rating with no reviews shown, an author name that appears nowhere. Google's guidelines call this out, it can cost the page its rich results, and if an assistant does read the block, it sees a page that says one thing and claims another.
  • FAQ spam. Ten questions bolted onto every page because an FAQ "helps AI". It does not; it dilutes the page, and the rich result it was meant to buy has largely gone. An FAQ belongs where readers have those questions, three to five of them, with real answers.
  • Copying a competitor's markup. Their sameAs links, their category, their offers. Every value must describe your page; a template is fine, the values are not.

A short example

This is the kind of block a pricing page for Fernwick, a made-up scheduling tool, would carry. It describes the product, its starting price and the company behind it, and nothing the page does not show.

{
  "@context": "https://schema.org",
  "@type": "SoftwareApplication",
  "name": "Fernwick",
  "applicationCategory": "BusinessApplication",
  "operatingSystem": "Web",
  "url": "https://fernwick.example",
  "description": "Scheduling tool for small sales teams and agencies.",
  "offers": {
    "@type": "Offer",
    "price": "8",
    "priceCurrency": "EUR",
    "url": "https://fernwick.example/pricing"
  },
  "publisher": {
    "@type": "Organization",
    "name": "Fernwick",
    "url": "https://fernwick.example",
    "sameAs": [
      "https://www.linkedin.com/company/fernwick-example",
      "https://github.com/fernwick-example"
    ]
  }
}

Note what it does and does not contain. The price is a number with a currency, matching the "from 8 euros per user per month" on the page. There is no rating, because the page shows none. The sameAs links go to profiles the company actually maintains. A blog post on the same site would carry an Article block instead, with the author and both dates, and both pages would share the Organization through the site template.

How to validate it

Validation takes minutes and catches most of the problems.

  1. Schema Markup Validator at validator.schema.org. Paste the URL or the code; it checks the markup is valid schema.org and shows what it parsed. Use this first, for every type.
  2. Google's Rich Results Test at search.google.com/test/rich-results. It checks the subset of types that can produce rich results in Google and tells you which the page qualifies for, with warnings for missing recommended fields.
  3. Search Console. After a few weeks, the Enhancements reports list the pages Google found markup on and any errors, across the whole site at once. This is where you catch a template change that broke every page.
  4. A view of the source. Confirm the block is in the HTML that a crawler receives, not injected by a script after load. If your site is a JavaScript application, check the server-rendered or static output, not the browser's inspector.

Re-validate whenever the template, a price or a plan name changes.

Where it sits in the order of priorities

For AI search, structured data comes fourth. The order most small teams should work in:

  1. Access. The AI crawlers can fetch the pages: robots.txt allows them, no bot challenge, the text is in the HTML. AI crawlers explained covers which agents to allow.
  2. Content. Each important page answers one question with facts, a date and an author, in passages that stand alone.
  3. Links. The pages are linked from your own cited pages and from the third-party pages assistants already cite.
  4. Structured data. The six types above, valid, matching the page.

An afternoon on step four is well spent, once. A week on it, while step two is undone, is the most common way a team convinces itself it has "done AI search" without changing what any assistant says. If you want a quick check of where you stand on all four, the AI score at /score reads a site's robots.txt, rendering, dates and structured data and reports them together.

Questions people ask

Does structured data help with ChatGPT or Perplexity citations?

Not directly, as far as anyone has documented. It helps the search engines that assistants rely on to understand and surface your pages, which is a precondition for being cited, but no assistant lists schema markup as a factor in choosing sources. Content, access and links are what move citations.

Which schema types should a SaaS company use?

Organization with sameAs links on the home page, SoftwareApplication with offers on product and pricing pages, Article with author and dates on blog posts, BreadcrumbList on every inner page, and FAQPage or HowTo only where the page really shows an FAQ or a procedure.

Should I add FAQ schema to every page?

No. Add it only where the page shows a real FAQ that readers need. Google has limited FAQ rich results to a small set of sites since 2023, so the visible benefit is largely gone, and an FAQ bolted onto every page dilutes the content and risks being treated as spam.

How do I check my structured data is correct?

Run the page through the Schema Markup Validator at validator.schema.org and Google's Rich Results Test, then watch the Enhancements reports in Search Console for site-wide errors. Also check the block is present in the HTML a crawler receives, not only in the browser after scripts run.

Written by Benjamin Libor, founder, echo.

Something wrong or out of date? Write to hello@echo-aeo.com and we'll fix it.

See what AI says about your brand. Free, in about a minute.

The free AI score checks how ready your website is for AI assistants. Echo itself tracks your customers' questions across the assistants every day and turns the answers into actions.