For the complete documentation index, see llms.txt. Every page on this site is also served as Markdown: append `.md` to any URL, or send `Accept: text/markdown`.

Structured data vs llms.txt: consumed versus proposed

One is a vocabulary four search engines document reading. The other is a proposal Google says no AI system currently uses. Compared on backing, format, consumers and what each actually earns you.

Last updated: 2026-09-03By Karl-Gustav Kallasmaa

Structured data

A standardised format for describing what a page is about and classifying its content, expressed in the schema.org vocabulary and most often embedded in the page as JSON-LD.

Checked 2026-09-03

llms.txt

A proposal to standardise on a markdown file that gives AI agents a curated list of a site's most useful pages, published as a separate document rather than embedded in the pages themselves.

Checked 2026-09-03

Which one should you choose?

One of these is a vocabulary with named backers and documented consumers; the other is a well-argued proposal whose consumption is, on the most specific public statement available, still hypothetical. That asymmetry decides the ordering, without saying anything unkind about the proposal.

Choose Structured data when

Invest in structured data when you have entities worth describing precisely - products, organisations, articles, FAQs, events - and when eligibility for enhanced display matters to you. It is read today by named systems, and it makes the meaning of a page explicit rather than inferred.

Choose llms.txt when

Publish an llms.txt when your platform generates it for free, when your documentation is large enough that a curated map is genuinely useful to a human as well, or when you want a cheap option on future adoption. The cost is minutes and the downside is a stale file.

When neither is the right answer

Neither is the answer if your pages do not contain clear, specific, extractable statements. Structured data describes content that already exists and llms.txt points at it; neither invents substance, and a precisely annotated vague page is still a vague page.

What is specific to this comparison

  • Only one side of this pairing has a named multi-vendor backing: schema.org states that its vocabularies are understood by the major search engines and names four of them, while the llms.txt proposal has a single author and no governing body.
  • The evidence runs in opposite directions on the two sides. Structured data has a documented consumer and no promised outcome; llms.txt has documented publishers, including three AI labs, and a Google statement in June 2026 that no AI system uses it.
  • They operate at incompatible granularities: structured data annotates entities and properties inside a page, while llms.txt lists whole pages, so one can describe a single product's availability and the other cannot describe anything smaller than a URL.
  • Google's language about structured data is unusually precise about limits, promising only eligibility for enhanced display, which makes it the rare web convention whose own documentation refuses to overstate what it earns you.

Structured data vs llms.txt, criterion by criterion

Definition
What it is
A standardised format for classifying page contentSource, checked 2026-09-03
Mechanics
Where it lives
Embedded in the page, usually as JSON-LDSource, checked 2026-09-03
Mechanics
Format
schema.org vocabulary, JSON-LD recommendedSource, checked 2026-09-03
Adoption
Named backers
YesGoogle, Microsoft, Yandex and YahooSource, checked 2026-09-03
Adoption
Published record of consumption
YesDocumented as read and acted on by GoogleSource, checked 2026-09-03
Adoption
Published by major AI labs for their own docs
Not documentedNot claimed either way by the vocabulary's own siteSource, checked 2026-09-03
Outcome
What correct implementation earns you
PartialEligibility for enhanced display, not a guaranteeSource, checked 2026-09-03
Outcome
Granularity of what it describes
Individual entities and properties within a pageSource, checked 2026-09-03

The comparison in one line

Structured data is a vocabulary that four search engines say they understand and that Google documents acting on. llms.txt is a proposal for a file that, on the most specific public statement available, no AI system currently reads. Both are cheap. Only one has consumers.

What each one actually is

Google defines structured data as a standardised format for providing information about a page and classifying the page content. The vocabulary is schema.org, which describes itself as a collection of shared vocabularies webmasters can use to mark up their pages in ways that can be understood by the major search engines, naming Google, Microsoft, Yandex and Yahoo. On format, Google is direct: it recommends JSON-LD where a site's setup allows it, calling it the easiest solution for website owners.

llms.txt describes itself as a proposal to standardise on using an /llms.txt file to provide information to help agents use a website. The specification is deliberately minimal: a markdown document whose only required section is an H1 with the name of the project or site, followed by a blockquote summary and optional H2-delimited lists of links. It lives at the site root or at any path within it, covering the pages beneath that path.

Note how differently the two describe themselves. One is a format with an articulated relationship to specific consuming systems. The other is a convention looking for consumers.

The evidence, in both directions

Structured data has a documented reader and a carefully limited promise. Google's own guidance says that all required properties must be present for an object to be eligible for appearance in Google Search with enhanced display, and warns that structured data breaking the general guidelines might be ineligible. The word "eligible" is doing deliberate work there. Google is not promising a rich result; it is promising that you are in the running for one.

That restraint is worth respecting rather than reading past. A convention that refuses to overstate its own payoff is unusual, and it means that anyone selling structured data as a guaranteed visibility gain is going beyond what the documentation supports.

llms.txt has the inverse evidence profile. Its adoption story is about publishers, not consumers: the proposal site states that OpenAI, Anthropic and Gemini publish llms.txt files for their own developer documentation, and that Mintlify, GitBook and Wix generate the file automatically. That is genuine momentum on the supply side.

On the demand side, the most specific public statement runs the other way. Search Engine Journal reported in June 2026 that Google's John Mueller called the file purely speculative and observed that it has existed for years without AI systems using it. That is one search engine's view, reported secondhand, and it is not a measurement of every assistant. But nobody has published a countervailing record of consumption, and the absence of one is itself informative for a convention that has been public since September 2024.

They are not even solving the same problem

The granularity difference makes these harder to compare than the surface similarity suggests, and it is the reason "which should I use" is usually the wrong question.

Structured data operates inside a page and below the level of the page. It can say that this specific thing is a product, that the product has this name, that the name is associated with this availability status, that this block is a question and that block is its answer. It attaches machine-readable meaning to fragments a reader would otherwise have to infer from layout.

llms.txt operates above the level of the page. Its unit is a link with a short description. It cannot say anything about what is inside a document; it can only say that the document exists and is worth reading. It is a table of contents, and a table of contents is a genuinely useful artefact for a large documentation set.

So the honest framing is not either-or. Structured data adds semantics to content an agent is already reading. llms.txt proposes a way for the agent to decide what to read. If the second worked reliably, the two would compose nicely.

What to do with your time

Spend the first hour on structured data for the entities you actually have. Organisation and product markup where you sell something, article markup where you publish, FAQ markup where you genuinely answer questions. Use JSON-LD because Google recommends it and because a script block is far easier to keep correct than attributes scattered through templates. Include every required property, because eligibility is conditional on completeness.

Then validate it, and validate it again after your next template change. Markup rots silently: a field renames, a component is refactored, and the JSON-LD keeps emitting with a missing required property that quietly removes eligibility. Nothing in your analytics will tell you.

Spend the second hour on the content itself, because both conventions describe substance rather than creating it. A precisely annotated page that says nothing specific is still a page that says nothing specific, and no retrieval system will prefer it.

Spend twenty minutes, at most, generating an llms.txt from the same source of truth as your documentation. Automate it so it cannot drift, point it at the pages you would want an agent to read first, and then stop thinking about it. If a major assistant announces tomorrow that it reads the file, you are already covered, and you will not have traded anything real to get there.

Why structured data still matters when the reader is a model

There is a popular argument that markup is obsolete because language models read prose perfectly well and do not need to be told that a number is a price. The argument is not stupid, and it is still worth resisting for three reasons that have nothing to do with rich results.

It removes ambiguity you did not know you had. A page with three numbers on it requires an inference about which one is the current price of the thing being asked about. Markup states it. Inference is where wrong answers come from, and reducing the number of inferences a system has to make is the cheapest accuracy work available.

It is machine-checkable, so it can be kept honest. Prose claims drift. A JSON-LD block can be validated in continuous integration against the page it describes, which means a stale price or a retired product becomes a build failure rather than a discovery six months later. Very little else about a marketing site has that property.

It forces you to have facts. Filling in required properties is a surprisingly effective audit. A team that cannot populate a product's fields without argument has learned something useful about how clearly it describes its own product, and that clarity is exactly what a retrieval system needs whether or not it ever parses the markup.

None of that requires believing any particular claim about how models consume structured data. It holds because the discipline improves the page.

What a good pairing looks like in practice

If you decide to do both, here is a shape that avoids the usual waste.

Generate the llms.txt from your documentation build, listing your reference pages, your changelog and your handful of genuinely authoritative explainers, each with a one-line description. Regenerate on every docs deploy. Never hand-edit it.

Emit JSON-LD from the same data your pages render from, so the markup and the visible content cannot disagree. Assert only what the page shows; markup that describes content a visitor cannot see is against Google's general guidelines and risks ineligibility rather than earning anything.

Keep both out of the way of the actual work, which is writing pages that contain specific, attributable, current statements. Both conventions are packaging. Packaging matters, and it does not substitute for having something in the box.

The trap to avoid

The trap is treating a new convention as evidence that the old one stopped mattering. The arrival of AI answers did not deprecate schema.org, and Google continues to document structured data as the way to classify page content for its systems. A team that skips markup because "AI does not need schema" is making an inference the documentation does not support, on behalf of systems that have not published what they use.

The symmetrical trap is dismissing llms.txt as vendor hype. It is not. It is a careful, minimal proposal from a credible author, and it costs almost nothing to adopt. It has simply not yet been shown to be read, and honest advice says so rather than selling it as a mechanism.

Rank them by evidence, implement both, and revisit the ranking when the evidence changes.

A reasonable review cadence is quarterly, and it takes about ten minutes. Re-read the llms.txt proposal to see whether the version has moved and whether any assistant vendor has published documentation describing itself as a consumer. Re-validate a sample of your structured data against the pages it annotates. Re-check whether Google's language about eligibility has changed. If any of those three answers shifts, the ordering on this page shifts with it, and the honest thing to do is change the advice rather than defend the previous version of it.

Questions people ask

Structured data, on the evidence. schema.org states its vocabularies are understood by the major search engines and names Google, Microsoft, Yandex and Yahoo, and Google documents reading it and acting on it. The llms.txt proposal has no equivalent record of consumption, and Google's John Mueller described it in June 2026 as purely speculative.

No, and Google is careful about the wording. Its documentation says including all the required properties makes an object eligible for enhanced display, and warns that structured data which breaks the general guidelines might be ineligible. Eligibility is the promise; appearance is not.

JSON-LD, unless something about your stack makes it impractical. Google says it recommends the format that is easiest to implement and maintain, which in most cases is JSON-LD, and that JSON-LD is the easiest solution for website owners because it sits in a script block rather than being woven through the markup.

Not useless, just unproven. It is cheap to generate, it does no harm, and the AI labs publish one for their own developer documentation. Treat it as a low-stake bet on future adoption rather than as a visibility mechanism, and do not let it displace work on things that are demonstrably read.

Barely. Structured data annotates the meaning of content inside a page for machines already reading that page. llms.txt is a separate markdown document listing links, intended to help an agent decide what to read at all. One adds semantics, the other proposes navigation.

Sources

Every claim on this page, with the page it came from and the date that page was read. Prices and feature lists change; these are what the source said on the date shown, not timeless facts.

  1. Google defines structured data as a standardised format for providing information about a page and classifying the page content.a standardized format for providing information about a page and classifying the page contenthttps://developers.google.com/search/docs/appearance/structured-data/intro-structured-data — read 2026-09-03
  2. Google recommends JSON-LD for structured data where a site's setup allows it, describing it as the easiest solution for website owners.Google recommends using JSON-LD for structured data if your site's setup allows it, as it's the easiest solution for website owners.https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data — read 2026-09-03
  3. Google frames structured data as conferring eligibility rather than a guarantee, stating that all required properties must be included for an object to be eligible for enhanced display.You must include all the required properties for an object to be eligible for appearance in Google Search with enhanced displayhttps://developers.google.com/search/docs/appearance/structured-data/intro-structured-data — read 2026-09-03
  4. schema.org describes itself as providing shared vocabularies that can be understood by the major search engines, and names Google, Microsoft, Yandex and Yahoo.Schema.org provides a collection of shared vocabularies webmasters can use to mark up their pages in ways that can be understood by the major search engines.https://schema.org/docs/gs.html — read 2026-09-03
  5. llms.txt is described by its own site as a proposal to standardise on a file providing information to help agents use a website.A proposal to standardise on using an /llms.txt file to provide information to help agents use a website.https://llmstxt.org/ — read 2026-09-03
  6. The proposal specifies a markdown document whose only required section is an H1 naming the project or site, followed by a summary blockquote and optional H2-delimited link lists.An H1 with the name of the project or site. This is the only required section.https://llmstxt.org/ — read 2026-09-03
  7. The proposal places the file at the site root or at any path within it, covering the pages beneath that path, with the most specific file taking precedence.at the site root, or at any path within it, covering the pages under that pathhttps://llmstxt.org/ — read 2026-09-03
  8. The proposal site states that OpenAI, Anthropic and Gemini publish llms.txt files for their own developer documentation, and that Mintlify, GitBook and Wix generate the file automatically.https://llmstxt.org/ — read 2026-09-03
  9. Search Engine Journal reported in June 2026 that Google's John Mueller called llms.txt purely speculative and observed that the file has existed for years without AI systems using it.https://www.searchenginejournal.com/google-says-llms-txt-is-purely-speculative-for-now/577576/ — read 2026-09-03