All blogs

AI marketing

llms.txt, Schema and Entity SEO: What Actually Works in the Technical Layer of AI Search

9 min readAugust 20, 2026
inX
llms.txt, Schema and Entity SEO: What Actually Works in the Technical Layer of AI Search

Explore what llms.txt, schema markup, and entity SEO actually do in AI search. Learn how these technical signals help AI systems understand your website, content, entities, and relationships—and where they fit within a practical GEO and AEO strategy.

Here is the number that should settle the llms.txt debate.

A study of 300,000 domains found roughly 10% had an llms.txt file. Among the fifty domains most frequently cited by AI engines, exactly one did.

The sites winning AI citations are, almost without exception, not doing the thing the industry is currently selling as the way to win AI citations.

So let's go through the technical layer properly: what does nothing, what quietly does a lot, and what order to do it in.

llms.txt: the honest verdict

The idea is sound. A Markdown file at your root giving AI systems a curated map of your important pages. Clean, simple, sensible.

The problem is that nobody is reading it.

Log analysis covering over 500 million AI bot events across a 90-day window found only a few hundred requests targeting the file directly. A separate study of 137,000 sites found the overwhelming majority of llms.txt files are never read at all. GPTBot, ClaudeBot, PerplexityBot and the rest crawl your HTML like everyone else.

Google has been explicit. Its June 2026 documentation states the file has no effect, positive or negative, on Search rankings or AI Overviews, and that Search simply ignores it. Gary Illyes confirmed Google does not support it and is not planning to. John Mueller compared it to the long-discredited keywords meta tag.

Mueller also gave the structural reason, which is the part worth understanding. A self-reported manifest cannot differentiate between sites, because every site would use it to declare itself the best one. Systems that rely on external consensus cannot take your word for what matters on your own domain. That is not a bug in the adoption curve. It is the reason adoption is not coming.

No major provider, OpenAI, Anthropic, Google, Meta or Mistral, has committed to using it in production.

So should you build one?

There is one real use case, and it is not marketing.

AI coding agents fetch it. Cursor, Copilot, Claude Code and similar tools use llms.txt to pull the right documentation without wasting tokens. Ahrefs found AI agents and the infrastructure serving them drive more llms.txt requests than any other bot type. Google added an llms.txt check to Lighthouse in May 2026 under a new agentic browsing category, alongside WebMCP.

So the rule is simple.

If you have developer documentation, an API, or technical resources that agents retrieve, build one. It genuinely helps, and that is what the format was designed for.

If you are a B2B services business hoping it improves your ChatGPT visibility, skip it. It costs twenty minutes and there is no penalty, so build it if it amuses you. Just do not let anyone bill you for it as a GEO strategy, and do not let it displace work that moves the number.

Schema: works, but not for the reason people think

Schema does help. Pages with structured data get cited more often. But the mechanism is not that markup is a ranking signal for AI answers, because Google is clear that no special markup is required for AI Overviews or AI Mode.

The mechanism is ambiguity removal.

When a system assembles an answer from a dozen sources, it has to decide what your page is, who published it, when, and what entity it describes. Schema answers those questions without inference. Pages that require less inference get used more.

Which reframes what good implementation looks like. The goal is not maximum coverage. It is accurate.

Organization is the one most companies get wrong and the one that matters most. Name, URL, logo, description, social profiles, contact points, address, founders. This single block does entity work across every other page you publish. Get it right once and it pays forever.

  • Person for your authors, connected to their content. If you are publishing expert commentary anonymously, you are throwing away the credibility you paid for.
  • Article for editorial content, with real published and modified dates.
  • Service for what you actually sell, described in your standard service names rather than five variations.

BreadcrumbList and WebSite for structural clarity.

The rule that matters more than any of the above: your markup must match what is visible on the page. Schema that describes things a reader cannot see is a guidelines violation and a trust problem, not a clever tactic.

And do not let a plugin decide. Most auto-generated schema is either incomplete or describes the wrong entity entirely.

Entity SEO: where the actual leverage is

This is the part of the technical layer nobody sells, because it is not a deliverable. It is housekeeping. And it does more than llms.txt and schema combined.

An entity is a thing a system can identify: your company, your people, your services, your locations. Entity SEO is the work of making those identifiable and consistent everywhere they appear.

Here is the failure mode, and almost every company has it.

Your homepage says AI-first digital marketing agency. Your LinkedIn says AI-powered growth partner. A directory lists you as a digital transformation consultancy. Your service pages call the same offering GEO on one page, generative engine optimization on another, and AI search optimization on a third. Your founder's title differs across three profiles.

A human reconciles all of that in two seconds. A system trying to build a confident answer about your category has to decide whether these are one company, three companies, or an ambiguity to route around. It routes around.

The fix has no tooling and no shortcut. Pick the canonical version of every important fact and enforce it:

One company description. Same sentence, same category words, everywhere. One name per service. Choose it and never vary it. Variants are separate entities to a machine.
One title per person, across your site, LinkedIn, bylines and conference bios.
One set of business facts. Founded, locations, size, contact.

Then audit the places you do not control: directories, publications that have covered you, partner sites, review platforms. Contradictions out there are worse than gaps, because they actively undermine the entity you are trying to establish.

This is an afternoon of work that most agencies will never bill you for, and it outperforms every emerging file format currently being marketed.

The order that actually matters

If you do these in the wrong sequence you will waste money, which is what most GEO programs do.

  • First, accessibility: Robots.txt, AI crawler access, indexation, status codes, canonicals, JavaScript rendering. A page that cannot be reached is not a candidate for anything. Also check you have not opted out of Google's generative AI features, which became a real setting in June 2026.
  • Second, content structure: One clear H1, descriptive subheadings, the answer in the first hundred words, visible dates, named authors, defined terminology, evidence for claims.
  • Third, structured data:  Accurate, matched to visible content, Organization first.
  • Fourth, entity consistency: On your site, then across the wider web.
  • Fifth, machine-readable extras: llms.txt sits here, at the end, as a hedge rather than a strategy.

Anyone who sells you step five before steps one through four is selling you the easy part.

Five ways this goes wrong

  • Publishing llms.txt instead of fixing content: The file is easy. The pages are hard. The pages are what get cited.
  • Bulk schema from a plugin: Irrelevant or inaccurate markup adds confusion and can breach guidelines. Less, correct, beats more, approximate.
  • Ten isolated service pages: Ten URLs targeting ten keywords can say almost nothing about how your services relate. Systems need relationships, not repetition.
  • Ignoring what third parties say about you: Your website is one source. If authoritative external sources contradict it, that is a bigger problem than anything on-page.
  • Optimizing before you are crawlable: Excellent schema on a blocked page achieves precisely nothing.

A quarterly technical review

Every three months, check four things.

  • Access: Crawlable, indexed, AI crawlers permitted, generative AI features not opted out, sitemap current, no orphaned commercial pages.
  • Markup: Organization schema complete and accurate. Author schema on editorial content. Markup matches visible content. No plugin junk.
  • Entities: Company description, service names, people and titles consistent on site, on LinkedIn, in directories, in bylines. Contradictions logged and fixed.
  • Evidence: Claims supported by something verifiable. Dates visible. Authors named. Definitions present for the terms you want to own.

That is the whole technical layer. It is unglamorous, it is mostly maintenance, and it beats every shortcut currently being marketed as GEO.

Conclusion

The technical layer is not about predicting how each model works. That knowledge expires in months.

It is about being unambiguous. Reachable, structured, consistently described, and backed by evidence, so that any system trying to build an answer in your category has enough reliable information to include you without guessing.

Build that, and you are positioned for whatever the engines do next. Build an llms.txt file, and you have a file.

Want to know which of these five layers is actually costing you visibility? Book a free growth audit. Thirty minutes, and you leave with a written first move, hire us or not.

There is no evidence that it does. Google states it has no effect on rankings or AI Overviews, and no major AI provider has committed to using it.

Only if you have developer docs or an API that AI coding agents retrieve. That is the format's real use case. As a marketing tactic, it is a hedge, not a strategy.

Because a self-reported file cannot differentiate between sites. Every site would claim to be the most important one, so it carries no signal.

Yes, indirectly. It removes ambiguity about what your page and organization are. Google is clear that no special markup is required, but accurate schema still helps.

Organization. It establishes your entity across every other page and is the one most companies implement incompletely.

Making your company, people and services consistently identifiable everywhere they appear, on your site and across the wider web. It is the highest-value and least-sold part of the technical layer.

Crawlability and indexation, then content structure, then schema, then entity consistency. Machine-readable files come last.