HomeSEOWhy Data Integrity Is The New Technical SEO: From Crawling To Trust

Why Data Integrity Is The New Technical SEO: From Crawling To Trust

Up to now two years, Google has dropped help for 9 ItemTypes from its wealthy end result search gallery. This has occurred not too lengthy after ChatGPT’s launch and when mass adoption started:

Picture from creator, July 2026

The query stays whether or not this decline will proceed, however the newest removing – FAQ/FAQPage – has since prompted some debate over the function of schema.org inside the way forward for Search.

Schema Is Useless, Proper?

Whereas some carry out exams and experiments to know whether or not schema actually makes a constructive affect on being cited inside platform responses, Gianluca Fiorelli notably noticed that we could also be performing these exams on restricted datasets. With that in thoughts, let’s remind ourselves of the wording of the deprecation message for FAQ wealthy outcomes:

“…We will probably be dropping the FAQ search look, wealthy end result report, and help within the Wealthy outcomes check in June 2026.”

Discover right here what they didn’t point out – which is that the usage of FAQ schema is not required. It is because the deprecation is that of wealthy outcomes solely – a show function. Schema itself is a comprehension layer – figuring out entities and the relationships between them. Is schema useless? In my view, it’s removed from it. Whereas some properties deprecate, others, equivalent to Product, are prolonged.

That being stated, I’m additionally conscious that including schema isn’t a magic bullet that contributes in direction of quotation progress. Nevertheless, that progress goes past the metrics we’ve been accustomed to depend on equivalent to citations, impressions, and so on. Suganthan Mohanadasan wrote an important piece about how schema has three “lives”:

  1. Google’s index pipeline.
  2. LLM pretraining (oblique).
  3. LLM runtime retrieval.

SEOs have been traditionally centered on No. 1 as one thing that may positively contribute in direction of success metrics. However schema goes past what we’re used to, or precisely, report on. Schema isn’t dying; one show function it benefited from is diminishing as an alternative.

An search engine optimization’s Largest Risk: Ambiguity

Schema is an ontology that, as an online normal, can contribute in direction of knowledge integrity. The danger to knowledge integrity is ambiguity. Ambiguity results in hallucinations. Hallucinations snowball. Ultimately the end result compounds, which might result in inaccurate outcomes and even incorrect LLM pre-training which might have longer-lasting results.

If an agent can misinterpret you, sooner or later it would. LLMs can then danger touring into “semantic drift” detracted from the information and in favor of narrative. This was explored inside a bit titled “Sangue e Grafi: Instructing a Small Mannequin to Learn the Bloodline” by Andrea Volpini and Chiara Carrozza the place frontier fashions tended to fall for narrative over information, whereas a small mannequin given knowledge-graph instruments drew degree with them.

Sangue e Grafi, by WordLift
Picture from creator, July 2026

→ Additional studying: Data Retrieval Half 1: Disambiguation

5 Layers Of Knowledge Integrity

All this corroborates my perception that an search engine optimization’s function is to maximise knowledge integrity, of which schema performs a job. Beneath, I illustrate 5 layers of what knowledge integrity can embody:

5 layers of data integrity
Picture from creator, July 2026
  1. Entities: What exists, and what that factor is. Factor, Group, and Particular person, stabilized with @ids and tied out to Wikidata, GS1, ISNI, or ORCID so an agent is aware of your “Apple” from the fruit.
  2. Relationships: How these issues join. @id and sameAs, RDF. Yoast search engine optimization’s schema aggregation function and EntityMap by Dixon Jones.
  3. Format: How construction is serialized and served. JSON-LD, RDFa, and Microdata. Markdown, too (encompassing LLMs.txt, brokers.md, OKF) and endpoints (content material negotiation, ARD, MCP).
  4. Actions: What could be finished, declared to brokers. Schema.org Actions equivalent to BuyAction, plus the newer WebMCP, ACP, and UCP.
  5. Notion: Grounding, third-party notion, sentiment, and so on.

Aggregation, Steering, And Consumption

In a publish I wrote in October final yr, I stated, “SEOs should contemplate either side of the net and easy methods to serve each.” The rising protocols (all of which have been launched previously two years) present this to be true, the place a brand new “agentic grounding stack” typically adopts considered one of three targets:

Protocol  Purpose  What It Does 
sitemap.xml  Aggregation Each canonical URL right into a single XML index.
llms.txt  Aggregation Abstract of a website’s content material with essential info and hyperlinks to additional studying.
Yoast Schema Aggregation  Aggregation Web page-level JSON-LD into one linked site-wide graph.
EntityMap  Aggregation A website’s entity declarations into one express map.
Information Catalog  Aggregation Structured, unstructured, and SaaS knowledge right into a ruled context engine.
OKF  Aggregation Web site data right into a markdown bundle at /okf/.
ARD · ai-catalog.json  Aggregation A site’s instruments and brokers right into a catalog; registries federate above it.
OpenKB  Aggregation Supply paperwork compiled right into a markdown wiki.
Schema.org  Steering The shared vocabulary that tells machines what issues imply.
brokers.md  Steering How brokers ought to signify and work together with you.
Markdown for Brokers  Consumption Similar URL served as clear markdown by way of content material negotiation.
Markdown alternate output  Consumption A separate .md model linked with rel=alternate.
/crawl endpoint  Consumption Renders a web page, or whole website, as clear markdown on demand.
WebMCP  Consumption Exposes a website’s actions as instruments an agent can invoke.
NLWeb  Consumption Ingests schema, feeds , and sitemaps to reply natural-language queries.
ACP  Consumption Agent checkout inside ChatGPT towards service provider product knowledge.
UCP  Consumption A typical language for agent commerce actions throughout surfaces.

These three targets assist cut back the variety of requests whereas rising token effectivity. Among the above protocols have been lined in additional element inside Search Engine Journal, together with my very own on ACP and UCP and Emina Demiri-Watson’s thorough article on OKF, ARD, and others earlier this month.

However there’s one thing none of those protocols have…

There Is No Consensus Or Agreed Commonplace

Schema.org was born out of consensus between Google, Microsoft/Bing, and Yahoo! (Yandex becoming a member of later) who launched it below joint governance. The identical occurred 5 years earlier with the XML sitemap. When the major search engines wanted a normal, they merely sat down and created one – collectively.

Nothing like that is occurring now, and it comes at the detriment of SEOs who genuinely need readability on what to implement and what to not implement for websites they work on. Even fundamental information about consumption are contested, the place the debate over markdown is a nice instance of this.

Whereas these debates proceed, there’s no room the place platforms are convening and agreeing to at least one common normal. The ecosystem has modified so dramatically that these firms should not within the enterprise of Search and the nice of the net, however should now navigate how their companies have an effect on jobs, economies, livelihoods, and the way forward for humanity as a complete. As such, I simply don’t consider questions posed by SEOs are on the high of their priorities.

What Can You Do About It Now?

Wanting again on the 5 layers of information integrity, the 4 you could have management over could be illustrated under when taking a look at what an agentic grounding stack can appear like:

Agentic grounding stack options
Picture from creator, July 2026

There’s quite a bit to think about, and all have completely different targets and technical debt. Determine that are most relevant to you, in addition to adopting something that ought to not require an excessive amount of technical debt.

If I needed to decide an order, it might be this.

  • Stabilize your @ids and add sameAs hyperlinks out to Wikidata and the opposite authorities first, as a result of all the things else stands on it.
  • Then check how you’re truly interpreted, with NLWeb, fairly than assuming the graph reads the way in which you supposed.
  • If you’re in ecommerce, audit the product feed earlier than touching something shiny, and have a look at BuyAction whereas you’re there: solely ReadAction and SearchAction are deployed at any actual scale at the moment, so the sphere is genuinely open. Look into the latest information about what has been added to the Product schema.
  • Try and implement WebMCP. It may be finished on any web site, and doesn’t must be ecommerce.
  • Markdown serving and content material negotiation can wait till you could have engineering capability to spare (except you should utilize Cloudflare’s Markdown for Brokers).
  • Look into OKF and ARD. When Google launches new protocols, I at all times take discover – particularly with regards to how an agent or LLM understands a website as a complete.

None of this can be a wager on a selected protocol. Implementing any of those reduces the danger that an LLM has to go “the good distance spherical” to kind the reply it desires to reply with. By hedging bets to welcome any agent from any platform, this may even allow you to suppose extra about precisely how your website could also be interpreted by them, and the way this improves when these protocols are adopted.

Even if you wish to make small experiments away from bigger websites, it’s value exploring not solely to see if there are constructive outcomes from it, but additionally to know how all of them work in apply. That is precisely what I’ve finished lately, the place I’ve rebuilt my private web site, which has a number of “agentic prepared” protocols operating, together with content material negotiation, markdown alternate, llms.txt and WebMCP.

Don’t Chase The Protocol, Personal The Layers Beneath

Proper now, it appears there isn’t any single “winner” that can progress from a proposal or protocol to an online normal. There’s no consortium to repeat what was finished with the XML sitemap and Schema.

The stack is now massive, however there’s one factor all of them share. Whether or not it aggregates, guides, or consumes, each is fed by the identical substrate: Correct entities, express relationships, and content material a machine can learn with out guessing or researching additional.

Rankings have been the success metrics of the outdated net. Belief, integrity, accuracy, and validity. Incomes that is nonetheless the function of an search engine optimization.

Extra Assets:


Featured Picture: hmorena/Shutterstock

RELATED ARTICLES

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Most Popular