The authority network is the machine-readable layer that lets a search engine or language model understand what a page is about and connect it to a known entity: valid JSON-LD structured data, Organization and Person nodes with sameAs links to Wikidata, LinkedIn and other corroborating profiles, consistent name and address data, and off-site mentions that agree with the on-site story. Google's Knowledge Graph holds about 54 billion entities, and a brand that is not one of them is invisible to entity-based retrieval. This hub links the Entity Authority framework, the Schema Scanner, the Knowledge Graph gap analysis and the pillar-four audit checks.
Before a machine can cite you it has to know who you are and trust it has the right entity. This pillar is the machine-readable layer: schema validity, entity and Knowledge Graph signals, consistent sameAs references and off-site corroboration. Third-largest commercial-intent cluster, led by schema markup and knowledge graph demand.
Which frameworks explain schema markup and entity authority?
Which free tools and lab builds cover schema markup and entity authority?
AEO Schema Scanner
A free Chrome extension that reads any page’s structured data and shows three things: what is detected, what Google has deprecated or demoted, and what is missing to be the trusted answer in AI search. It also flags incomplete Organization and Person entities (missing sameAs, @id, logo, knowsAbout, jobTitle). The scan is 100% on-device. Open source.
Knowledge Graph Gap Analysis
A method for finding what is missing in a knowledge base by visualising it as a graph and identifying structural holes, dense clusters of related ideas that are not connected to each other. Bridging a structural hole tends to produce original, non-generic insight, because the territory between two developed clusters is by definition under-explored.
Which Google search systems bear on schema markup and entity authority?
Knowledge Graph
The Knowledge Graph is Google's database of facts about people, places, organisations and things, launched in May 2012. Google confirms it and says it held over 500 billion facts about five billion entities by 2020. It is infrastructure rather than a ranking system: it powers knowledge panels, answers factual questions and is named as a data source for AI Mode. For a site, the practical questions are whether its organisation is a recognised entity and whether the facts Google shows about it are correct.
Link analysis systems and PageRank
Link analysis systems are the Google ranking systems that read how pages link to each other to work out what pages are about and which are most useful for a query. PageRank is the best known of them. Google confirms both in its ranking systems guide and says PageRank has evolved a lot but remains part of its core ranking systems. For a site, this means crawlable links, descriptive anchor text and an internal link graph that points weight at the pages that matter are still worth auditing.
Penguin
Penguin was a Google system designed to combat link spam. Google announced it in April 2012 and integrated it into its core ranking systems in September 2016, when it became real-time and started devaluing spam links rather than demoting whole sites. Google now lists Penguin among retired systems. Its job continues through link spam handling in core ranking and SpamBrain, so sites should audit for manipulative links and over-optimised anchors, not for a Penguin penalty.
Reasonable surfer model
The reasonable surfer model is described in Google patent US7716225, filed in 2004 by Jeffrey Dean, Corin Anderson and Alexis Battle. It replaces PageRank's random surfer, who clicks any link with equal probability, with one who is more likely to follow prominent, relevant links than footer or advert links. Google has not confirmed it uses this model. The practical lesson is still testable: put links to important pages where readers will actually use them, not only in boilerplate.
Topic-sensitive PageRank
Topic-sensitive PageRank is a 2002 research method by Taher Haveliwala at Stanford that computes a separate PageRank score for each of a set of topics, then blends them by how well a query matches each topic. It is general information retrieval theory. Google acquired Haveliwala's startup Kaltix in 2003 and holds later personalisation patents by the same team, but has never confirmed using topic-sensitive PageRank. The useful takeaway for sites is to earn links from pages about your topic.
TrustRank
TrustRank is a link analysis method from a 2004 paper by Zoltán Gyöngyi and Hector Garcia-Molina of Stanford and Jan Pedersen of Yahoo. It starts from a small set of human-reviewed trustworthy seed pages and propagates trust through links, so pages far from the seeds are more likely to be spam. It is not a Google system. Google filed a trademark for the name for an anti-phishing filter, and Matt Cutts said Google has nothing specifically called trust rank. Proximity to reputable sites is still a sensible link audit.
Hilltop algorithm
Hilltop is a ranking method from Krishna Bharat and George Mihaila, developed at Compaq's Systems Research Center and published at WWW10 in 2001. It ranks pages by links from independent expert pages, meaning topic directories that link to many unaffiliated sources. Bharat later joined Google, but the Hilltop patent is held by HP's successors, not Google, and Google has never confirmed using Hilltop. The durable lesson is that links from independent, topical resource pages carry more meaning than links from related sites.
What has Laurelin published on schema markup and entity authority?
Which of the 387 audit checks belong to schema markup and entity authority?
66 checks in the Authority Network pillar of the 387-check AEO audit, each cited to a primary source. The most consulted first.
- Invalid JSON-LD syntax
- Unparseable / bad escaping
- Missing required properties
- Missing recommended properties
- Wrong @type
- Markup for invisible content
- Structured data on noindex page
- Inconsistent vocabularies
- Microdata/RDFa conflicting with JSON-LD
- Unstable/duplicate @id
- Deprecated types/properties
- Schema stuffing
- Conflicting duplicate entities
- No Organization schema
- No WebSite schema/SearchAction
- No Article markup on editorial
- Article missing author/date/headline
- Product missing price/availability/currency
- Product missing/invalid review/rating
- Offer missing priceValidUntil
- Missing BreadcrumbList
- FAQ/QA markup misused
- No LocalBusiness for a local business
- LocalBusiness missing geo/hours/address
Frequently asked questions about schema markup and entity authority
What is entity authority?
The degree to which a brand, product or person exists as a distinct, verified entity in the knowledge graphs that search engines and AI models draw on. It is built by consistent structured data, resolvable sameAs identifiers, a Wikidata item where the notability rules allow one, and third-party mentions that corroborate the same facts.
Which schema types matter most for AI search?
Organization and Person with sameAs (entity resolution), Article or TechArticle with author, dates and about (provenance), FAQPage and speakable (passage extraction), and BreadcrumbList (site hierarchy). Types the page is not genuinely eligible for add nothing and can trigger a manual action.