On 16 July I rearranged five pages on this site without changing a single sentence on them, and registered it as a controlled test. The results land on 27 August. This post is the design, the numbers going in, and the limits I want on the record first.
The idea being tested comes from Koray Tugberk Gubur, writing in Search Engine Land about visual semantics. His argument is that search engines read structure and function, not only wording, and that this is the part of topical authority most practitioners have skipped. I did not want to restate that argument. I wanted to find out whether it survives contact with a real site.
What is visual semantics?
Visual semantics is a meaning model that segments and classifies a document by its layout and functional components, alongside the text. The claim underneath it is that search has moved some of its attention from wording to arrangement. Google's Quality Rater Guidelines already treat human effort and involvement as a top quality principle, and design effort counts as one aspect of that.
The consequence that matters is about segmentation. Splitting a document into passages is not purely a language operation, because some process must first work out which pixels and tags form one coherent block. That process reads markup. Entities, predicates and triples all sit downstream of it, which is why a tidy semantic network can still fail to register. The block that carries most of this weight has a name.
What is a centerpiece annotation?
A centerpiece annotation is the primary visual block that expresses a page's purpose, function and context. Martin Splitt has publicly described Google identifying and annotating the primary content of a page. Exhibits from the Department of Justice case show a comparable mechanism applied to news documents, working within a window of a few hundred characters.
Two operating rules follow, and both live in a template rather than a style guide:
- Lead with the primary content or the working component. The area above the fold carries the macro context, and the rater guidelines weight main content hardest.
- Do not thread furniture through it. Navigation, social widgets and calls to action sitting inside the primary block break the extraction, and the documented failure cases are exactly that: boilerplate landing where the content should be.
The case attached to this idea involves a converter site of more than 100,000 pages, where moving a calculator from the foot of the page to the top produced the largest single gain of nineteen changes: clicks up 30.5%, impressions up 98.6%, average position from 8.9 to 8.5. I want to be blunt about the evidence grade. That is one practitioner's self-reported account with no independent audit, which is the reason I built my own version rather than quoting the figure and moving on.
What am I actually testing?
I am testing whether promoting the working component above the fold moves performance, across five of my own comparison pages, against an untouched control. Codename Visual Topical Layer, live since 16 July 2026, readout 27 August, registered as SEOtesting group tests 27551 and 27552.
The treatment is small on purpose, because a modest change with clean attribution beats an ambitious one you cannot read afterwards. Two edits per page:
- An "At a glance" verdict box placed after the opening and before the first H2, above the fold, with nothing cutting through it.
- The existing comparison table lifted to sit directly underneath, so the comparison occupies the centerpiece slot instead of trailing the prose.
No fresh copy. No schema edits. No new links. The verdict box restates an answer that already appeared further down, so what is being measured is position and prominence rather than added information.
Here is the state of those five pages beforehand. The figure counts characters of prose sitting between the top of the article and the comparison table, which is the element that actually settles the question the page exists to answer:
| Page | Characters of prose before the table |
|---|---|
| ai-seo-consultant-vs-traditional-seo-consultant | 9,216 |
| ai-seo-consultant-vs-agency | 5,928 |
| anthropic-mcp-vs-webmcp | 3,004 |
| head-terms-vs-long-tail-keywords | 1,843 |
| claude-vs-chatgpt-for-seo | 1,225 |
Over nine thousand characters ahead of the answer, on a page whose only job is to answer a comparison. I wrote that page. Nobody forced the decision.
The control is two comparison pages where the table already sat high, at 803 and 619 characters of preceding prose, left entirely alone.
Why did the cohort drop from eight pages to five?
Three of the eight candidates already belonged to other live tests, so including them would have spoiled two readouts at once. Those three were sitting inside the Mushroom method cohort or the google-vs-bing title test.
A page can carry one live test at a time. Apply a second treatment and neither result can be attributed cleanly afterwards. So the usable pool fell from eight to five, and a smaller honest sample beats a padded one I would then have to trust.
This part rarely appears in published case studies, which is roughly where many of them quietly come apart. It also caps what my own result can claim.
Why does retrieval cost decide whether you get evaluated?
Retrieval cost governs evaluation because Google weighs the expense of ranking a document against the expense of leaving it alone. A page that fails to identify itself cheaply, particularly at the centerpiece, costs more to fetch, render and assess, and gets dropped instead of debated.
Two public data points support it. The HTML size ceiling was reduced to 2MB, with widescale deindexing following the December 2025 core update. Separately, Pandu Nayak testified in the DOJ proceedings that the costly ranking systems are held back for results already showing at least one click, solid topicality, and annotations that warrant the expense.
The practical version: inexpensive structural checks happen first. Clear those and the richer systems run. Fail them and the richer systems never see the page at all, so the content is bypassed rather than assessed. Among the earliest of those cheap judgements is what kind of page this is.
How does Google work out what type of page you are?
Google sorts a site into a category, then evaluates the content within that category, and layout is what drives the sorting. Affiliate, aggregator, service provider, ecommerce, software. The capability a page offers is expressed through its components, which is why identical wording performs differently depending on the category it lands in.
Misleading functionality is now written into Google's spam policies. Implying a comparison, filter, calculator, booking flow or review capability without delivering it is a stated risk. Most failures here are accidental rather than dishonest: a filter that materialises only after scripts execute, a comparison folded into a tab that never opens, review content sealed inside an embedded frame from another domain.
This is the area I feel reasonably comfortable about here, because the tools on this site genuinely run. The low quality content checker, the AI search revenue calculator and the agent friendliness audit all execute for real. Working functionality is a moat when the classifier is hunting for capability. The remaining question is whether that capability is visible where it counts, which is a question about responsiveness.
What separates relevance from responsiveness?
Relevance is entity match and responsiveness is task completion, and a page has to satisfy both. Neural matching connects the entity in a query to documents carrying that same entity, so a mismatch rarely competes. Responsiveness asks the harder question of whether the underlying job can be finished on the page.
A document can carry the correct entity, discuss the correct subject, and still leave the job untouched. Where the intent is to compare, filter, calculate, book or review, describing the activity is not the same as providing it.
There is also a constraint operating at the results-page level. Google caps how many results of a given type appear together, so a page can be relevant enough to qualify and still be squeezed out by the mix the results page wants. Losing a slot to page-type composition is a structural outcome, not a content one, and it argues for deciding page type deliberately instead of accepting whatever the CMS produces.
What does this change in a topical map?
A topical map now needs to name the page format and the working elements each entry must ship with, decided before any brief gets written. Entities, attributes and relationships stop short of the question of what a page has to be able to do.
The maps I have built carry topic, query, intent and priority. What they have lacked is a commitment recorded up front about which working elements the page must ship with, and a willingness to hold the page back when those elements are not going to exist. Where an intent does not warrant its own page, prune it. Where two pages perform the same job, merge them. A tighter set of well-shaped pages lowers retrieval cost and concentrates authority, which cuts against the reflex to mint a page for every phrasing.
That links straight to what I already publish on topical authority and on building a social topical map. The map gains a field. The underlying strategy holds.
What I expect to be wrong about
Five treated pages is a small sample with real limits, and I want them stated before the result arrives.
- The two edits shipped together, so a positive result identifies the treatment as a whole and cannot separate the verdict box from the table move.
- Comparison pages are the easiest possible subject because the working component is obvious. A guide with no natural centerpiece is a harder case and my findings will not transfer to it cleanly.
- An off-site authority ceiling sits over this whole site. Most of these pages rank well beyond the top 20, and arrangement does not repair an entity problem. If the treatment does anything, I expect impressions and AI citations to move ahead of clicks.
- A null result is entirely possible, and it gets published on 27 August either way.
FAQ
What is visual semantics in SEO?
Visual semantics is the way search engines segment, classify and interpret a page through its arrangement and working components rather than its wording alone. Structure operates as a retrieval and classification signal, so where a block sits affects how the page is understood.
What is a centerpiece annotation?
A centerpiece annotation is the primary visual block expressing what a page exists to do. Google extracts it from markup within a limited character window, and furniture cutting through that block corrupts what gets extracted.
Does page layout affect rankings?
Arrangement affects rankings indirectly but materially, because it drives classification and retrieval cost. Cheap structural checks determine whether a page earns expensive evaluation, and page category determines which results slot it competes for.
Is visual semantics a design job or a technical SEO job?
Visual semantics falls mostly to technical SEO, because almost every change it calls for happens in a template. Component order, above-the-fold composition and whether a function renders without scripts are engineering decisions.
How do I check my own centerpiece?
Turn scripts off, read the opening stretch of meaningful content in the markup, and judge whether it describes the page itself or the surroundings. Then confirm the page's core capability still exists in that same scripts-off state.
Sources
- Koray Tugberk Gubur, Search Engine Land, July 2026: https://searchengineland.com/visual-semantics-topical-authority-482254
- Google Search Quality Rater Guidelines (human effort and involvement as a quality principle)
- United States v. Google, Department of Justice trial exhibits and testimony (Pandu Nayak on ranking cost and result-type limits)
Soaring Above Search
Weekly AI search insights from the front line. One newsletter. Six sections. Everything that actually moved this week, with a practitioner's take.