Rewriting your pages can make you less visible
The standard GEO service, rewriting page bodies to be more citable, has now been measured end to end. On the largest test so far it reduced visibility.
Published 2026-08-19 · last reviewed 2026-08-19 · 4 sources
The most common thing sold under the name GEO is a rewrite: take the pages you have, restructure the body so an answer engine is more likely to quote it, and wait. In July 2026 a survey of 45 studies collected the first measurement of that service end to end, from the web index through retrieval to the final citation, rather than inside a fixed test harness. Across 171,003 documents and 2,700 queries in the SAGEO Arena benchmark, body-only optimisation reduced average presence in the top twenty retrieved documents by about 9%, reduced presence in the top ten after reranking by 16%, and reduced final citation by 6%. Not a smaller gain than promised. A loss. If your agency's deliverable is a rewritten page body, that is the result you are buying, and the sections below explain why the number comes out negative.
Two probabilities, multiplied
A page is cited in an AI answer only if two things happen in order. First it has to be retrieved: the engine's search step has to pull it into the handful of documents the model reads. Second, given that it was retrieved, the model has to choose to quote it. Write the first as P(retrieved) and the second as P(cited given retrieved). The probability of the thing you care about, P(cited), is the product of the two. A rewrite that makes a page more quotable raises the second probability. But the same rewrite changes the words on the page, and the retrieval step is a search engine scoring those words against the query. Change the words in the service of the model and you can lower the score the search step gives you. The composed effect is the product of a number that went up and a number that went down, and nothing guarantees the product rises. On the largest end-to-end test, it fell.
Why most benchmarks missed it
Almost every GEO benchmark before 2026 measured only the second half. The method was to take a fixed set of documents, usually the top five Google results for a query, inject them into the model's context, and measure how much of the answer each document received before and after an edit. Retrieval was held constant by construction, so a rewrite could only ever look neutral or positive. The 252,000-trial factorial experiment by Vishwakarma and colleagues, across six models and eighteen factors, found relevance and position in the context to be the primary determinants of the first citation. That is consistent with the harness design: when the documents are already in front of the model, what matters is which one sits where. It says nothing about whether a rewritten page would have been in front of the model at all, which is the question a business is paying to have answered.
What the famous 40% actually measured
The figure most vendors quote comes from the original 2023 GEO paper and is usually stated as "up to 40% more visibility". Read the paper and the number describes something narrower. The metric is position-adjusted word count, which discounts passages that appear later in the answer. The test supplies the top five Google results for each query to GPT-3.5-turbo, then measures the share of attributed text each source receives. Adding a quotation to a source raised its position-adjusted word count from 19.3 to 27.2, a relative increase of about 41%. So the claim is this: a source that was already handed to the generator receives a larger position-weighted share of the answer text after editing. That is a real finding about answer composition. It establishes neither that the edited page gets retrieved from the open web, nor that anyone visits it. Quoting it as a visibility gain is quoting a different measurement.
What does move the number
The same survey collects the controlled evidence on individual tactics into a single table, and the summary is that little of it is strong. Keyword stuffing is null or negative across several benchmarks and is simply to be avoided. Formatting changes on their own generalise poorly; the gains seen are occasional and local to the test. An authoritative tone is weak and unstable as a lever and can conflict with credibility. The levers with moderate support are specific facts: recency, explicit prices, explicit dates, on the time-sensitive and commercial queries where a reader would want them. None of these is a rewrite of the body for the model's benefit. They are the page carrying the concrete information a buyer came for, in a form the search step can still find. That is also why our own diagnostic reports what the quoted page has that yours lacks, fact by fact, rather than handing you a rewritten draft.
What to do instead
Measure before touching anything, because a rewrite you cannot measure is a change you cannot reverse with confidence. Freeze the questions, record who is quoted today and from which page, and keep the answers. Then prefer changes that raise retrieval without lowering it: the specific price, the dated update, the comparison a buyer searches for, the page that exists for a question you currently have no page for. Check the page still ranks for its own query afterwards, because an AI Overview cannot quote a page Google does not surface. And treat any agency that promises a percentage lift from a body rewrite as someone quoting a harness result. The simulator below lets you set the two probabilities yourself and watch where the product turns negative. The ratio between them is your assumption, not a measured constant; the point is that it exists.
PIPELINE SIMULATOR
A page is quoted only if it is first retrieved and then chosen. A rewrite can raise the second chance while lowering the first. Move the sliders and watch the product.
Chance of being quoted, before
12.0%
After the rewrite
12.6%
Net effect
higher by 0.6 points
The effect turns negative once the retrieval loss ratio exceeds 0.67 at this lift. The ratio is a user-set assumption, not a measured constant; the SAGEO Arena result is that, on 171,003 documents, the net effect was negative.
SOURCES
Body-only optimisation reduced average top-20 presence by about 9%, top-10 presence after reranking by 16%, and final citation by 6% across 171,003 documents and 2,700 queries (SAGEO Arena).
Martinez, Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026), arXiv:2607.14035 · 2026-07-15 · 171,003 documents, 2,700 queries
The widely quoted 'up to 40%' improvement is a relative gain in position-adjusted word count (19.3 to 27.2 for quotation addition) inside a fixed five-document context.
Martinez, Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026), arXiv:2607.14035 · 2026-07-15 · 5 documents per query, GPT-3.5-turbo
Relevance and position are the primary determinants of first citation in a 252,000-trial factorial experiment across six LLMs and eighteen factors (Vishwakarma et al. 2026).
Martinez, Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026), arXiv:2607.14035 · 2026-07-15 · 252,000 trials, 6 LLMs, 18 factors
Keyword stuffing is null or negative; formatting alone generalises poorly; authoritative tone is weak and unstable; prices and dates have moderate, non-universal support (survey Table 4).
Income Lab (Pepform Pty Ltd). "Rewriting your pages can make you less visible". https://incomelab.me/findings/rewriting-your-pages-can-make-you-less-visible. Accessed 2026-09-22. Licensed CC BY 4.0.
READ NEXT