Explainer
What the Original GEO Research Found, and What a UK Business Should Take From It
The paper that coined GEO tested nine content tactics against a benchmark of generative engines. Adding quotations and statistics performed best, with the best methods improving visibility by up to 41% on one measure, while keyword stuffing offered little to no improvement. Lower-ranked sites gained most. It was a 2023 benchmark on the engines of its day, so treat it as a useful hypothesis rather than a rulebook for Google today.
Almost every GEO article you will read cites the same source, usually vaguely. 'Research shows adding statistics boosts AI visibility by 40%.' I wanted to know exactly what the research said, so I read it. It is more careful, and more limited, than the way it is quoted.
This is a plain-English account of what the paper tested and found, the limits it states itself, and how I would apply it to a business in London without overclaiming. It matters because the paper is the closest thing GEO has to evidence, and a lot of money is being spent on the strength of a headline number.
What is the GEO paper?
It is a research paper titled GEO: Generative Engine Optimization, first submitted in November 2023 and accepted to the KDD 2024 conference. Its authors are Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, and it was last revised in June 2024.
The authors introduce GEO as a flexible framework to help content creators improve their visibility in generative engine responses, and they build a benchmark called GEO-bench, a large set of diverse user queries across multiple domains, with the web sources relevant to answering them. They then test content changes against it. That is a reasonable method. It is also a lab setting, which I come back to.
Which tactics did it test?
Nine content changes, applied to existing pages. Authoritative, Statistics Addition, Keyword Stuffing, Cite Sources, Quotation Addition, Easy-to-Understand, Fluency Optimization, Unique Words and Technical Terms.
| Tactic | What it changes in the content | Result reported by the paper |
|---|---|---|
| Quotation Addition | Adds relevant quotations | Among the best performers |
| Statistics Addition | Adds relevant statistics | Among the best performers |
| Cite Sources | Adds citations to sources | Strong, especially for lower-ranked sites |
| Authoritative | Makes the tone more persuasive and authoritative | Helped most on debate-style and historical queries |
| Keyword Stuffing | Adds more keywords, as in classic SEO | Little to no improvement |
| Easy-to-Understand, Fluency, Unique Words, Technical Terms | Style and vocabulary changes | Not highlighted among the top performers |
The paper states that the best methods improved on the baseline by 41% and 28% on two of its measures, position-adjusted word count and subjective impression. Its abstract summarises this as GEO boosting visibility by up to 40%. That 'up to' is doing real work. It means the best tactic on the best measure, not the typical result.
What did it say about keyword stuffing?
That it offered little to no improvement in generative engine responses. The paper notes that keyword stuffing is widely used in classic SEO and finds such methods do not help here. It scored below the unmodified baseline in the paper's main table.
This is the finding I find most practically useful, because it rules out a whole genre of lazy advice. If someone proposes inserting a list of location keywords or repeating a phrase so an AI picks you up, the primary research evidence points the other way. Evidence-bearing text outperformed repetition.
Why did quotations and statistics work?
The paper does not claim to know the mechanism, but the pattern is consistent with how these systems work. A generative engine builds an answer from many sources and has to justify what it says. Text containing a specific figure, a named source or a direct quotation is easier to lift into an answer and easier to attribute than a general statement.
That is my interpretation, not the authors' finding. It also matches the commonsense reading for a human reader: a page with a sourced figure is more useful than one without. I return to this in the post on writing content AI quotes.
Who benefits most from GEO?
The paper found that lower-ranked websites benefit significantly more. It reports that the Cite Sources method led to a 115.1% increase in visibility for websites ranked fifth in search results, while the visibility of the top-ranked website decreased by 30.3% on average.
For a smaller UK business, that is the headline worth keeping. If the pattern holds, answer engines give a well-evidenced smaller site a route to visibility that classic rankings may deny it. I would not bank on the precise figures, but the direction is encouraging, and it is consistent with why I think AI search is a realistic channel for independents in a crowded London market.
Does it work the same in every subject?
No, and the authors say so. Their abstract states that the efficacy of these strategies varies across domains, underscoring the need for domain-specific optimisation methods. Within the results, adding statistics was especially helpful for law and government topics and opinion-style questions, adding citations helped factual questions, and quotations helped most in people, society, explanation and history topics.
- Law and government topics: statistics helped most.
- Factual questions: citations helped most.
- Debate and historical questions: an authoritative tone helped.
- People, explanation and history: quotations helped most.
For a UK solicitor, accountant or dentist, that suggests the right evidence differs by what the page is about. A page on a legal process may benefit from concrete figures and references, while a page about a person or a practice may benefit from a quoted client or colleague, with consent. Do not apply one tactic to everything.
How could I test this on my own site?
Run a small before-and-after test on a few pages, with untouched pages as a control. The paper is a benchmark, but you can borrow its logic on a modest scale. It will not be statistically rigorous, and it does not need to be, as long as you are honest about what it can show.
- Pick six pages on similar topics. Choose three to change and leave three alone.
- Write your prompt set and run it three times on separate days to get a baseline appearance and citation rate for all six pages.
- Add evidence to the three: a real figure with its basis, a named source, and a consented quotation where it fits. Change nothing else.
- Wait six to eight weeks and re-run the identical prompts three times.
- Compare the changed pages with the controls. If they improved and the controls did not, you have a hint. If both moved, the cause is elsewhere.
With so few pages and such variable answers, treat the result as a nudge, not a verdict. But it replaces borrowing a headline from a 2023 paper with seeing what happens in your own market, which is a better basis for spending money.
What are the limits of the paper?
It is a 2023 benchmark on the generative engines of that period, not a test of Google's current AI features or of your own market. The experiments used the authors' own benchmark and engines available at the time. AI products change quickly, and results from one system and date do not carry over automatically.
- Benchmark, not the wild. GEO-bench queries are diverse but they are not UK local searches.
- Engines have changed. The systems tested are not the ones answering your customers now.
- Measures are proxies. Position-adjusted word count and subjective impression are the paper's metrics of visibility, not enquiries.
- No local business angle. The paper is not about Maps, profiles or reviews, which dominate local recommendations.
How does this square with Google's advice?
Google says you do not need to write in a special way for generative AI, and the paper's best tactics are things good content does anyway. Google's guide lists rewriting content just for AI systems among the things to ignore. At first glance that contradicts a paper recommending changes to content.
I think the tension is only apparent. Google is warning against AI-specific rewriting, such as spinning variations or restructuring into tiny pieces. The paper's winners are adding verifiable evidence: a figure, a source, a quotation. That is not writing for a model. It is writing better for a reader, and it happens to help the model too. I discuss Google's position in the post on Google's AI guide.
How would I apply it honestly in a UK business?
- Add real figures to your key pages, drawn from your own work, with the basis stated. 'Typically eight to twelve weeks' is better than 'quickly'.
- Cite your sources. Link to the regulator, the statute, the study or the standard you are relying on, so a claim and its support sit together.
- Use real quotations, from clients who have consented or from named colleagues, never invented ones.
- Stop keyword stuffing, including repeated location lists. The evidence says it does not help.
- Adjust by subject. Match the kind of evidence to the kind of page.
- Measure the result across repeated prompts, as in how to check your AI visibility, not by a single run.
Straight answers
Questions
What is generative engine optimisation?
The practice of improving your content's visibility in generative engine responses. The term was formalised in a 2023 research paper, accepted to KDD 2024, which introduced a benchmark called GEO-bench and tested nine content tactics.
Does adding statistics really boost AI visibility by 40%?
The paper says GEO can boost visibility by up to 40%, and that the best methods improved on baseline by 41% and 28% on two of its measures. That is the best case on a research benchmark in 2023, not a typical result for your site.
Does keyword stuffing work for AI search?
The paper found keyword stuffing offered little to no improvement on generative engine responses, scoring below the unmodified baseline in its main results.
Do small websites benefit from GEO?
The paper reports that lower-ranked websites benefit significantly more, citing a 115.1% visibility increase from adding citations for sites ranked fifth, with the top-ranked site's visibility falling 30.3% on average.
Is the GEO paper still valid?
It is a useful hypothesis rather than a rulebook. It tested a 2023 benchmark on the engines of that time, and the authors themselves note results vary across domains. Test any tactic on your own prompts before relying on it.
