What the Princeton GEO study actually found
Every pitch for "generative engine optimization" traces back to one academic paper. Here is what it measured, the nine edits it tested, the numbers behind the "up to 40%" claim, where it was tested, and the limits the authors put on it themselves.
The 2024 KDD paper "GEO: Generative Engine Optimization" tested nine content edits across 10,000 queries. Adding statistics, quotations and cited sources lifted a page's share of the AI answer by 30–40%; keyword stuffing lowered it. Gains were largest for low-ranked sources. It was tested on a research engine and on Perplexity via uploaded files, not the live web.
What are the key facts about the GEO study?
| Fact | Detail | Source |
|---|---|---|
| Who wrote it | Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, from Princeton, IIT Delhi, Georgia Tech and the Allen Institute for AI | arXiv |
| When | Posted to arXiv in November 2023; published at KDD '24, the ACM SIGKDD conference, Barcelona, 25–29 August 2024 | arXiv, ACM |
| What it coined | The term "Generative Engine Optimization" and a black-box framework for measuring and improving a source's visibility in an AI answer | arXiv |
| Test set | GEO-bench: 10,000 queries from a diverse set of domains and sources | arXiv |
| Tactics tested | Nine: Authoritative, Keyword Stuffing, Statistics Addition, Cite Sources, Quotation Addition, Easy-to-Understand, Fluency Optimization, Unique Words, Technical Terms | arXiv |
| Headline result | "GEO can boost visibility by up to 40% in generative engine responses" | arXiv |
| Best three | Cite Sources, Quotation Addition and Statistics Addition: 30–40% relative gain on Position-Adjusted Word Count, 15–30% on Subjective Impression | arXiv |
| Keyword stuffing | Position-Adjusted Word Count fell from 19.3 (baseline) to 17.7 | arXiv |
| Who gains most | Sources ranked lower in the answer's source list: with Cite Sources, the fifth-ranked source gained 115.1% on average while the top-ranked source lost 30.3% | arXiv |
| Live engine | Methods also tested on Perplexity.ai by uploading sources as files; visibility gains up to 37% | arXiv |
| Google's position | "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary" |
What is the GEO paper and who wrote it?
It is the academic paper that named the field. "GEO: Generative Engine Optimization" was written by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, researchers at Princeton University, IIT Delhi, Georgia Tech and the Allen Institute for AI. arXiv
The first version went up on arXiv in November 2023. The peer-reviewed version appeared in the Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '24), held in Barcelona on 25–29 August 2024. ACM
The paper's argument is simple. Generative engines such as Bing Chat, Google's AI features and Perplexity answer a question by reading several web pages and writing one response that cites them. A page that is not quoted in that response is invisible, even if it is in the source list. So the authors proposed a "black-box optimization framework": you cannot see inside the engine, but you can change your page, re-run the query and measure how much more of the answer comes from you.
How did the study measure visibility?
By how much of the AI's answer came from each source, not by rankings or clicks. The paper defines two metrics. arXiv
- Position-Adjusted Word Count. How many words of the answer are attributed to your source, weighted so that words near the top of the answer count more than words near the bottom.
- Subjective Impression. A score, produced by a language model, for how prominent, relevant and persuasive your source looked within the answer.
To run the test, the authors built a benchmark called GEO-bench: 10,000 queries drawn from a diverse set of domains and sources, each paired with the web pages a generative engine would read to answer it. For each query they took one source, applied one of the nine edits to it, had a generative engine answer the question again, and compared the edited source's share of the answer with its original share. The experiments were repeated across five random seeds and reported with standard deviations.
The generative engine in the main experiments was one the authors assembled from a search engine's results plus a large language model, built to behave like the deployed engines of 2023. That detail matters later.
Which of the nine tactics worked?
Adding evidence worked best. Rewriting for tone barely moved the needle. Stuffing keywords moved it the wrong way. These are the paper's own numbers for the baseline and the five most-discussed methods, on GEO-bench: arXiv
| Edit | Position-Adjusted Word Count | Subjective Impression | What the edit does |
|---|---|---|---|
| Baseline (no change) | 19.3 | 19.3 | The source as found |
| Keyword Stuffing | 17.7 | 20.2 | Adds more query keywords to the text |
| Fluency Optimization | 24.7 | 21.9 | Rewrites for smoother, clearer prose |
| Cite Sources | 24.6 | 21.9 | Adds citations to credible outside sources |
| Statistics Addition | 25.2 | 23.7 | Replaces vague claims with numbers |
| Quotation Addition | 27.2 | 24.7 | Adds quotes from relevant sources |
The paper reports that the best-performing methods improved on the baseline by 41% on Position-Adjusted Word Count and close to 30% on Subjective Impression. Taken as a group, Cite Sources, Quotation Addition and Statistics Addition gained 30–40% on the first metric and 15–30% on the second. Fluency Optimization and Easy-to-Understand, which only change how the text reads, still produced a 15–30% boost in visibility.
The remaining edits, Authoritative, Technical Terms and Unique Words, produced smaller or inconsistent gains. The authors note that generative engines were already fairly robust to a more authoritative tone on its own.
Did keyword stuffing hurt?
Yes, on the metric that counts. Keyword Stuffing was the only edit whose Position-Adjusted Word Count fell below the untouched baseline: 17.7 against 19.3. Its Subjective Impression score ticked up slightly, to 20.2, but the engine quoted less of the stuffed text and placed it lower. The authors' summary is that "Keyword Stuffing, traditionally used in SEO, does not perform very well" for generative engines. arXiv
This is the clearest break the paper found between old search habits and AI answers. A classic search engine matched query words to page words, so repeating the words could help. A generative engine reads the page and decides what to quote; a paragraph padded with repeated phrases has less in it worth quoting.
Do the same tactics work in every topic?
No. The paper broke results down by the domain of the query and found that the best edit changed with the subject. arXiv
- Statistics Addition did best for Law & Government questions and for Opinion questions, where a number settles an argument.
- Quotation Addition did best in People & Society, Explanation and History, where a direct quote adds a human voice or a primary source.
- Authoritative tone, weak overall, helped noticeably for debate-style questions and for historical questions.
- Combining edits helped: Fluency Optimization plus Statistics Addition gave the highest combined result in the paper's tests.
The authors' conclusion is that site owners "should strive towards making domain-specific targeted adjustments" rather than applying one recipe everywhere. For a business, that means: a pricing or legal question wants a number, a how-we-did-it story wants a quote, and both want to read cleanly.
Does GEO help small sites more than big ones?
In this study, yes. The paper reports gains by the source's original rank in the answer's source list, from rank 1 to rank 5. Sources that started near the bottom gained far more from the same edit than sources that started at the top. With Cite Sources, the fifth-ranked source's visibility rose by 115.1% on average, while the top-ranked source's fell by 30.3%. arXiv
The authors frame this as a "democratizing potential" for smaller content creators: in a classic results page, position ten is nearly invisible, but inside an AI answer a lower-ranked source that supplies the best quotable fact can end up contributing a large part of the response.
Was it tested on a real AI engine?
Partly. The main results come from the engine the authors built. To check that the findings carried over, they also ran the methods on Perplexity.ai, a deployed engine with a large user base. Because Perplexity does not let a user specify which web pages to answer from, the authors uploaded the sources as files and had Perplexity answer using only those files. On that test, visibility improvements reached 37%. arXiv
That is a reasonable check of the mechanism, how an engine quotes a set of sources once it has them, but it is not a test of live web search, where the engine also decides which pages to fetch.
What does the GEO study not prove?
Several things that it is often quoted as proving. None of these are criticisms of the paper; they are the boundaries the authors drew and the gaps between a 2023 experiment and the engines of 2026.
- It does not prove the numbers hold in today's engines. The main experiments used an engine the authors assembled in 2023 and models of that era. ChatGPT search, Google AI Mode and Perplexity have all changed how they retrieve and rank since then. The direction of the findings has been widely repeated; the exact percentages are specific to that setup.
- It does not test retrieval. Every edited page was already in the source list. The paper measures what happens after an engine has your page, not how to make it fetch your page.
- It does not measure clicks, leads or sales. Both metrics measure your share of the answer text. Being quoted more is a reasonable goal, but it is not the same as being chosen.
- It does not check whether the added facts were true. The edits were applied automatically with a language model. The study measured visibility, not accuracy. A real business should only add statistics and quotes it can stand behind, because a wrong number in an AI answer about you is worse than no answer.
- It is one engine's view of 'quality'. Google's page on AI features and your website says there are no additional requirements to appear in AI Overviews or AI Mode and no special optimizations necessary, and that no AI text files or special structured data are needed. Google Its guide to optimizing for those features asks for non-commodity content, not content that "could easily be produced by a generative AI model." Google Bolting statistics onto thin content is exactly the kind of commodity page that guidance warns against.
How should a business apply the GEO findings?
This is our advice, built on what the paper measured and on what Google and the AI companies publish. The goal is to give an engine something true and specific to quote, on a page it can already find.
- Start with pages AI already retrieves. Ask ChatGPT, Perplexity and Google a handful of customer questions and note which of your pages, if any, get cited. Improve those first; the paper's gains assume the page is in the source list.
- Add real numbers. Prices, timelines, counts, before-and-after results, survey findings. Statistics Addition was the strongest single edit for law, government and opinion questions. Label where each number came from and when it was measured.
- Add real quotes. A sentence from a customer who agreed to be quoted, a line from your own expert, a passage from a public standard or regulation. Quotation Addition was the top edit overall. Never invent a quote; see Do reviews affect AI recommendations? for how consented reviews do this job.
- Cite your sources in the text. Link to the study, the regulation, the manufacturer page. Cite Sources was a top-three edit and it costs nothing.
- Rewrite for fluency. Short sentences, one idea each, the answer in the first sentence under each heading. Fluency Optimization gained 15–30% on its own and more when combined with statistics.
- Match the edit to the question. A pricing page wants numbers; a case study wants a quote; a 'how does this work' page wants a cited explanation.
- Stop stuffing keywords. It was the one edit that reduced visibility. Write the phrase a customer would use once, in the heading, and move on.
- Re-run the questions. Ask the same prompts a month later and see whether more of the answer now comes from you.
What should be on my GEO content checklist?
| Check | Done when | Source |
|---|---|---|
| Page is retrievable | At least one AI engine already cites the page for a customer question | Our advice |
| Numbers present | The page carries at least three specific figures you measured or can attribute | GEO paper |
| Numbers labelled | Each figure says who measured it and when | Our advice |
| Quote present | At least one real, consented quotation from a customer, expert or public document | GEO paper |
| Sources linked | Outside claims link to the document they came from | GEO paper |
| First sentence answers | Every H2 is a question and the first sentence under it is the answer | Our advice |
| Fluency pass | Sentences average under 20 words; no paragraph over five lines | GEO paper |
| No keyword padding | The target phrase appears in the heading and naturally after that, not repeated | GEO paper |
| Not commodity | The page contains something only you could write: your data, your work, your customers | |
| Re-measured | The same prompts were run before and after the edit, and the results recorded | Our advice |
How do I know if the changes worked?
The same way the paper did: run the question, see how much of the answer comes from you, and compare with last time. The Prompt Simulator runs one question across ChatGPT, Gemini, Perplexity and Google's AI answers and shows which sources each engine quotes. Our free AI visibility check shows where your business is named today, so you have a baseline before you edit anything. For the full list of what to fix on a page before adding evidence, see The 12-point GEO audit checklist.
Sources
- GEO: Generative Engine Optimization (abstract) — arXiv. Read Oct 5, 2026.
- GEO: Generative Engine Optimization (full text, v3) — arXiv. Read Oct 5, 2026.
- GEO: Generative Engine Optimization, Proceedings of KDD '24 — ACM Digital Library. Read Oct 5, 2026.
- GEO: Generative Engine Optimization (publication record) — Princeton University. Read Oct 5, 2026.
- AI Features and Your Website — Google Search Central. Read Oct 5, 2026.
- Google's Guide to Optimizing for Generative AI Features on Google Search — Google Search Central. Read Oct 5, 2026.
- Creating Helpful, Reliable, People-First Content — Google Search Central. Read Oct 5, 2026.
Common questions
What is the Princeton GEO paper?
"GEO: Generative Engine Optimization" is a 2024 paper by Pranjal Aggarwal and colleagues at Princeton, IIT Delhi, Georgia Tech and the Allen Institute for AI, published at the ACM KDD '24 conference. It coined the term GEO, built a 10,000-query benchmark called GEO-bench, and tested nine content edits to see which made a page more visible in AI-generated answers. arXiv
Where does the 'GEO boosts visibility by 40%' claim come from?
From the paper's abstract, which says GEO "can boost visibility by up to 40% in generative engine responses." In the results, the best methods beat the baseline by 41% on Position-Adjusted Word Count and close to 30% on Subjective Impression. The figure applies to the authors' 2023 test engine and benchmark, not to any particular live engine today. arXiv
Which GEO tactics worked best in the study?
Quotation Addition, Statistics Addition and Cite Sources. On Position-Adjusted Word Count they scored 27.2, 25.2 and 24.6 against a baseline of 19.3, a 30–40% relative gain as a group. Fluency Optimization scored 24.7. Keyword Stuffing scored 17.7, below the baseline. arXiv
Does keyword stuffing help with AI search?
No. In the GEO study it was the only edit that lowered a source's Position-Adjusted Word Count, from 19.3 to 17.7. The authors wrote that keyword stuffing, traditionally used in SEO, does not perform very well for generative engines. arXiv
Was the GEO study tested on ChatGPT or Google?
No. The main experiments ran on a generative engine the authors built from search results and a language model. A second test ran on Perplexity.ai by uploading the source pages as files, where gains reached 37%. Google's own guidance says no special optimizations are needed for its AI features. arXiv Google
Should I add statistics to every page?
Add real ones to pages that answer questions where a number matters: prices, timelines, outcomes, legal or regulatory facts. The study found Statistics Addition strongest for Law & Government and Opinion questions, and Quotation Addition stronger for People & Society, Explanation and History. Invented or unsourced numbers are a risk, because the study measured visibility, not whether the numbers were true. arXiv
Related pages.
See how much of the AI answer comes from you.
Enter your domain. The free AI visibility check shows where ChatGPT, Gemini, Perplexity and Google's AI answers mention you, and which of your pages they quote.