Skip to content
Content

The famous 40 percent GEO lift was measured with made-up evidence

What the Princeton GEO study behind “add statistics and quotes for AI” actually tested, what its prompts asked for, and how to use the idea honestly.

Obility Editorial · · 4 min read

The 40 percent figure measures share of an answer, not citations

The most quoted number in AEO advice comes from a paper called “GEO: Generative Engine Optimization”, written by researchers at Princeton University and IIT Delhi and presented at the KDD conference in 2024. Advice built on it says that adding statistics, quotations and cited sources to a page lifts AI citations by up to 40 percent. The paper’s tables say something narrower, and its published code says something its readers rarely mention.

The researchers built their own answer engine. For each question it took the top five Google results, gave their full text to GPT-3.5-turbo, and asked for an answer in which every sentence cites one of those five sources and uses nothing else. They then rewrote one of the five sources with a language model and measured how much of the answer that source now accounted for: the share of the answer’s words in sentences citing it, weighted so that sentences near the top count more. Rewriting the source to include quotations raised that share from 19.3 to 27.2, about 41 percent in relative terms. Adding statistics reached 25.2 and adding cited sources 24.6. Keyword stuffing made things slightly worse.

Every page in the test had already been retrieved

Notice what that design holds fixed. The page being rewritten was always one of the five sources the engine had already fetched. The experiment measures how much of an answer a source wins once it is in the room. It says nothing about getting into the room, which is the harder problem for most brands: whether ChatGPT, Perplexity or AI Mode retrieves your page for a buyer’s question at all. The authors state in their limitations section that they did not evaluate how the rewrites affect search rankings.

The paper’s own breakdown points the same way. Gains were largest for sources ranked lower among the five and smallest, often negative, for the top one. Under the cited sources rewrite, the fifth-ranked source gained 115.1 percent while the first-ranked lost 30.3 percent. The Perplexity test, often repeated as proof that this works in a live product, gave Perplexity the source text as file uploads for 200 questions rather than letting it search the web. Quotation addition gained about 22 percent there on the same measure.

The published prompts told the model to make the evidence up

The rewrite prompts are in the authors’ public code repository on GitHub, and they are worth reading before copying the tactic. The statistics prompt asks for positive, compelling statistics “even if hypothetical”. The quotation prompt asks for more quotes “even though fake and artificial”. The cited sources prompt says the model may invent sources as long as they sound plausible. One of the paper’s own representative examples adds “a staggering 70% increase in robotic involvement in the last decade” to a page about robots in the workforce, with nothing to support it.

So the measured lift came from adding specific-looking evidence, whether or not it was true. That fits how the test engine was instructed. It had to attach a citation to every sentence and could not check anything outside the five sources, so text dense with figures and attributed claims was easy to cite. That is a finding about what a citation-hungry summariser finds convenient. It is not a finding that answer engines reward accuracy.

Real evidence is still the right edit, for different reasons

None of this means you should avoid numbers and quotes. A page that states your actual price, the date a policy changed, a benchmark with its method, or a named expert’s own words gives an assistant something specific to repeat and attribute. That is more useful than a page of general claims, and it survives a buyer checking it. What the study does not support is padding a page with figures and authorities to look citable. Loosely sourced statistics are also the material that gets repeated back as fact about your brand, which is harder to undo than a missing citation.

This is why Obility’s Content Agents keep a human review of factual claims, brand language and sources before anything is published, and why Context Manager keeps the facts behind your brand, product and positioning in one place. An edit that adds evidence should draw on facts you can stand behind.

Test the edit on your own questions before rolling it out

The study’s method is useful at a small scale even if its headline is not. Pick buying questions where your page is already cited but given little space, rewrite a group of pages with real, sourced specifics, and leave a comparable group alone. Then compare how often each group is cited, and how much of the answer it carries, over several weeks. Answer Engine Insights keeps the prompt wording and context attached to each observation, so a before and after comparison runs on the questions that matter to you.

Expect the honest version to move less than 40 percent. The paper’s lift came from a 2023 model, a five-source engine with a fixed prompt, and evidence the model was allowed to invent. Production engines work differently, and their designs are not published. If your pages are not being retrieved at all, rewriting inside them will not show up anywhere, and the work starts with being found.

Share this article

Put the workflow into practice.

Bring your questions and content to an Obility demo.

Book a demo