Information gain SEO works when a page covers the question fully and adds one useful fact that the current results do not have.
I checked 238 live US Google searches across 12 unrelated niches and read 1,334 publisher articles that ranked in their top 10. Pages in the most complete fifth were quoted by Google AI Overviews 49% of the time. Pages in the least complete fifth were quoted 35% of the time.
Backlinks moved far less in this snapshot. Pages with no referring domains were quoted 43% of the time. Pages with 342 referring domains were quoted 45% of the time.
This was an observational study from one day. It shows what appeared together and does not prove what caused a citation.
If you want to run the research and publishing workflow from one place, try Distribb.
![]()
Eighty-six percent of the informational searches in the sample showed an AI Overview. For every search that showed one, I split the ranking publisher pages into two groups: pages the AI Overview quoted and pages it skipped.
Both groups had already earned a top-10 position for the same search on the same day. That made it possible to compare coverage, backlinks, and original information without comparing unrelated queries.
I also had 435 ranking pages judged from their body text. The judges did not see the URL, domain, title, or citation status. They scored how much of the question each page covered and how much it said that no rival page said.
Ranking position could have explained the first result. Pages that cover a topic well can rank higher, and higher-ranking pages can get quoted more often.
So I held position still and compared pages again.
Among pages ranked 1 to 3, the best-covered third was quoted 74% of the time, compared with 63% for the least-covered third. In positions 4 to 6, the comparison was 48% to 39%. In positions 7 to 10, it was 34% to 24%.
The gap became smaller lower on the results page, but it remained in every ranking band.
Original information on its own did not predict citation rates. The four bands from the least original to the most original were quoted 46%, 48%, 46%, and 45% of the time.
The combination mattered. Pages that covered the topic and added a point no rival made were quoted 53% of the time. Pages that added something new while missing part of the topic were quoted 29% of the time.
The practical order is simple. Cover the question first, then add the result, example, or experience only you can provide.
Start with the full text of the ten organic results, the transcripts from the top five relevant YouTube videos, the People Also Ask questions, and the related searches.
Save each source as plain text or Markdown. Keep a list of URLs that failed to download. In my study, 387 of 2,998 pages could not be fetched. A blocked page is unknown, so do not count it as a page that says nothing.
For a query such as cold brew ratio, the folder can contain roughly 40,000 words. Claude Code can collect it in about ten minutes.
Turn the corpus into a table. Put each distinct claim or sub-question on a row and the ten ranking pages in columns.
Merge claims that say the same thing in different words. Mark the pages that answer each claim. If six or more pages cover one claim, treat it as required coverage for your page.
This table sets the floor for the article. It also stops the draft from repeating one popular point while missing the rest of the question.
Read Reddit threads, Quora answers, comments below the videos you transcribed, and reviews when the topic involves a product. Add questions from your own support inbox or sales calls when you have them.
Keep the wording people used. Then compare every question with the corpus from step 1 and remove any question a ranking page already answers.
An unanswered question still needs demand. Some gaps exist because very few people care about the answer.
Put the coverage table next to the unanswered-question list and pick one gap.
Choose a question you can answer with evidence. A number should settle the question, and your business should have a reasonable way to collect that number. Check your own site before you begin so the new page does not compete with something you already published.
How does cold brew change after 24 hours in the fridge? can be measured. More detail on brewing cannot.
Original research can start with data you already own. Support tickets, delivery times, quote ranges, intake forms, and the objections in recent sales calls all count when they answer the question.
Public datasets are another option. You can also assemble a corpus from product reviews, competitor pricing pages, or posts in a relevant community. The information gain study in this article used the search results themselves.
Write the rules before looking at the results: what you will count, what you will exclude, how you will choose the sample, and what result would prove your assumption wrong.
Run the study, then put one short result paragraph near the top of the article.
The paragraph needs the sample, the comparison, the result, and one plain limitation. A reader should understand it without knowing the model name or statistical method.
Publish a result that disagrees with your original idea. Choosing which studies to publish after seeing the answer turns the research into marketing.
Use the coverage table as a checklist. Give each required point the space it needs, put the new finding near the top, source every number, and answer the questions from step 3 using the language people used.
Publish the underlying data when you can. A public method and dataset make the result easier to check and cite.
Google Search Console has a generative AI performance report for sites with eligible data. Track appearances there and record links or citations that point to the study. Google also states that third-party tools do not have access to its internal ranking or AI systems, so treat outside citation trackers as samples rather than a complete record.
The study supports one practical rule: complete coverage was associated with a larger change in AI Overview citation rates than backlink count in this sample. Original information helped most when the page also covered the question well.
It does not support a ranking guarantee. The data came from one US snapshot across 238 searches, and an observational study cannot prove that one page characteristic caused the citation.
Information gain in SEO is the useful information your page adds beyond what the current ranking pages already say. It can be a first-party result, a checked example, an expert observation, or a new analysis of public data. This study found that original information performed best when the page also covered the full question.
Read the full top 10 and compare every claim in a coverage table. Then collect real questions from communities, video comments, reviews, support conversations, or sales calls. Remove any question that a ranking page already answers and choose one remaining gap that you can prove with evidence.
You need a useful contribution, and it can come from data you already have or public data you reanalyse. A support-ticket count, a review corpus, competitor pricing pages, or an open government dataset can answer a real question. The method and sample should be stated clearly so a reader can check the result.
The study found many quoted pages with few or no backlinks, but it does not prove that backlinks do not matter. Pages with zero referring domains were quoted 43% of the time, compared with 45% for pages with 342 referring domains. One in five quoted pages had no backlinks in this snapshot.
The coverage and question-mapping steps can fit into one afternoon, while the study can take an afternoon or several days. The source workflow budgets about ten minutes for collection, 15 minutes for the coverage table, 20 to 30 minutes for unanswered questions, and 15 minutes to choose the measurable gap. Study design and data collection take the remaining time.
Use Google Search Console's generative AI performance report when it is available for your property. Record the page, date, appearances, and any links or human citations that follow. Outside tools can sample prompts, but Google says they do not have access to its internal ranking or AI systems.
Which step would slow your team down today: building the coverage table or collecting a result that no competing page has?
If you want to run this research and publishing workflow from one place, try Distribb.