Full research findings
Everything four parallel research threads turned up on tools, open-source projects, and the evidence base for analyzing the top 20 results and outranking them. This is the source material. The decisions drawn from it live in The SERP Teardown Method.
Four threads ran in parallel: the commercial SaaS market, the open-source and DIY layer, the evidence base for what actually moves a page into the top 10, and the evidence on search intent and SERP diagnosis. Every price below was fetched from the vendor's own live pricing page on 26 July 2026, not from a review site. GitHub star counts and last-commit dates came from the GitHub API, package versions from the npm and PyPI registries.
Claims are graded, because most of what circulates as SEO fact is not:
| Grade | Meaning |
|---|---|
| T1 | Sworn testimony, court exhibits, or first-party Google documentation |
| T2 | Controlled split tests or matched difference-in-differences designs |
| T3 | Large-sample observational research, causation unproven |
| T4 | Folklore. Repeated widely, no traceable source, or the source does not say what people claim |
When a statistic sounds decisive, find the denominator before believing it. Three worked examples. "Wikipedia is 47.9% of ChatGPT's citations" is really 7.8% of citations and 47.9% of the top-ten source share. "52% of AI Overview citations come from the top 10" is 52% of the roughly half that ranked anywhere at all, so about 26% of citations. "Schema makes you 3x more likely to be cited" is a selection effect that vanishes under matched controls. The field's most-repeated numbers are mostly denominator errors.
This is the category that sells "analyze the top pages and tell me what to write." Two structural things happened in the last year: Surfer was acquired by Positive Group (announced October 2025, now branded Positive Surfer), and MarketMuse was absorbed by Siteimprove (completed 31 October 2024). The category has also split, with Frase and NeuronWriter going agent-native while Surfer, Clearscope and Scalenut pivoted their messaging to AI-visibility tracking. Nobody leads with "TF-IDF content editor" anymore.
| Tool | Entry price | Reads | Briefs included | Per brief | API / MCP |
|---|---|---|---|---|---|
| Surfer (Positive Surfer) | $99/mo yearly | Top 20 | 360 docs | $0.28 | API at $299 tier; MCP announced, undated |
| Clearscope | $129/mo | Top 30 | 20 | $6.45 | Neither. API hosts do not resolve |
| Frase | $39/mo yearly | Top 20 | 10 articles | $3.90 | Official MCP, pushed daily. 50+ REST endpoints, public docs |
| NeuronWriter | $19/mo yearly | Top 30 | 25 to 150 | $0.76 | Official MCP + public API at Gold ($57/mo) |
| Page Optimizer Pro | $40/mo | User-selected | credit-based | varies | API is a $10/mo add-on on any plan, docs post-purchase |
| Scalenut | $59/mo (promo $24) | Top 30 | 5 to 75 | varies | Neither. Helpdesk returns Payment Required |
| MarketMuse | no public price | Top 20 | 5 to 20 | unknown | Gateway alive but undocumented |
| Outranking | unverifiable | - | - | - | Every host fails TLS. Appears to be failing |
Surfer analyses the top 20 by default and now splits its Content Score into an SEO Score (keywords, NLP terms, "true density," heading structure, length versus SERP, images) and an AI Search Score (facts coverage, whether the page answers the primary question early). Notably its own documentation warns against the product's own headline metric: aim for 70 to 85, because "over-optimizing can actually hurt your final content."
Clearscope is refreshingly plain about its mechanism: it "calculates the importance of each term by how much the keyword appears in the competitors' articles." That is term frequency, stated by the vendor. It has the best practitioner reputation in the category for restraint, and a genuinely novel feature in AI Term Presence, which flags which suggested terms actually appear in Gemini and GPT answers for the query.
Page Optimizer Pro is the most differentiated and the least fashionable. It is the only tool here that pulls competitors' schema and compares it against yours, and it runs competitor content through the Google Cloud NLP API to show entities and categories. Its epistemology is single-variable live testing rather than correlation, which is a more serious claim than the others make and also the least independently verifiable, since the tests are not published.
Frase is the only one scoring author and E-E-A-T signals as a first-class metric, and its documentation is unusually honest: "Scores are feedback, not goals. Don't chase perfect 100s. Focus on beating competitor averages."
Short answer: it is a real but weak correlation, computed close to circularly, with no causal evidence in either direction.
It comes from Ahrefs, a vendor in the category, which is what makes it usable: "Keyword density is not topic coverage."
Semrush was acquired by Adobe, completed 28 April 2026, all cash at $1.9 billion. No price rises yet. The warning worth carrying is Cyrus Shepard's: Adobe's tooling leans enterprise, which historically means drift away from small operators.
| Tool | Entry | What it does for top-N analysis | Verdict for a solo consultant |
|---|---|---|---|
| Ahrefs | $29 Starter $129 Lite | Best raw SERP and backlink data. AI Content Helper deliberately rejects keyword-density scoring and grades topical coverage instead | Buy one seat when a client needs it. The backlink index is the one asset you cannot rebuild |
| Semrush (Adobe) | $139/mo | On Page SEO Checker aggregates a fixed, non-configurable top 10 across six factors. Uniquely includes referring domains shared across the top 10 | Skip unless a client recognizes the brand. MCP needs the $199 tier |
| seoClarity | from $2,500/mo | Most brief-shaped output in the category, plus a genuine first: an AI Overviews Analysis Mode that analyses what performs inside the AIO specifically | Out of budget by an order of magnitude |
| Conductor | median ~$49k/yr | Undisclosed methodology. Writing Assistant caps at 60 drafts per year on the entry tier | No. Contracts are 1 to 3 years, no monthly |
| Similarweb | ~$125/mo | Real rank tracking from the Rank Ranger acquisition, but no content optimizer and no brief of any kind | No. Estimates break below 5,000 visits a month, which is every Front Range local business |
Semrush's MCP server ships with 50,000 API units included on the Starter and Pro+ tiers, but not on the more expensive Advanced tier, where you must buy unit packages separately. Buying the pricier plan removes the free units. The cheapest path to Semrush MCP is the $199 Starter plan.
Independent testing puts Semrush traffic estimates as reasonable above 50,000 organic visits a month and "not accurate at all" below 5,000, with a 30 to 60 percent under-report on small sites. Ahrefs' estimates swing between minus 80 and plus 80 percent on newer domains. For local service businesses, treat all third-party traffic estimates as directional only and lean on Search Console.
One index-size correction, since the marketing comparison is misleading: DataForSEO's 2.8 trillion links counts live links only, while Ahrefs' 35 trillion and Semrush's 43 trillion count cumulative including lost links. The gap is real but nowhere near twelve-fold.
Version 24.3, £199 per year (shown as $279). It cannot fetch a Google SERP. Its "SERP mode" is a snippet pixel-width preview where you upload titles and meta descriptions, not organic-results retrieval. You must source the top-20 URL list elsewhere.
What it does superbly is list mode plus custom extraction: paste twenty competitor URLs and pull headings, word counts, schema blocks, author bylines and dates via XPath or regex. Version 24 shipped a built-in MCP server, but with a catch confirmed by their own staff: list mode is not yet supported over MCP, which kills the obvious "hand Claude twenty competitor URLs" workflow for now.
Analyses up to 100 results, extracts headings, word counts and on-page data, covers AI Overviews and AI-engine citation sources, and ships both a SERP API and a Brief API. Starter $13/mo yearly, Pro $33, Agency $66, with published per-unit prices of $0.47 per brief at the Pro tier. A 100-result brief at 47 cents with an API is close enough to DIY cost that it is the honest comparator for a build decision, not Surfer.
Both are the wrong category for this job. Sitebulb has no SERP retrieval, no competitor analysis and no brief; it is a site crawler with the best plain-English audit output on the market. Rank Math's Content AI surfaces suggested keywords but exposes no top-N breakdown, and its credits are metered at one credit per word generated with keyword research costing 500 credits per operation.
| Endpoint | Price |
|---|---|
| SERP Google Organic, standard queue (~5 min) | $0.0006 per 10 results |
| SERP Google Organic, live (~6 sec) | $0.0020 per 10 results |
| On-Page instant pages / content parsing | $0.00015 per page |
| On-Page with JavaScript loaded | $0.0015 per page |
| On-Page Lighthouse | $0.005 per page |
| Labs keyword endpoints | $0.012 per task + $0.00012 per item |
| Backlinks, all endpoints | $0.024 per request + $0.000036 per row |
| LLM Scraper (real ChatGPT interface) | $0.0012 to $0.004 per page |
1. depth: 20 returns page one only unless you also set max_crawl_pages: 2. A naive build silently analyses ten results and calls it twenty.
2. Depth multiplies cost per each ten results, so top 100 costs ten times top 10.
3. Each advanced search operator such as site: or inurl: multiplies the task cost by five, and they stack.
Blunt finding: not one of them does top-20 ranking-page analysis or produces an outrank plan. The category runs on a different primitive entirely, which is firing prompts at LLMs repeatedly, parsing the answers for your brand and the cited URLs, and charting it.
| Tool | Price | Status |
|---|---|---|
| Profound | $99 and $399/mo | Raised $96M Series C at a $1B valuation, February 2026. API is Enterprise only |
| Peec AI | ~$95 to $495 | Uses UI scraping rather than APIs, deliberately. Has a real MCP |
| Scrunch AI | $250/mo | Acquired by Sitecore for $225M, 3 June 2026. API, MCP and CLI all Enterprise only |
| Athena HQ | $295/mo | Cheapest real API in the pure-play group |
| Otterly.AI | ~$29 to $489 | Google AI Mode and Gemini are paid add-ons |
| Evertune | quote only | Methodologically the most honest. Mass sampling, which is what the research says is required |
| Goodie AI | - | Dead. The domain is a parking page for sale |
Profound's own integrations page lists DataForSEO as its partner for real-time SERP data. The billion-dollar category leader buys its SERP layer from the same commodity API already wired into this setup. Profound Growth at $399 a month buys 9,000 responses; the same 9,000 through DataForSEO's LLM Scraper costs about $11.
Rand Fishkin and Patrick O'Donnell ran 12 identical brand-recommendation prompts 2,961 times through 600 volunteers across ChatGPT, Claude and Google AI. The odds of the same brand list appearing twice were under 1 in 100. The same list in the same order: under 1 in 1,000. What survived as a stable measure was membership in the consideration set, expressed as a visibility percentage across many runs, never a rank.
The variance stacks in ways nobody discloses. Querying ChatGPT via the official API returns an average of 7 sources while scraping the actual web interface returns 16, with only 8% source overlap on Perplexity between the two collection methods. Google's own AI Overviews and AI Mode cite the same URLs only about 13.7% of the time. Roughly 11% of domains are cited by both ChatGPT and Perplexity. Every dashboard in this category renders a single draw from that distribution to two decimal places.
My read: the measurement layer is a bubble consolidating into incumbents and DXP vendors, while the execution layer is not. Profound pivoting into Agents and Evertune into a ChatGPT Ad Agent are both admissions that dashboards will not sustain those valuations. Structurally, Ahrefs bundling prompt tracking at $129 and Semrush at $199 caps what a standalone tracker can ever charge.
There is no maintained open-source library that does "analyze the top 10 to 20 ranking pages and generate a plan to outrank them." Every component exists. The assembly does not. For a consultancy that already pays for DataForSEO and runs Claude Code, that is a build opportunity rather than a gap.
| Project | Stars | License | What it is |
|---|---|---|---|
| OpenSEO | ~8,100 | MIT | Self-hostable on Cloudflare's free plan, bring-your-own DataForSEO key, ships its own MCP server plus Claude Code skills. The highest-leverage item found. |
| searchsolved (Lee Foot) | 407 | unclear | ~60 standalone tools. A July 2026 commit migrated twelve of them onto DataForSEO. Includes a SERP crossover analyzer and a keyword-gap analyzer |
| Agentic-SEO-Skill | 791 | MIT | 16 sub-skills and 88 Python evidence-collector scripts. The scripts are the value |
| seranking/seo-skills | 100 | MIT | 26 Claude Agent Skills. Its seo-content-brief spec is the most thorough brief structure published anywhere. Read it even if you never install it |
| iannuttall/seo | 45 | Apache-2.0 | Three weeks old and committing daily. Records the API cost of every report and caches so you never pay twice. Steal that design decision |
| Qforia | - | open | iPullRank's query fan-out simulator. Publishes its production prompt verbatim, so you can lift it straight into Claude |
| advertools | 1,423 | MIT | SERP ingestion, Scrapy-based crawling and crawl diffing in one dependency. Its new serp_claude function tracks what Claude's web search surfaces, and nothing else does that |
| ContentSwift | 161 | open | Explicitly positioned as a free alternative to Surfer, NeuronWriter and Frase |
trafilatura (v2.1.0, June 2026) is the default choice, precision-oriented and admitting only about 6.6% boilerplate. rs-trafilatura scores highest on the current benchmark and is fastest at 44ms per page, but carries real bus-factor risk. crawl4ai (75,000 stars, Apache-2.0) is the right pick when you need JavaScript rendering or schema-driven extraction.
The repo shows 15,125 stars and a July 2026 push, which looks healthy. The last six commits touch only the README, adding and removing proxy sponsors. The package on PyPI is still v0.2.8, uploaded September 2018. Eight years with zero code changes. Use newspaper4k if you want that API.
Several SEO-skill repos created in 2026 show 9,000 to 41,000 stars while linking out to paid communities. Genuinely good work in this category sits at 45 to 250 stars. Star counts in this niche are being gamed.
DEJAN's Grounding Snippet Extraction Tool uses Gemini's live search grounding and returns the exact URLs and exact sentences Google actually extracted for a query. Every other tool in this survey guesses at retrievability using a public embedding model. This one shows you the real output, and it is free. It is the closest thing to ground truth available to anyone outside Google.
The job has two halves: get a positional SERP, and get the ranking pages in analysable form. Most tools do one.
| Tool | Positional SERP? | Scrapes ranking pages? | Geo-targeting | Verdict |
|---|---|---|---|---|
| DataForSEO | Yes, plus AI Overview and all features | Yes | City level | The only one doing the whole job in one vendor |
| Bright Data | Yes, four engines | Yes | Yes | Strong second. 5,000 free credits a month, no card |
| Firecrawl | No, it is a search proxy not a ranked SERP | Best-in-class markdown | Country only | Great scraper, wrong tool for rank |
| Exa / Tavily | No, semantic indexes | Yes | Limited | Wrong tool for SERP work |
| Jina Reader | Not Google truth | Yes, excellent | No | Cheapest page fetcher. 20 requests/min with no key at all |
| Ahrefs / Semrush | Tracked-rank data, not live fetch | No | Yes | Metrics layer, not a pipeline |
Ahrefs archived its local npm MCP server on 24 February 2026 and the README now tells you not to use it. Any 2025 tutorial installing that package sends you somewhere broken. The current product is a remote OAuth-based hosted server.
Among the paid content tools, only Frase and NeuronWriter ship official MCP servers today. NeuronWriter's evaluate-content method is worth singling out: it scores a draft without consuming an analysis credit or creating a revision, which makes it the cheapest correct scoring primitive available anywhere if you want a tight iteration loop.
T1 From DOJ trial exhibits, Google's own engineers describe topicality as built from ABC signals: Anchors, Body, and Clicks, where clicks means how long a user stayed on a linked page before returning to the SERP. That is dwell time in all but name. Ex-Googler Eric Lehman on NavBoost: "it's just a big table. It says for this search query, this document got two clicks." Trained on 13 months of rolling data.
Pandu Nayak, asked directly by the judge where NavBoost sits in importance: "navboost is important, right. So I don't want to minimize it in any way. But I will also say that there are plenty of other signals that are also important." He ranked the document itself first, then topicality, page quality, reliability, localization, then NavBoost.
T1 The leaked Content Warehouse documentation contains siteAuthority. Google spokespeople had denied domain authority for years. The DOJ deposition is blunter still, on the page-quality signal: "Q is largely static and largely related to the site rather than the query."
This is the mechanical reason a low-authority site underperforms on merit. A static, site-attached quality score gates you before page content is even assessed. Combined with homepagePagerankNs being copied onto new pages as a proxy, and NavBoost having no table entry for a page with no click history, a new page on a weak domain is not merely unranked, it is unmeasured.
T2 SearchPilot's controlled split tests, which use control and variant groups and so rule out seasonality and algorithm updates:
Zyppy's 23-million-link study found traffic rising with incoming internal links up to roughly 45 to 50, then reversing, explained by sitewide navigation links pointing at low-traffic URLs.
The cleanest illustration comes from one researcher finding opposite results in two datasets. Zyppy's 23-million-link study found anchor-text variety "highly correlated with higher search traffic." Zyppy's 50-site study of Google update winners and losers found internal anchor variations correlating at minus 0.337 with traffic change, and external anchor variations at minus 0.352, both graded strong evidence.
The likely reconciliation is that anchor diversity proxies for "a site that does deliberate SEO," which was rewarded before the helpful-content era and penalized after. Correlation studies measure the current equilibrium, not a lever.
| Figure | What it means |
|---|---|
| 96.55% | of all pages get zero Google search traffic. The sample skews to the quality side of the web, so reality is worse |
| 1.74% | of newly published pages reach the top 10 within a year for even one keyword, down from 5.7% in 2017 |
| 40.82% | of pages that do reach the top 10 got there within one month. The distribution is fast-or-never, not a slow grind |
| 84.8% | of top-three entrants after the May 2026 core update came from pages already in the top 20. Improving near-misses beats publishing new pages |
| 32.2% | of domains that lost top-10 positions in March 2026 had recovered by May. Recovery is the exception |
| 49% | of the time, the top-ranking page gets the most search traffic. Ranking number one is overrated |
BrightLocal classified the first ten organic results for local-intent terms. Overall: business websites 47%, directories 31%, business mentions 16%, forums 7%. The per-keyword breakdown is the useful part, because one word flips it:
| Search term | Business website | Directory | Forum |
|---|---|---|---|
| dentist | 88% | 3% | 0% |
| best dentist | 37% | 34% | 9% |
| electrician | 59% | 35% | 2% |
| best electrician | 24% | 57% | 11% |
| attorney | 41% | 52% | 0% |
| best attorney | 11% | 68% | 7% |
| vet clinic | 96% | 3% | 0% |
Note this data also refutes the strong form of the content-type rule. Business websites still hold 53% for "best chiropractor" and 63% for "best vet clinic." Mismatch is a heavy prior against you, not a hard block. Google's own rater guidelines instruct raters to reward format diversity, and to treat a more-specific brand page as Highly Meets for a broad category query if the brand is popular and prominent. The operative variables are prominence and specificity distance, not format.
Quality Rater Guidelines, section 12.2, verbatim: "If you research the query on Google, please do not rely on the top results on the SERP. A query may have other meanings not represented on Google's search results pages." Reading the SERP is still the best available proxy for intent. It is a proxy.
| Study | Sample | Finding |
|---|---|---|
| Ahrefs, July 2025 | 1.9M citations | 76.1% of AI Overview citations rank top 10 |
| Ahrefs, March 2026 | 863K SERPs, 4M URLs | 38% top 10, 31% at 11 to 100, 31% beyond 100 |
| BrightEdge, 16 months | 9 industries | Overlap grew 32.3% to 54.5%. Opposite direction |
| seoClarity | 362,000 keywords | 33% citation rate for position one; 94% of AIOs cite at least one top-20 result |
| Originality.AI | YMYL queries | 52%, but of the half that ranked at all, so ~26% of citations |
| Ahrefs cross-engine | 15,000 prompts | Only 12% of ChatGPT, Gemini and Copilot citations appear in Google's top 10. Perplexity 28.6% |
These are not measuring the same thing. Ahrefs counts SERP blocks including features, BrightEdge counts organic-only overlap, and the cross-engine figure covers non-Google assistants. Ahrefs states plainly that its own parser improved between its two studies so they are not directly comparable, and that AI Overviews moved to Gemini 3 in January 2026. Anyone quoting "AI citations collapsed from 76% to 38%" as a trend is over-reading the vendor's own caveat.
DEJAN analyzed 883,262 snippets across 7,060 queries and 2,275 tokenised pages:
Petrovic's conclusion, and it contradicts what the industry still sells: density beats length.
The Princeton GEO paper (KDD 2024) is cited everywhere for "up to 40% visibility improvement." Its measured lifts on Position-Adjusted Word Count: quotation addition +41%, statistics +31%, fluency +28%, citing sources +27%, and keyword stuffing at minus 8%, worse than doing nothing.
Table 3 is the one that matters, and almost nobody quotes it:
| Tactic | Rank 1 | Rank 2 | Rank 3 | Rank 4 | Rank 5 |
|---|---|---|---|---|---|
| Cite sources | -30.3% | +2.5% | +20.4% | +15.5% | +115.1% |
| Quotation addition | -22.9% | -7.0% | +3.5% | +25.1% | +99.7% |
| Statistics addition | -20.6% | -3.9% | +8.1% | +10.0% | +97.9% |
These tactics roughly double visibility for the weakest source and actively hurt the strongest. It is a leveling effect. Caveats: the engine was a researcher-built pipeline over 2023-era results, not any production system; "visibility" means share of words in the answer, not clicks; and no independent replication exists on 2026 engines.
C-SEO Bench, the one benchmark with open code, data, and an adversarial design, found that "most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking," and that traditional SEO aimed at improving the source's rank within the LLM context was significantly more effective. It also found that as more actors adopt these tactics the gains shrink, describing the problem as congested and zero-sum. That is the same shape as local map-pack SEO, and it is relevant to how a GEO engagement should be priced and scoped.
T3 Across 6 million URLs, AI-cited pages were almost 3x more likely to have JSON-LD. That is the stat on the conference slides.
T2 Then Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against 4,000 matched control pages: Google AI Overviews minus 4.6% (small but statistically significant), AI Mode +2.4% and ChatGPT ~0, neither significant. Four separate statistical tests pointed the same way.
Ahrefs publishes its own limitation: every page already had 100-plus citations, so this tests whether schema lifts an already-cited page, not whether it helps an invisible page get discovered. Separately, searchVIU tested five assistants during direct real-time fetch and found none of them used JSON-LD, hidden Microdata or hidden RDFa; all extracted visible HTML only. Keep schema for rich results, entity clarity, and Bing, which has been explicitly more schema-friendly. Do not sell it as an AI citation lever.
Ahrefs' correlations across 75,000 brands: YouTube mentions 0.71 to 0.74, branded web mentions 0.66 to 0.71, branded anchors 0.527, Domain Rating 0.326, referring domains 0.295, backlinks 0.218, and number of site pages 0.17, which is effectively no relationship between content volume and AI visibility. Mention frequency beat mention reach, so twenty small creators may register more than one viral video.
The methodology flag Ahrefs discloses: brands were defined as domains above DR 40 whose top keyword has 800-plus monthly volume. That filter selects for large established brands and almost certainly inflates the correlation between "is a big brand" and every other variable. Directional, not a coefficient to plan against.
Be honest with clients about the economics. Pew Research, using metered browsing behavior from 900 US adults rather than vendor SERP scraping, found users who encountered an AI summary clicked a traditional result in 8% of visits versus 15% without one. Seer Interactive found being cited lifts click-through from roughly 0.6% to 1.08%, which is a real doubling of a very small number. And Google still sends roughly 345 times more traffic than ChatGPT, Gemini and Perplexity combined.
AI citation is a brand and defensive play. It is not yet a traffic channel.
Claims checked against primary sources and found wanting. Several of these are actively sold in Boulder right now.
| Claim | Status |
|---|---|
| Average word count of the top 10 is a target | False. Google has denied word count repeatedly. Danny Sullivan: "not a thing, it doesn't exist" |
| Optimal keyword density is 1 to 2% | False. Mueller, January 2023: Google has no notion of optimal keyword density. No Google origin for the figure |
| TF-IDF optimization | False. Mueller calls it an old metric. Search Engine Land: "Can you optimize for it? No" |
| Illyes confirmed topical authority is a ranking factor, January 2025 | No primary source exists. Widely repeated with no citation. Treat as fabricated |
| Link velocity is a ranking signal | Not supported. Ahrefs looked across 10,000 SERPs and found no observable relationship |
| Information gain is a Google ranking factor | Overreach. The patent claims a per-user assistant deduplication mechanism, not a corpus-wide signal |
| Optimal AI passage length is 134 to 167 words | Unsourced. No dataset, sample size, method or author attached anywhere |
| Content decays 1.21% per week | Unsourced. Traces to vendor marketing |
| llms.txt improves AI visibility | Contradicted. 97% of files got zero requests in May 2026 across 137,000 domains' logs |
| Schema improves AI citations | Contradicted by the only controlled test. Minus 4.6% on AI Overviews |
| E-E-A-T is something you add to pages | False. Google states rater data is not used directly in ranking algorithms |
| 93% of local-intent searches trigger a Local Pack | No traceable source. Recycled across dozens of 2026 statistics pages |
| 46% of all Google searches have local intent | Traces to a reported Google remark, not a study |
| 73% of queries have mixed intent | A corruption of Broder's 2002 finding that 73% were informational |
| Position 1 pages have a 10% higher CWV pass rate | Unsourced. No linked methodology anywhere |
| Google updated the Quality Rater Guidelines in June 2026 | False. The live PDF is dated 11 September 2025 and its change log ends there |
A "position-controlled replication" study that circulates in this space, attributed to "Lee (2026c)," traces to preprints hosted on a server for AI-generated research with no verifiable authorship or peer review. Its conclusion happens to align with a legitimate benchmark, which makes it more tempting to cite, not less. Do not.
Google disabled the num=100 URL parameter around 10 to 12 September 2025. Rank-tracking bots had been generating impressions across positions 1 to 100 in a single request, so removing it removed the bot impressions. Across 319 analyzed properties, 87.7% of sites lost Search Console impressions and 77.6% lost unique ranking terms. Clicks were unaffected, average position improved arithmetically, and click-through rate rose mechanically.
Any longitudinal analysis crossing that date is contaminated. Use October 2025 as the new baseline. This is also a free credibility play, since most agencies quietly let clients believe that September cliff was somebody's fault.
Assumes one keyword, top 20 pages, DataForSEO pay-as-you-go, cheap models for the language work.
| Step | What happens | Cost |
|---|---|---|
| 0 | Expand the keyword into 12 to 15 fan-out sub-queries. Run two different generators and treat the overlap as the priority signal | $0.001 |
| 1 | Fetch the real SERP at depth 20 with city-level geo-targeting, AI Overview forced live, People Also Ask expanded | $0.007 |
| 2 | Pull the ranking pages' content and on-page metrics, 20 URLs | $0.006 |
| 3 | Chunk each page at heading boundaries, embed, and build the coverage matrix: 20 pages by 15 sub-queries, each cell yes, partial or no | $0.00 local |
| 4 | Score structure and citation vulnerability per page. Flag any competitor over 2,000 words as structurally diluted | $0.011 |
| 5 | Layer competitive metrics: SERP competitors, search intent, referring domains, Lighthouse | $0.067 |
| 6 | Synthesise the plan. Score every gap on intent times extractability times consistency opportunity | $0.02 |
| 7 | Baseline and re-measure after 30 days, per platform, never blended | $0.01 |
Three configurations: lean at $0.02 (enough for the free redacted Snapshot), standard at $0.13 (the right depth for a paid Sprint), full at $0.58 (only for commercially load-bearing terms). A hundred keywords at standard depth costs about $13.
1. The coverage matrix. Twenty pages by fifteen fan-out sub-queries. Every practitioner describes this artifact. Nobody ships it.
2. Rank-conditional GEO advice. The same tactic helps a position-five page and hurts a position-one page, per the GEO paper's third table. No tool conditions on where the client already ranks.
3. A grounding-budget-aware brief. One that targets under 1,500 words on purpose and competes for share of a fixed 2,000-word pie. Every content tool in the market currently optimizes in the wrong direction on this.