SEO Believer Internal method note · not for client distribution yet

Research and method spec

The SERP Teardown Method

How to take apart the top 20 results for a search term, and turn what you find into a defensible plan to reach the top 10. Plus an honest answer to whether you should buy a tool for this or build one. The full underlying research, including every tool price, study and folklore claim checked, is in the research appendix.

Researched 26 July 2026 Four parallel research threads Live test keyword: SEO consultant Boulder
01

The short version

Buy or build
Build
But not for cost. For control and for a method you can publish.
Cost per teardown
6 to 8¢
On your existing DataForSEO account. Clearscope charges $6.45.
Your current stack
One sentence
That is the entire SERP analysis capability across all four SEO skills.
Before you build
$33/mo
Rent Thruuu first to find out if your brief format is even right.

Four research threads ran in parallel: the commercial tool market, the open source and DIY layer, the evidence on what actually moves a page into the top 10, and the evidence on search intent and SERP diagnosis. The findings converge on something uncomfortable and useful: most of what this category sells is a thin wrapper around data you already pay for, and several of the things it sells hardest are contradicted by the only controlled tests anyone has run.

The one-line recommendation

Stand up OpenSEO against your existing DataForSEO key this week (free, MIT, ships its own MCP server and Claude Code skills). Rent Thruuu at $33/mo for sixty days to pressure-test your brief format against a real product. Then build the SEO Believer teardown skill around the schema in section 5, because that is the part nobody sells and the part that makes the business its own proof of concept.

02

The hole in your current stack

You have four SEO skills installed. Here is what each one actually does:

SkillWhat it doesLooks at the SERP?
seo-analysisAudits your own site through Search Console, PageSpeed, technical crawlNo
keyword-researchDiscovers and prioritizes termsNo
content-writerWrites posts and landing pagesOne sentence
geo-optimizerTunes a page for AI answer enginesNo

The one sentence, quoted in full from content-writer:

"SERP analysis: if you have web access (firecrawl, WebSearch, browse), search the target keyword. Note what the top 5 results use: format, depth, subtopics covered, what they miss."

That is the whole thing. Top 5, not top 20. No data pull, no scoring, no gap quantification, no link-authority read, no page-type classification, no kill criteria. It is a reminder to look, not a method.

This matters more than a normal tooling gap, because looking at a SERP and explaining how to win it is the product. It is the free Snapshot, and it is the first half of the $4,500 Sprint. The capability sitting closest to the revenue is the one with the least machinery behind it.

Also worth fixing

research/keyword-targets.md is built entirely on estimated volume and difficulty. It says so, and adds that real numbers would arrive "once a paid tracking tool is in place." That tool is now in place and connected. That file is due a rebuild on measured data.

03

A live teardown of your own keyword

I ran the method against SEO consultant Boulder, geo-located to Boulder, desktop, on 26 July 2026. The result is a good demonstration of why "study the top pages and write a better one" is the wrong instinct.

Read this as one capture, not as the SERP

Everything below is a single pull from one location, one device, one moment. SERPs move, sometimes weekly. Before any of this goes in front of a paying client, re-pull it on three separate days across desktop and mobile and report how much the slot count, the page-type mix and the median link bar actually moved. At four thousandths of a cent per capture there is no excuse not to. A $4,500 engagement should not rest on a sample size of one.

Page one is not ten organic slots

It is a three-result local pack, then nine organic results, then related searches. Here is what occupies those nine:

PosResultWhat kind of page
1LinkedIn company page, Boulder SEO MarketingSocial profile
2boulderseomarketing.com homepageCompetitor site
3Thumbtack, SEO consultants near BoulderDirectory
4socialseo.com/boulder-seoCompetitor site
5Yelp, Boulder SEO MarketingDirectory
6taylorscherseo.com, "Best 5 Boulder SEO Companies That Don't Suck in 2026"Third-party listicle
7Facebook, Boulder SEO MarketingSocial profile
8lessermedia.com/seo-consulting-services/boulder-coCompetitor site
9Clutch, Top SEO Companies in BoulderDirectory

What that actually tells you

Five of nine slots are directories and social profiles. One is somebody else's listicle. Only three are service pages. That is a strong prior against a service page, not a hard ceiling of three: a service page can displace a directory if Google's read of relevance changes. But you should plan against the prior. And Chris Raulf already holds four of the nine through his own site plus his LinkedIn, Yelp and Facebook profiles. Site diversity matters here: Google has said you usually will not see more than two results from the same site, though it is a tendency rather than a hard cap. Profiles on other people's domains do not count against it at all, so he has effectively built a moat out of third-party presence.

The strategic read, which no content tool would produce

There are four routes into this top 10, and only one of them is a better service page. Get listed and ranked on Clutch and Thumbtack. Build the LinkedIn, Yelp and Facebook profiles that occupy those slots. Get into Taylor Scher's listicle, since he ranks a competitor roundup and is reachable. Or build the page. A single SERP capture cannot tell you which is genuinely fastest, because they differ in cost, control and eligibility, so treat this as four hypotheses to sequence and test rather than a ranked list. What it does tell you is that three of the four are invisible to a content optimizer, which would have recommended only the fourth.

The incumbent page, pulled apart in one API call

One call against Raulf's /boulder-seo/ page returned the full heading tree, 5,359 words, 43 internal links, 13 external, readability indices, title-to-content consistency of 0.875, an on-page score of 95.24, and a last-modified date of 28 November 2025. His H3s are the interesting part, because they are the objection list he built the page around: We tried SEO before and got burned. Our Boulder competitors dominate Google and we can't figure out why. We hear about AI search changing everything. We don't have time to become SEO experts.

Two things to be precise about. That URL is his local-pack landing page, not one of the nine organic results above; his organic entry is the homepage. And at 5,359 words it is a long page, which sits in tension with the extraction research in section 6. Both can be true: length appears to work against being quoted by an AI answer, and it clearly is not stopping him ranking. Those are different systems, and section 6 keeps them apart.

Either way, that heading tree is the content brief, and it took one request costing fifteen thousandths of a cent.

04

Buy or build, with real numbers

Cost per brief, verified against live pricing pages

OptionEntry pricePer briefReadsVerdict
Clearscope$129/mo$6.45Top 10Skip
Frase$39/mo$3.90Top 10Skip
NeuronWriter$57/mo$0.76Top 10Skip
Thruuu$33/mo$0.47Top 100Rent first
Surfer$99/mo$0.28Top 10Skip
DataForSEO pipelinepay as you go$0.06 to $0.08Any depthBuild

The DIY figure is real arithmetic, not a guess. Costed end to end at three depths:

DepthCost per keywordWhat it includes
Lean$0.02SERP plus AI Overview, parse 20 pages, cheap LLM synthesis. Enough for the free redacted Snapshot
Standard$0.13Adds fan-out SERPs, embeddings, structural scoring, SERP competitors, search intent, backlinks, Lighthouse. The right depth for a paid Sprint
Full$0.58Adds ranked keywords for three rivals and page intersection. Only for commercially load-bearing terms

At Standard depth, a hundred keywords costs about $13. Against a $4,500 Sprint the data cost is a rounding error, which is the real point: the entire economic argument sits in the analysis quality, not the API bill. One line item dominates and should stay behind a flag, which is pulling ranked keywords for rival domains at roughly $0.40 a keyword.

One caveat on that per-brief column, because the units are not the same

Surfer's $0.28 divides $99 by 360 "documents created or optimized", while Clearscope's $6.45 divides $129 by 20 "drafts" and Thruuu's $0.47 is a published per-brief price. Those are three different products of work. The comparison is directionally sound, since the spread is two orders of magnitude and no unit definition closes that, but do not quote the individual figures to a client as like-for-like.

The honest caveat, and it is the whole decision

The API bill is not the decision variable. The engineering time is. A pipeline that hits five endpoints, normalizes twenty pages of parsed HTML, survives JavaScript-only pages and failed fetches, and prompts an LLM into a consistent format is realistically 10 to 20 hours to build plus ongoing breakage. At your consulting rate that is 15 to 30 months of Surfer. Build for control and differentiation, not to save money.

What is genuinely differentiated, and what is a wrapper

Three things in this market cannot be rebuilt on commodity data:

Everything else in the content-optimizer category is the same loop: fetch the top 10, count words, extract headings, compute term frequency, render a score out of 100. Semrush's On Page SEO Checker is exactly that computation behind a $139 to $549 a month interface.

The tell

Profound's own integrations page lists DataForSEO as its partner for real-time SERP data. The billion-dollar category leader buys its SERP layer from the same commodity API already wired into your repo. Separately, OpenSEO's README states its hosted business model outright: it charges 28% on top of every DataForSEO request. That is the entire value-add of a wrapper business, admitted by the wrapper.

05

What to actually extract from the top 20

This is the part worth owning, because the evidence says most tools measure the wrong things. Fields are grouped by whether the research supports them.

Capture per SERP

Capture per result

Do not capture, or capture only as curiosity

Average word count of the top 10. Google has denied word count as a factor repeatedly. It is also a summary statistic over a bimodal distribution: a SERP holding one 400-word definition page and one 6,000-word guide has a mean that describes nothing that exists.

Keyword density and TF-IDF targets. John Mueller said in January 2023 that Google has no notion of optimal keyword density. The 1 to 2% figure has no Google origin.

A content score to beat. Surfer's own published study, across 10,000 queries and a million SERP entries, puts the correlation between their Content Score and rankings at Spearman 0.28. That is a weak relationship. Squaring it gives about 0.08, though be careful quoting that as "variance explained", since Spearman works on ranks and does not carry the ordinary least-squares interpretation that phrase implies. It is also correlational and close to circular: the score is computed from the pages currently ranking, so correlating it against current rank partly correlates a variable with its own input.

Ahrefs ran the only independent head-to-head across five tools and found "weak correlations everywhere." Their most damaging sentence, from inside the category: "in some tools, you can literally copy-paste the entire list, draft nothing else, and get an almost perfect score." Their summary line is the one to keep for skeptical prospects: keyword density is not topic coverage. Use coverage tools as a checklist for missed subtopics, never as a target number.

Meta description keyword usage. A quarter of top-ranking pages have no meta description at all.

06

Turning the teardown into a top-10 plan

Nine steps, and the first four are gates: if any fails you stop, change the target, or switch to a different playbook entirely. Saying so is the most valuable thing in the deliverable.

  1. Gate zero: what kind of SERP is this, and what does the query want?

    Before anything else, classify two things. Query intent (informational, commercial investigation, transactional, navigational, or local) determines which page types can rank at all and whether an AI Overview will eat the clicks. And SERP shape: is this organic-dominant, local-pack-dominant, AI-Overview-dominant, shopping, or big-publisher territory?

    This routes everything downstream. A local-pack-dominant SERP does not run this playbook at all, it runs the one in section 7. A navigational SERP for somebody else's brand is usually not winnable at any effort, and you should say so rather than sell a plan. Do not skip this because the keyword "looks" commercial. On my own Boulder run, DataForSEO classified "Boulder SEO" as navigational, and the SERP confirmed it: five of the first seven results were one competitor's brand assets.

    One caution from Google's own rater guidelines, section 12.2: "If you research the query on Google, please do not rely on the top results on the SERP. A query may have other meanings not represented on Google's search results pages." Google tells its own raters not to read intent off the SERP. It is still the best proxy available. It is a proxy.

  2. Gate one: can your page type rank here at all?

    Classify all twenty results by archetype. If eight or more of the top ten are a format you cannot credibly produce, stop. If the mix is genuinely varied, a differently-typed page can hold a slot. Be careful with the strong version of this rule though, because it is folklore: BrightLocal's data shows business websites still holding 53% of the top ten for "best chiropractor" and 63% for "best vet clinic". Mismatch is a heavy prior against you, not a hard block.

  3. Gate two: is the link bar reachable in twelve months?

    Take the median referring domains across the organic ten, counted per ranking URL rather than per domain, since that is the unit that tells you whether a specific page is displaceable. Then read link quality too, not just count: topical relevance, local relevance, and whether the linking pages themselves rank. Ten editorial links from Colorado business media are not ten directory profiles.

    For a site under DR 30 these are the working bands. They are calibration heuristics, not measured win rates, and the honest move is to replace them with your own outcomes as engagements complete: 0 to 9 winnable on content and internal linking alone, 10 to 19 needs a handful of real links, 20 to 29 genuinely uncertain, 30 to 49 mostly fails as a single-page attempt, 50-plus should not be chased directly. Chase the long tail underneath instead and let cluster strength accrue.

  4. Gate three: can the page even be indexed and served?

    Cheap, boring, and skipped constantly. A page can pass both gates above, target a genuinely winnable SERP, and still fail on crawlability, a stray noindex, a canonical pointing somewhere else, JavaScript-only content, or simply being orphaned with no internal links pointing at it. Check indexation and serving before writing a ranking plan, not after three months of wondering. Google's AI features have the same prerequisite: the page must be indexed and snippet-eligible or it cannot be cited at all.

  5. Count the slots that are actually available

    Subtract the directories, social profiles and third-party listicles you cannot displace with a page. What remains is your real target count. On SEO consultant Boulder that number is three, out of nine. Then ask the separate question the content tools never ask: can you occupy the non-page slots instead, by getting listed, profiled or included?

    And then answer the follow-up question, which is the one that actually gets it done. "Get listed on Clutch" is not a plan. Each of those platforms has its own internal ranking system, and you need the specifics: what review count and velocity moves position, how complete the profile has to be, whether there is a paid inclusion tier and whether it is worth it, and what the category and service-line taxonomy rewards. Treat every directory that holds a slot as its own small SERP with its own algorithm. For third-party listicles the mechanism is different again and simpler: the author is a person, they are reachable, and being genuinely worth including is most of the work.

  6. Build the coverage gap, not the word count gap

    Take the union of every H2 and H3 across the top 20 and subtract what your page covers. That set, ranked by how many of the twenty results address each item, is the legitimate core of what content optimizers sell.

    Then add the layer nobody ships. Google confirms AI Overviews issue a fan of related sub-queries, so expand the keyword into 12 to 15 of them, pull the real SERP for five, and build a coverage matrix: twenty pages down the side, fifteen sub-queries across the top, each cell scored yes, partial or no. Every serious practitioner describes this artifact and none of them sells it. iPullRank's Qforia publishes its fan-out prompt in full, so you can lift it straight into Claude. One useful refinement from DEJAN's research: feeding the actual SERP results into the fan-out prompt substantially beats prompting from the bare keyword.

  7. Find the angle nobody has

    The heading union tells you the table stakes. The differentiator is what none of the twenty have: original data, first-hand testing, named expertise, a local specific. On the Boulder SERP, not one of the nine results shows a real client outcome with numbers. You have Bone Voyage. That is the gap.

  8. Write for sentence-level extraction, and aim short on purpose

    Google's grounding extraction operates on individual sentences with a heavy bias toward opening paragraphs. Front-load the answer, write self-contained sentences that survive being lifted out of context, and keep bloated tables of contents out of the main content area where they compete for extraction slots.

    Then the finding that should change your word-count advice permanently. DEJAN analyzed 883,262 grounding snippets across 7,060 queries and found the average extracted chunk is 15.5 words, that there is a roughly fixed 2,000-word grounding budget per query regardless of how many sources are used, and that the budget is split by rank: the top source gets a median 531 words, the fifth gets 266. You are competing for share of a fixed pie, not expanding it.

    And coverage collapses as pages get longer:

    Page lengthWords actually groundedCoverage
    Under 1,000 words37061%
    1,000 to 2,00049235%
    2,000 to 3,00053222%
    3,000 plus54413%

    In that dataset, grounded words rose only slightly past about 540 while the share of the page being used collapsed. For extraction, density beats length.

    Do not turn this into a word-count rule, which is the mistake I nearly made here

    This is evidence about what gets quoted into an AI answer. It is not evidence about what ranks. A lower share of a long page being cited does not show that shorter pages rank better, convert better, or serve the reader better, and the incumbent in section 3 ranks first with 5,359 words. Run the two as parallel tracks: set organic length by what the task and the winning archetypes actually require, and win extraction by building self-contained, front-loaded answer passages inside whatever length that turns out to be. A blanket 1,500-word ceiling would walk a client into a thin service page on a query that wants proof density.

    Same caution on the diagnostic: a competitor over 2,000 words is a candidate for redundancy, not proven diluted. Check its passages before claiming an opening.

  9. Audit for demotions before adding boosts

    Anchor mismatch, exact-match-domain demotion, nav demotion, conflicting date signals. Checking whether you are being held down is usually higher-yield than adding another optimisation.

Label every rule with the evidence behind it

The failure mode this method is most exposed to is presenting useful hypotheses as validated rules. The fix is cheap: carry a grade next to each one, so a client can see which parts are load-bearing and which are working assumptions you intend to test. This is the same T1 to T4 scale used in the research appendix.

RuleEvidence behind itGradeRevisit
Gate zero: classify intent and SERP shape firstGoogle's rater guidelines define the intent taxonomy; SERP shape is directly observedT1Stable
Gate one: page-type mismatch is a heavy priorBrightLocal per-keyword page-type mix, plus the NavBoost mechanism from DOJ testimonyT3Annually
Gate two: median referring domains as the entry barDescriptive statistic. No study establishes it as a thresholdT3Replace with your own outcomes
Gate three: indexation before planningGoogle documentation, prerequisite for both ranking and AI citationT1Stable
Internal linking lifts trafficSearchPilot controlled split tests, plus 5 to 7% measuredT2Annually
Front-load answers for extractionObserved positional bias in grounding snippets. An empirical pattern, not a confirmed mechanismT3Every 6 months
DR bands for a low-authority sitePractitioner heuristic calibrated from referring-domain benchmarksT4Replace with your own outcomes
Rank-conditional GEO effectsOne simulated-engine study, unreplicated on production systemsT3Every 6 months

The two T4 and T3-heuristic rows are the honest weak points, and naming them is what separates this from a sales document. Every band in the link table should be replaced with your own measured outcomes as engagements complete, which turns the weakest row into the most defensible one within about a year.

Kill criteria, stated plainly in the deliverable

A plan that never says "do not chase this" is a sales document. Fire these:

The timing fact that should shape every engagement

Only 1.74% of newly published pages reach the top 10 within a year for even one keyword. But among pages that do make it, 40.82% got there within one month. The distribution is fast-or-never, not a slow grind. Combined with SE Ranking's finding that 84.8% of top-three entrants came from pages already in the top 20, the operational conclusion is clear: improving near-misses beats publishing new pages, and a page that has not moved in six months needs a different keyword, not more patience.

07

When the SERP is a local pack

Everything above is an organic playbook. For a Front Range service business it is frequently the wrong one, and the demonstration query proves it: SEO consultant Boulder puts a three-result local pack above every organic result, and the method as written walks straight past it.

The rule

If a local pack sits above the fold, the map pack is the plan and organic is the supporting act, not the other way round. Do not run the organic teardown and bolt local on as an afterthought. And never use a competitor's organic ranking as a proxy for their map-pack ranking. They are separate systems with different inputs, and a firm can dominate one while being invisible in the other.

What actually determines pack position

Three inputs, and only one of them is on your website.

The audit, in order

  1. Eligibility and verification first

    Is the profile verified, unsuspended, and eligible for the category at all? Service-area businesses have different rules from storefronts. This is the gate, and it fails more often than people expect.

  2. Primary category, then secondary

    The primary category does most of the work and is the highest-leverage single field in the whole profile. Pull the primary category of every business in the pack and compare. If all three competitors share a category the client does not have, that is usually the finding of the engagement.

  3. Measure on a grid, never at a point

    Because proximity dominates, a single coordinate tells you almost nothing. Sample rank across a grid of coordinates spanning the actual service area and report the map, not the number. The honest deliverable is "you rank 2 to 4 within two miles and drop out entirely past five," which is a completely different conversation from "you rank 3."

  4. Reviews as a rate, not a total

    Count, average rating, velocity over the last 90 days, recency of the most recent, whether the owner responds, and whether review text mentions the services you want to rank for. Compare against the pack incumbents rather than an abstract target. A competitor with 28 reviews is a very different problem from one with 400.

  5. NAP consistency and the citation set

    Name, address and phone consistency across the directories that matter in the market. This overlaps directly with the directory work in the organic plan, which is convenient: the same Clutch, Thumbtack and Yelp placements that occupy organic slots also feed prominence here. One effort, two systems.

  6. The landing page the profile points at

    This is the only place the organic and local playbooks genuinely converge. The page the profile links to should be locally relevant and match the primary category. Note that this is exactly what the Boulder incumbent does: the page I pulled apart in section 3 is his local-pack landing page, not his organic entry.

  7. Check for competitor spam, and decide whether to report it

    Keyword-stuffed business names, fake addresses, and lead-gen listings are common in local service categories and they are reportable. Sometimes the highest-return action in a local engagement is a redressal filing rather than anything you build.

And measure the right outcome

Pack rank is a means. The deliverable metrics are calls, direction requests, bookings and qualified leads, all of which are available in the profile's own performance data. A client who moves from position 3 to position 2 and gets no more calls has not been helped, and you will only know that if you were measuring calls.

08

After you reach the top 10

The method above ends at arrival, which is a real gap: a page that reaches position 8 and falls to 15 within six weeks is a failed engagement, not a successful one. Two numbers set the stakes.

Recovery after a loss
32.2%
of domains that lost top-10 positions in the March 2026 core update had recovered by May. Recovery is the exception.
Where entrants come from
84.8%
of top-three entrants came from pages already in the top 20. Defending a position is cheaper than re-earning it.

The maintenance protocol

Two deliverables, not one

Keep the organic ranking plan and the AI-visibility plan as separate documents with separate KPIs and timelines, even though one teardown feeds both. They are different systems, they move independently, and merging them is how you end up optimizing a page for extraction in ways that cost it rank.

Organic plan: position, clicks, conversions, measured in Search Console. AI-visibility plan: citation share and appearance frequency across repeated runs, per platform, never blended into one score and never reported as a rank. Different timelines too, since organic moves on core updates and AI citation moves whenever a model ships.

09

What not to sell, because the evidence kills it

This section is worth more than the rest combined, because every competitor in Boulder is selling at least two of these.

ClaimWhat the evidence actually says
Schema improves AI citations Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 matched controls. Result: minus 4.6% on AI Overviews, statistically significant, and nothing on AI Mode or ChatGPT. Correlational studies find a small positive because sites that add schema do everything else well too. Keep schema for rich results and Bing. Do not sell it as an AI lever.
llms.txt helps AI visibility 97% of llms.txt files received zero requests in May 2026, across 137,000 domains' server logs. The top requesters are the SEO audit tools checking whether the file exists. Google states it ignores them outright. You already left this off Ben's page. The evidence caught up with you.
Link velocity is a ranking signal Ahrefs went looking across 10,000 SERPs and states plainly they "failed to distill any observable relationship." Velocity appears in the leaked documentation only as a spam signal.
Core Web Vitals will lift your rankings Google's own FAQ: "trying to get a perfect score just for SEO reasons may not be the best use of your time." Fix genuinely broken LCP and INP. Do not chase 100s.
We track your rank in ChatGPT SparkToro ran 12 identical prompts 2,961 times across 600 volunteers. Odds of the same brand list appearing twice: under 1 in 100. Same list in the same order: under 1 in 1,000. Any tool selling a rank position in an AI answer is selling noise. Report a visibility percentage across many runs, never a rank.
Rank top 10 and the AI will cite you Ahrefs' AI Overview citation overlap with the top 10 fell from 76% to 38% between July 2025 and March 2026, though they honestly flag that their own parser improved and the datasets are not comparable. Off Google it is far weaker: only 12% of URLs cited by ChatGPT, Gemini and Copilot appear in Google's top 10.
Read the vendor numbers with suspicion, including the ones above

Nearly all AI-citation research comes from companies selling AI visibility tools. On the single question of how often AI Overview citations come from the top 10, Ahrefs says 38%, BrightEdge says 17%, seoClarity says 33%, and Originality says 52%. Same question, threefold spread. Treat the direction as real and the decimals as unusable in a client deck. The field's most-repeated numbers are mostly denominator errors.

And one thing worth selling that almost nobody is

The strongest correlate anyone has measured for AI visibility is YouTube brand mentions, at roughly 0.71 to 0.74, against backlinks at 0.218. Frequency of mention beat reach of mention, meaning twenty small creators mentioning a client may register more than one viral video. Unlike link building, this is not gated by domain authority, which makes it the single best lever available to a small local business. Caveat honestly: Ahrefs' sample was filtered to domains above DR 40, which inflates the correlation between "is a big brand" and everything else. Directional, not a coefficient to plan against.

The finding hiding in plain sight, and the best sales argument you have

The Princeton GEO paper is widely cited for its headline, that adding statistics, quotations and cited sources lifts visibility by up to 40%. Almost nobody quotes its third table, which breaks the same result out by where the page already ranked:

TacticRank 1Rank 3Rank 5
Cite sources-30.3%+20.4%+115.1%
Add quotations-22.9%+3.5%+99.7%
Add statistics-20.6%+8.1%+97.9%

These tactics roughly double visibility for the weakest source in the set and measurably hurt the one already in front. It is a leveling effect, not a universal uplift. That is a genuinely good argument to make to a Boulder business sitting at position four or five behind an entrenched incumbent, and it is also the reason no off-the-shelf tool gives correct advice here: not one of them conditions its recommendations on where the client already ranks. Yours can.

Keep the caveat with it. The study ran on a simulated engine over 2023-era results, "visibility" means share of words in the generated answer rather than clicks, and nobody has replicated it on 2026 engines. Also note keyword stuffing was the one tactic that scored below doing nothing.

10

Two traps in the data layer

The depth parameter does not do what you expect

Asking DataForSEO for depth 20 returns page one only, because max_crawl_pages defaults to 1. Depth also multiplies cost per ten results, and each search operator such as site: multiplies the task cost by five. Budget accordingly.

Any trend line crossing September 2025 is contaminated

Google disabled the num=100 URL parameter around 10 to 12 September 2025. Rank-tracking bots had been generating impressions across positions 1 to 100 in a single request, so removing it removed the bot impressions. Across 319 analyzed properties, 87.7% of sites lost Search Console impressions and 77.6% lost unique ranking terms. Clicks were unaffected, average position "improved" arithmetically, and click-through rate rose mechanically. Use October 2025 as the new baseline, and if a client shows you a chart with a cliff last September, that is what it is.

A free credibility play

Most Boulder agencies will have quietly let clients believe that September drop was their fault or someone else's failure. Explaining it correctly, with the mechanism, is a five-minute conversation that positions you as the person who actually reads the primary sources.

11

Recommended sequence

  1. This week: OpenSEO against your existing key

    Free, MIT licensed, self-hostable on Cloudflare's free plan, bring-your-own DataForSEO key, ships its own MCP server plus prebuilt Claude Code skills. Roughly 8,100 stars and actively maintained. An afternoon of setup for keyword research, rank tracking, competitor insight, backlinks, site audit and AI visibility. Highest leverage item in this entire report.

  2. This week: Screaming Frog, $279 a year

    The cheapest deep page extraction that exists, and version 24 ships a built-in MCP server. Two caveats: it cannot fetch a SERP at all (its "SERP mode" is a snippet pixel-width preview, not results retrieval), so you feed it URLs from DataForSEO. And list mode is not yet supported over MCP, which kills the obvious "hand Claude twenty competitor URLs" workflow for now.

  3. Next sixty days: rent Thruuu at $33 a month

    Reads up to 100 results, has both a SERP API and a Brief API, costs $0.47 a brief. Use it to find out whether your brief format is right before spending fifteen hours building one. If it turns out you need things Thruuu does not carry, which it will, you will then build the right thing instead of the obvious thing.

  4. Then: build the SEO Believer teardown skill

    Around the schema in section 5 and the seven steps in section 6. The reason to build is not cost. It is that you own the prompt, so the deliverable carries your methodology and your voice rather than a vendor's Content Score, and you can fold in link authority, keyword intersection and AI Overview citation data that no per-brief tool exposes. It is also the only thing here that makes SEO Believer its own proof of concept while every competitor resells the same eight dashboard logos.

  5. Only when a client needs it: Ahrefs Lite at $129

    For the backlink index, which you genuinely cannot rebuild, and because clients recognize the DR number. Its MCP is included on every paid tier. Do not stack Content Kit and Brand Radar on top, which takes it to $427 a month for things you can build.

  6. Rebuild keyword-targets.md on measured data

    Every number in it is currently an estimate, by its own admission. Run the real volumes, difficulties and SERP compositions, and record which Tier 1 terms fail the two gates in section 6. Some of them will.

The build stack, if and when you build

If you would rather buy something agent-drivable

Only two tools in this entire market can be driven from Claude Code today. Frase ships an official MCP server that is pushed to daily, plus 50-plus documented REST endpoints, from $39 a month. NeuronWriter ships an official MCP server and a public six-method API at $57 a month yearly, and its evaluate-content method scores a draft without consuming credits, which is the cheapest correct scoring primitive available anywhere. NeuronWriter also happens to be the one tool that scored best in Ahrefs' independent cross-tool test. Surfer's own API is locked to the $299 tier and its MCP is announced with no date.

Three things worth building that genuinely nobody sells

1. The coverage matrix. Twenty pages by fifteen fan-out sub-queries, scored yes, partial or no. Every practitioner describes it. Nobody ships it.

2. Rank-conditional GEO advice. The same tactic helps a position-five page and hurts a position-one page. No tool conditions on this.

3. A grounding-budget-aware brief. One that targets under 1,500 words on purpose and competes for share of a fixed 2,000-word pie. Every content tool in the market currently optimizes in the wrong direction on this.

One positioning caution to carry into client work

With only 38% of AI Overview citations and 12% of ChatGPT, Gemini and Copilot citations coming from Google's top 10, do not sell "outrank the top 10" as an AI visibility deliverable. Sell it as what it is, an organic ranking deliverable, and price AI visibility separately as an honestly caveated measurement product. Google still sends roughly 345 times more traffic than ChatGPT, Gemini and Perplexity combined. The traffic math still says win the blue link.