Research and method spec
How to take apart the top 20 results for a search term, and turn what you find into a defensible plan to reach the top 10. Plus an honest answer to whether you should buy a tool for this or build one. The full underlying research, including every tool price, study and folklore claim checked, is in the research appendix.
Four research threads ran in parallel: the commercial tool market, the open source and DIY layer, the evidence on what actually moves a page into the top 10, and the evidence on search intent and SERP diagnosis. The findings converge on something uncomfortable and useful: most of what this category sells is a thin wrapper around data you already pay for, and several of the things it sells hardest are contradicted by the only controlled tests anyone has run.
Stand up OpenSEO against your existing DataForSEO key this week (free, MIT, ships its own MCP server and Claude Code skills). Rent Thruuu at $33/mo for sixty days to pressure-test your brief format against a real product. Then build the SEO Believer teardown skill around the schema in section 5, because that is the part nobody sells and the part that makes the business its own proof of concept.
You have four SEO skills installed. Here is what each one actually does:
| Skill | What it does | Looks at the SERP? |
|---|---|---|
seo-analysis | Audits your own site through Search Console, PageSpeed, technical crawl | No |
keyword-research | Discovers and prioritizes terms | No |
content-writer | Writes posts and landing pages | One sentence |
geo-optimizer | Tunes a page for AI answer engines | No |
The one sentence, quoted in full from content-writer:
"SERP analysis: if you have web access (firecrawl, WebSearch, browse), search the target keyword. Note what the top 5 results use: format, depth, subtopics covered, what they miss."
That is the whole thing. Top 5, not top 20. No data pull, no scoring, no gap quantification, no link-authority read, no page-type classification, no kill criteria. It is a reminder to look, not a method.
This matters more than a normal tooling gap, because looking at a SERP and explaining how to win it is the product. It is the free Snapshot, and it is the first half of the $4,500 Sprint. The capability sitting closest to the revenue is the one with the least machinery behind it.
research/keyword-targets.md is built entirely on estimated volume and difficulty. It says so, and adds that real numbers would arrive "once a paid tracking tool is in place." That tool is now in place and connected. That file is due a rebuild on measured data.
I ran the method against SEO consultant Boulder, geo-located to Boulder, desktop, on 26 July 2026. The result is a good demonstration of why "study the top pages and write a better one" is the wrong instinct.
Everything below is a single pull from one location, one device, one moment. SERPs move, sometimes weekly. Before any of this goes in front of a paying client, re-pull it on three separate days across desktop and mobile and report how much the slot count, the page-type mix and the median link bar actually moved. At four thousandths of a cent per capture there is no excuse not to. A $4,500 engagement should not rest on a sample size of one.
It is a three-result local pack, then nine organic results, then related searches. Here is what occupies those nine:
| Pos | Result | What kind of page |
|---|---|---|
| 1 | LinkedIn company page, Boulder SEO Marketing | Social profile |
| 2 | boulderseomarketing.com homepage | Competitor site |
| 3 | Thumbtack, SEO consultants near Boulder | Directory |
| 4 | socialseo.com/boulder-seo | Competitor site |
| 5 | Yelp, Boulder SEO Marketing | Directory |
| 6 | taylorscherseo.com, "Best 5 Boulder SEO Companies That Don't Suck in 2026" | Third-party listicle |
| 7 | Facebook, Boulder SEO Marketing | Social profile |
| 8 | lessermedia.com/seo-consulting-services/boulder-co | Competitor site |
| 9 | Clutch, Top SEO Companies in Boulder | Directory |
Five of nine slots are directories and social profiles. One is somebody else's listicle. Only three are service pages. That is a strong prior against a service page, not a hard ceiling of three: a service page can displace a directory if Google's read of relevance changes. But you should plan against the prior. And Chris Raulf already holds four of the nine through his own site plus his LinkedIn, Yelp and Facebook profiles. Site diversity matters here: Google has said you usually will not see more than two results from the same site, though it is a tendency rather than a hard cap. Profiles on other people's domains do not count against it at all, so he has effectively built a moat out of third-party presence.
There are four routes into this top 10, and only one of them is a better service page. Get listed and ranked on Clutch and Thumbtack. Build the LinkedIn, Yelp and Facebook profiles that occupy those slots. Get into Taylor Scher's listicle, since he ranks a competitor roundup and is reachable. Or build the page. A single SERP capture cannot tell you which is genuinely fastest, because they differ in cost, control and eligibility, so treat this as four hypotheses to sequence and test rather than a ranked list. What it does tell you is that three of the four are invisible to a content optimizer, which would have recommended only the fourth.
One call against Raulf's /boulder-seo/ page returned the full heading tree, 5,359 words, 43 internal links, 13 external, readability indices, title-to-content consistency of 0.875, an on-page score of 95.24, and a last-modified date of 28 November 2025. His H3s are the interesting part, because they are the objection list he built the page around: We tried SEO before and got burned. Our Boulder competitors dominate Google and we can't figure out why. We hear about AI search changing everything. We don't have time to become SEO experts.
Two things to be precise about. That URL is his local-pack landing page, not one of the nine organic results above; his organic entry is the homepage. And at 5,359 words it is a long page, which sits in tension with the extraction research in section 6. Both can be true: length appears to work against being quoted by an AI answer, and it clearly is not stopping him ranking. Those are different systems, and section 6 keeps them apart.
Either way, that heading tree is the content brief, and it took one request costing fifteen thousandths of a cent.
| Option | Entry price | Per brief | Reads | Verdict |
|---|---|---|---|---|
| Clearscope | $129/mo | $6.45 | Top 10 | Skip |
| Frase | $39/mo | $3.90 | Top 10 | Skip |
| NeuronWriter | $57/mo | $0.76 | Top 10 | Skip |
| Thruuu | $33/mo | $0.47 | Top 100 | Rent first |
| Surfer | $99/mo | $0.28 | Top 10 | Skip |
| DataForSEO pipeline | pay as you go | $0.06 to $0.08 | Any depth | Build |
The DIY figure is real arithmetic, not a guess. Costed end to end at three depths:
| Depth | Cost per keyword | What it includes |
|---|---|---|
| Lean | $0.02 | SERP plus AI Overview, parse 20 pages, cheap LLM synthesis. Enough for the free redacted Snapshot |
| Standard | $0.13 | Adds fan-out SERPs, embeddings, structural scoring, SERP competitors, search intent, backlinks, Lighthouse. The right depth for a paid Sprint |
| Full | $0.58 | Adds ranked keywords for three rivals and page intersection. Only for commercially load-bearing terms |
At Standard depth, a hundred keywords costs about $13. Against a $4,500 Sprint the data cost is a rounding error, which is the real point: the entire economic argument sits in the analysis quality, not the API bill. One line item dominates and should stay behind a flag, which is pulling ranked keywords for rival domains at roughly $0.40 a keyword.
Surfer's $0.28 divides $99 by 360 "documents created or optimized", while Clearscope's $6.45 divides $129 by 20 "drafts" and Thruuu's $0.47 is a published per-brief price. Those are three different products of work. The comparison is directionally sound, since the spread is two orders of magnitude and no unit definition closes that, but do not quote the individual figures to a client as like-for-like.
The API bill is not the decision variable. The engineering time is. A pipeline that hits five endpoints, normalizes twenty pages of parsed HTML, survives JavaScript-only pages and failed fetches, and prompts an LLM into a consistent format is realistically 10 to 20 hours to build plus ongoing breakage. At your consulting rate that is 15 to 30 months of Surfer. Build for control and differentiation, not to save money.
Three things in this market cannot be rebuilt on commodity data:
Everything else in the content-optimizer category is the same loop: fetch the top 10, count words, extract headings, compute term frequency, render a score out of 100. Semrush's On Page SEO Checker is exactly that computation behind a $139 to $549 a month interface.
Profound's own integrations page lists DataForSEO as its partner for real-time SERP data. The billion-dollar category leader buys its SERP layer from the same commodity API already wired into your repo. Separately, OpenSEO's README states its hosted business model outright: it charges 28% on top of every DataForSEO request. That is the entire value-add of a wrapper business, admitted by the wrapper.
This is the part worth owning, because the evidence says most tools measure the wrong things. Fields are grouped by whether the research supports them.
Average word count of the top 10. Google has denied word count as a factor repeatedly. It is also a summary statistic over a bimodal distribution: a SERP holding one 400-word definition page and one 6,000-word guide has a mean that describes nothing that exists.
Keyword density and TF-IDF targets. John Mueller said in January 2023 that Google has no notion of optimal keyword density. The 1 to 2% figure has no Google origin.
A content score to beat. Surfer's own published study, across 10,000 queries and a million SERP entries, puts the correlation between their Content Score and rankings at Spearman 0.28. That is a weak relationship. Squaring it gives about 0.08, though be careful quoting that as "variance explained", since Spearman works on ranks and does not carry the ordinary least-squares interpretation that phrase implies. It is also correlational and close to circular: the score is computed from the pages currently ranking, so correlating it against current rank partly correlates a variable with its own input.
Ahrefs ran the only independent head-to-head across five tools and found "weak correlations everywhere." Their most damaging sentence, from inside the category: "in some tools, you can literally copy-paste the entire list, draft nothing else, and get an almost perfect score." Their summary line is the one to keep for skeptical prospects: keyword density is not topic coverage. Use coverage tools as a checklist for missed subtopics, never as a target number.
Meta description keyword usage. A quarter of top-ranking pages have no meta description at all.
Nine steps, and the first four are gates: if any fails you stop, change the target, or switch to a different playbook entirely. Saying so is the most valuable thing in the deliverable.
Before anything else, classify two things. Query intent (informational, commercial investigation, transactional, navigational, or local) determines which page types can rank at all and whether an AI Overview will eat the clicks. And SERP shape: is this organic-dominant, local-pack-dominant, AI-Overview-dominant, shopping, or big-publisher territory?
This routes everything downstream. A local-pack-dominant SERP does not run this playbook at all, it runs the one in section 7. A navigational SERP for somebody else's brand is usually not winnable at any effort, and you should say so rather than sell a plan. Do not skip this because the keyword "looks" commercial. On my own Boulder run, DataForSEO classified "Boulder SEO" as navigational, and the SERP confirmed it: five of the first seven results were one competitor's brand assets.
One caution from Google's own rater guidelines, section 12.2: "If you research the query on Google, please do not rely on the top results on the SERP. A query may have other meanings not represented on Google's search results pages." Google tells its own raters not to read intent off the SERP. It is still the best proxy available. It is a proxy.
Classify all twenty results by archetype. If eight or more of the top ten are a format you cannot credibly produce, stop. If the mix is genuinely varied, a differently-typed page can hold a slot. Be careful with the strong version of this rule though, because it is folklore: BrightLocal's data shows business websites still holding 53% of the top ten for "best chiropractor" and 63% for "best vet clinic". Mismatch is a heavy prior against you, not a hard block.
Take the median referring domains across the organic ten, counted per ranking URL rather than per domain, since that is the unit that tells you whether a specific page is displaceable. Then read link quality too, not just count: topical relevance, local relevance, and whether the linking pages themselves rank. Ten editorial links from Colorado business media are not ten directory profiles.
For a site under DR 30 these are the working bands. They are calibration heuristics, not measured win rates, and the honest move is to replace them with your own outcomes as engagements complete: 0 to 9 winnable on content and internal linking alone, 10 to 19 needs a handful of real links, 20 to 29 genuinely uncertain, 30 to 49 mostly fails as a single-page attempt, 50-plus should not be chased directly. Chase the long tail underneath instead and let cluster strength accrue.
Cheap, boring, and skipped constantly. A page can pass both gates above, target a genuinely winnable SERP, and still fail on crawlability, a stray noindex, a canonical pointing somewhere else, JavaScript-only content, or simply being orphaned with no internal links pointing at it. Check indexation and serving before writing a ranking plan, not after three months of wondering. Google's AI features have the same prerequisite: the page must be indexed and snippet-eligible or it cannot be cited at all.
Subtract the directories, social profiles and third-party listicles you cannot displace with a page. What remains is your real target count. On SEO consultant Boulder that number is three, out of nine. Then ask the separate question the content tools never ask: can you occupy the non-page slots instead, by getting listed, profiled or included?
And then answer the follow-up question, which is the one that actually gets it done. "Get listed on Clutch" is not a plan. Each of those platforms has its own internal ranking system, and you need the specifics: what review count and velocity moves position, how complete the profile has to be, whether there is a paid inclusion tier and whether it is worth it, and what the category and service-line taxonomy rewards. Treat every directory that holds a slot as its own small SERP with its own algorithm. For third-party listicles the mechanism is different again and simpler: the author is a person, they are reachable, and being genuinely worth including is most of the work.
Take the union of every H2 and H3 across the top 20 and subtract what your page covers. That set, ranked by how many of the twenty results address each item, is the legitimate core of what content optimizers sell.
Then add the layer nobody ships. Google confirms AI Overviews issue a fan of related sub-queries, so expand the keyword into 12 to 15 of them, pull the real SERP for five, and build a coverage matrix: twenty pages down the side, fifteen sub-queries across the top, each cell scored yes, partial or no. Every serious practitioner describes this artifact and none of them sells it. iPullRank's Qforia publishes its fan-out prompt in full, so you can lift it straight into Claude. One useful refinement from DEJAN's research: feeding the actual SERP results into the fan-out prompt substantially beats prompting from the bare keyword.
The heading union tells you the table stakes. The differentiator is what none of the twenty have: original data, first-hand testing, named expertise, a local specific. On the Boulder SERP, not one of the nine results shows a real client outcome with numbers. You have Bone Voyage. That is the gap.
Google's grounding extraction operates on individual sentences with a heavy bias toward opening paragraphs. Front-load the answer, write self-contained sentences that survive being lifted out of context, and keep bloated tables of contents out of the main content area where they compete for extraction slots.
Then the finding that should change your word-count advice permanently. DEJAN analyzed 883,262 grounding snippets across 7,060 queries and found the average extracted chunk is 15.5 words, that there is a roughly fixed 2,000-word grounding budget per query regardless of how many sources are used, and that the budget is split by rank: the top source gets a median 531 words, the fifth gets 266. You are competing for share of a fixed pie, not expanding it.
And coverage collapses as pages get longer:
| Page length | Words actually grounded | Coverage |
|---|---|---|
| Under 1,000 words | 370 | 61% |
| 1,000 to 2,000 | 492 | 35% |
| 2,000 to 3,000 | 532 | 22% |
| 3,000 plus | 544 | 13% |
In that dataset, grounded words rose only slightly past about 540 while the share of the page being used collapsed. For extraction, density beats length.
This is evidence about what gets quoted into an AI answer. It is not evidence about what ranks. A lower share of a long page being cited does not show that shorter pages rank better, convert better, or serve the reader better, and the incumbent in section 3 ranks first with 5,359 words. Run the two as parallel tracks: set organic length by what the task and the winning archetypes actually require, and win extraction by building self-contained, front-loaded answer passages inside whatever length that turns out to be. A blanket 1,500-word ceiling would walk a client into a thin service page on a query that wants proof density.
Same caution on the diagnostic: a competitor over 2,000 words is a candidate for redundancy, not proven diluted. Check its passages before claiming an opening.
Anchor mismatch, exact-match-domain demotion, nav demotion, conflicting date signals. Checking whether you are being held down is usually higher-yield than adding another optimisation.
The failure mode this method is most exposed to is presenting useful hypotheses as validated rules. The fix is cheap: carry a grade next to each one, so a client can see which parts are load-bearing and which are working assumptions you intend to test. This is the same T1 to T4 scale used in the research appendix.
| Rule | Evidence behind it | Grade | Revisit |
|---|---|---|---|
| Gate zero: classify intent and SERP shape first | Google's rater guidelines define the intent taxonomy; SERP shape is directly observed | T1 | Stable |
| Gate one: page-type mismatch is a heavy prior | BrightLocal per-keyword page-type mix, plus the NavBoost mechanism from DOJ testimony | T3 | Annually |
| Gate two: median referring domains as the entry bar | Descriptive statistic. No study establishes it as a threshold | T3 | Replace with your own outcomes |
| Gate three: indexation before planning | Google documentation, prerequisite for both ranking and AI citation | T1 | Stable |
| Internal linking lifts traffic | SearchPilot controlled split tests, plus 5 to 7% measured | T2 | Annually |
| Front-load answers for extraction | Observed positional bias in grounding snippets. An empirical pattern, not a confirmed mechanism | T3 | Every 6 months |
| DR bands for a low-authority site | Practitioner heuristic calibrated from referring-domain benchmarks | T4 | Replace with your own outcomes |
| Rank-conditional GEO effects | One simulated-engine study, unreplicated on production systems | T3 | Every 6 months |
The two T4 and T3-heuristic rows are the honest weak points, and naming them is what separates this from a sales document. Every band in the link table should be replaced with your own measured outcomes as engagements complete, which turns the weakest row into the most defensible one within about a year.
A plan that never says "do not chase this" is a sales document. Fire these:
Only 1.74% of newly published pages reach the top 10 within a year for even one keyword. But among pages that do make it, 40.82% got there within one month. The distribution is fast-or-never, not a slow grind. Combined with SE Ranking's finding that 84.8% of top-three entrants came from pages already in the top 20, the operational conclusion is clear: improving near-misses beats publishing new pages, and a page that has not moved in six months needs a different keyword, not more patience.
Everything above is an organic playbook. For a Front Range service business it is frequently the wrong one, and the demonstration query proves it: SEO consultant Boulder puts a three-result local pack above every organic result, and the method as written walks straight past it.
If a local pack sits above the fold, the map pack is the plan and organic is the supporting act, not the other way round. Do not run the organic teardown and bolt local on as an afterthought. And never use a competitor's organic ranking as a proxy for their map-pack ranking. They are separate systems with different inputs, and a firm can dominate one while being invisible in the other.
Three inputs, and only one of them is on your website.
Is the profile verified, unsuspended, and eligible for the category at all? Service-area businesses have different rules from storefronts. This is the gate, and it fails more often than people expect.
The primary category does most of the work and is the highest-leverage single field in the whole profile. Pull the primary category of every business in the pack and compare. If all three competitors share a category the client does not have, that is usually the finding of the engagement.
Because proximity dominates, a single coordinate tells you almost nothing. Sample rank across a grid of coordinates spanning the actual service area and report the map, not the number. The honest deliverable is "you rank 2 to 4 within two miles and drop out entirely past five," which is a completely different conversation from "you rank 3."
Count, average rating, velocity over the last 90 days, recency of the most recent, whether the owner responds, and whether review text mentions the services you want to rank for. Compare against the pack incumbents rather than an abstract target. A competitor with 28 reviews is a very different problem from one with 400.
Name, address and phone consistency across the directories that matter in the market. This overlaps directly with the directory work in the organic plan, which is convenient: the same Clutch, Thumbtack and Yelp placements that occupy organic slots also feed prominence here. One effort, two systems.
This is the only place the organic and local playbooks genuinely converge. The page the profile links to should be locally relevant and match the primary category. Note that this is exactly what the Boulder incumbent does: the page I pulled apart in section 3 is his local-pack landing page, not his organic entry.
Keyword-stuffed business names, fake addresses, and lead-gen listings are common in local service categories and they are reportable. Sometimes the highest-return action in a local engagement is a redressal filing rather than anything you build.
Pack rank is a means. The deliverable metrics are calls, direction requests, bookings and qualified leads, all of which are available in the profile's own performance data. A client who moves from position 3 to position 2 and gets no more calls has not been helped, and you will only know that if you were measuring calls.
The method above ends at arrival, which is a real gap: a page that reaches position 8 and falls to 15 within six weeks is a failed engagement, not a successful one. Two numbers set the stakes.
Keep the organic ranking plan and the AI-visibility plan as separate documents with separate KPIs and timelines, even though one teardown feeds both. They are different systems, they move independently, and merging them is how you end up optimizing a page for extraction in ways that cost it rank.
Organic plan: position, clicks, conversions, measured in Search Console. AI-visibility plan: citation share and appearance frequency across repeated runs, per platform, never blended into one score and never reported as a rank. Different timelines too, since organic moves on core updates and AI citation moves whenever a model ships.
This section is worth more than the rest combined, because every competitor in Boulder is selling at least two of these.
| Claim | What the evidence actually says |
|---|---|
| Schema improves AI citations | Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 matched controls. Result: minus 4.6% on AI Overviews, statistically significant, and nothing on AI Mode or ChatGPT. Correlational studies find a small positive because sites that add schema do everything else well too. Keep schema for rich results and Bing. Do not sell it as an AI lever. |
| llms.txt helps AI visibility | 97% of llms.txt files received zero requests in May 2026, across 137,000 domains' server logs. The top requesters are the SEO audit tools checking whether the file exists. Google states it ignores them outright. You already left this off Ben's page. The evidence caught up with you. |
| Link velocity is a ranking signal | Ahrefs went looking across 10,000 SERPs and states plainly they "failed to distill any observable relationship." Velocity appears in the leaked documentation only as a spam signal. |
| Core Web Vitals will lift your rankings | Google's own FAQ: "trying to get a perfect score just for SEO reasons may not be the best use of your time." Fix genuinely broken LCP and INP. Do not chase 100s. |
| We track your rank in ChatGPT | SparkToro ran 12 identical prompts 2,961 times across 600 volunteers. Odds of the same brand list appearing twice: under 1 in 100. Same list in the same order: under 1 in 1,000. Any tool selling a rank position in an AI answer is selling noise. Report a visibility percentage across many runs, never a rank. |
| Rank top 10 and the AI will cite you | Ahrefs' AI Overview citation overlap with the top 10 fell from 76% to 38% between July 2025 and March 2026, though they honestly flag that their own parser improved and the datasets are not comparable. Off Google it is far weaker: only 12% of URLs cited by ChatGPT, Gemini and Copilot appear in Google's top 10. |
Nearly all AI-citation research comes from companies selling AI visibility tools. On the single question of how often AI Overview citations come from the top 10, Ahrefs says 38%, BrightEdge says 17%, seoClarity says 33%, and Originality says 52%. Same question, threefold spread. Treat the direction as real and the decimals as unusable in a client deck. The field's most-repeated numbers are mostly denominator errors.
The strongest correlate anyone has measured for AI visibility is YouTube brand mentions, at roughly 0.71 to 0.74, against backlinks at 0.218. Frequency of mention beat reach of mention, meaning twenty small creators mentioning a client may register more than one viral video. Unlike link building, this is not gated by domain authority, which makes it the single best lever available to a small local business. Caveat honestly: Ahrefs' sample was filtered to domains above DR 40, which inflates the correlation between "is a big brand" and everything else. Directional, not a coefficient to plan against.
The Princeton GEO paper is widely cited for its headline, that adding statistics, quotations and cited sources lifts visibility by up to 40%. Almost nobody quotes its third table, which breaks the same result out by where the page already ranked:
| Tactic | Rank 1 | Rank 3 | Rank 5 |
|---|---|---|---|
| Cite sources | -30.3% | +20.4% | +115.1% |
| Add quotations | -22.9% | +3.5% | +99.7% |
| Add statistics | -20.6% | +8.1% | +97.9% |
These tactics roughly double visibility for the weakest source in the set and measurably hurt the one already in front. It is a leveling effect, not a universal uplift. That is a genuinely good argument to make to a Boulder business sitting at position four or five behind an entrenched incumbent, and it is also the reason no off-the-shelf tool gives correct advice here: not one of them conditions its recommendations on where the client already ranks. Yours can.
Keep the caveat with it. The study ran on a simulated engine over 2023-era results, "visibility" means share of words in the generated answer rather than clicks, and nobody has replicated it on 2026 engines. Also note keyword stuffing was the one tactic that scored below doing nothing.
Asking DataForSEO for depth 20 returns page one only, because max_crawl_pages defaults to 1. Depth also multiplies cost per ten results, and each search operator such as site: multiplies the task cost by five. Budget accordingly.
Google disabled the num=100 URL parameter around 10 to 12 September 2025. Rank-tracking bots had been generating impressions across positions 1 to 100 in a single request, so removing it removed the bot impressions. Across 319 analyzed properties, 87.7% of sites lost Search Console impressions and 77.6% lost unique ranking terms. Clicks were unaffected, average position "improved" arithmetically, and click-through rate rose mechanically. Use October 2025 as the new baseline, and if a client shows you a chart with a cliff last September, that is what it is.
Most Boulder agencies will have quietly let clients believe that September drop was their fault or someone else's failure. Explaining it correctly, with the mechanism, is a five-minute conversation that positions you as the person who actually reads the primary sources.
Free, MIT licensed, self-hostable on Cloudflare's free plan, bring-your-own DataForSEO key, ships its own MCP server plus prebuilt Claude Code skills. Roughly 8,100 stars and actively maintained. An afternoon of setup for keyword research, rank tracking, competitor insight, backlinks, site audit and AI visibility. Highest leverage item in this entire report.
The cheapest deep page extraction that exists, and version 24 ships a built-in MCP server. Two caveats: it cannot fetch a SERP at all (its "SERP mode" is a snippet pixel-width preview, not results retrieval), so you feed it URLs from DataForSEO. And list mode is not yet supported over MCP, which kills the obvious "hand Claude twenty competitor URLs" workflow for now.
Reads up to 100 results, has both a SERP API and a Brief API, costs $0.47 a brief. Use it to find out whether your brief format is right before spending fifteen hours building one. If it turns out you need things Thruuu does not carry, which it will, you will then build the right thing instead of the obvious thing.
Around the schema in section 5 and the seven steps in section 6. The reason to build is not cost. It is that you own the prompt, so the deliverable carries your methodology and your voice rather than a vendor's Content Score, and you can fold in link authority, keyword intersection and AI Overview citation data that no per-brief tool exposes. It is also the only thing here that makes SEO Believer its own proof of concept while every competitor resells the same eight dashboard logos.
For the backlink index, which you genuinely cannot rebuild, and because clients recognize the DR number. Its MCP is included on every paid tier. Do not stack Content Kit and Brand Radar on top, which takes it to $427 a month for things you can build.
Every number in it is currently an estimate, by its own admission. Run the real volumes, difficulties and SERP compositions, and record which Tier 1 terms fail the two gates in section 6. Some of them will.
trafilatura as the default, crawl4ai when you need JavaScript rendering. Deliberately not Firecrawl if any of this ever becomes a hosted client-facing tool, because it is AGPL-3.0 and the network clause bites.fastembed with 256-dimension truncation. No PyTorch dependency, which is a materially better Windows install story.seo-content-brief skill. It is the most thoroughly specified brief structure published anywhere, and it is a better starting spec than anything you would write cold.Only two tools in this entire market can be driven from Claude Code today. Frase ships an official MCP server that is pushed to daily, plus 50-plus documented REST endpoints, from $39 a month. NeuronWriter ships an official MCP server and a public six-method API at $57 a month yearly, and its evaluate-content method scores a draft without consuming credits, which is the cheapest correct scoring primitive available anywhere. NeuronWriter also happens to be the one tool that scored best in Ahrefs' independent cross-tool test. Surfer's own API is locked to the $299 tier and its MCP is announced with no date.
1. The coverage matrix. Twenty pages by fifteen fan-out sub-queries, scored yes, partial or no. Every practitioner describes it. Nobody ships it.
2. Rank-conditional GEO advice. The same tactic helps a position-five page and hurts a position-one page. No tool conditions on this.
3. A grounding-budget-aware brief. One that targets under 1,500 words on purpose and competes for share of a fixed 2,000-word pie. Every content tool in the market currently optimizes in the wrong direction on this.
With only 38% of AI Overview citations and 12% of ChatGPT, Gemini and Copilot citations coming from Google's top 10, do not sell "outrank the top 10" as an AI visibility deliverable. Sell it as what it is, an organic ranking deliverable, and price AI visibility separately as an honestly caveated measurement product. Google still sends roughly 345 times more traffic than ChatGPT, Gemini and Perplexity combined. The traffic math still says win the blue link.