
Everybody in SEO has an opinion about anchor text and almost nobody has a number. So I bought one. I pulled 46 complete search results, scored 303 ranking pages against their own backlink anchors, and tested two things people argue about constantly: whether it helps to have your anchors varied, and whether it helps to have them match the keyword exactly.
Neither one predicted where a page ranked. What did predict it was looser than either: anchors that use the query’s words in any order, with the small words thrown away. Matching the phrase exactly did nothing. Being recognisably about the thing did.
The short version
- Exact-phrase anchors: nothing. The anchor that reproduces your keyword word-for-word showed no relationship with position.
- Token-match anchors: a real effect. Ignore word order and stopwords, and the relationship holds up under the strictest test I ran.
- Anchor variety: nothing. Four different ways of measuring how varied a profile is, four zeros.
- How many different sites link to you still matters more than any of it. That was true at the start and it is still the biggest number on the page.
What the people who build links for a living already told me
I have recorded a hundred-plus interviews with SEOs, and the useful thing about running a study afterwards is that you can hold the result up against what practitioners already believe. Sometimes you contradict them. This time you mostly do not, which is its own kind of finding.
“It’s not some stupid third-party metric like DA or DR or trust flow or anything else that matters about whether a link is valuable or not. What matters is whether it’s relevant.”
Bradley Benner · Semantic Links
Bradley places roughly 90% of his tier-one links as brand or compound anchors. He is deliberately not building exact-match anchors, and he has been saying so for years. The data now agrees with him.
“Anchor text and surrounding supporting content are definitely important too. But those come secondary to that initial review to make sure that this is even a domain that’s worth going after.”
Jeremy Moser · uSERP
He ranks anchor text third, behind whether the site is a real business and whether its authority is healthy. The regression puts them in the same order.
“Whoever has the bigger footprint and the more people talking about them across the web in a positive way — that is the ultimate be-all and end-all of winning and dominating the SERPs.”
Rob Bonham · The SEO Advisory
Footprint, in plain speech, is how many different domains link to you. It beat everything else I measured.
“No one’s going to tell us to what degree that patent is still in effect, how much impact it has, or whether it connects with something else. If you know how to test it and you know how to be skeptical, then it’s still valuable.”
Chris Green · Torque Partnership
This is the whole argument for measuring instead of citing. Chris also named the failure mode: “I’ve proved this patent proves this to be true, therefore I’m going to keep doing it.”
There is one place a guest and the data part company, and it is the more interesting half. Malte Landwehr, who wrote his undergraduate thesis on PageRank, told me footer links are probably discounted now — “links that are in white colour on white background somewhere in the footer, these are probably not counted anymore as strongly as in the past.” That is the consensus view and he may well be right. I could not test it. Across 233,845 classified backlinks pointing at pages that rank, only 4.34% sit in navigation, footer, header or sidebar at all. There is not enough variation in the data to correlate against, which is a different answer from “he is wrong”.
The numbers, exactly as they came out
I am not going to narrate statistics at you in a voice I do not have. Here is the output, and you can argue with it.
| Measure | Effect on rank | p | Verdict |
|---|---|---|---|
| Exact match, stemmed (word order and stopwords ignored) | b = −0.257 | 0.008 | Real |
| Exact match, literal (anchor equals the keyword) | ρ = −0.127 | 0.087 | Null |
| Anchor diversity (normalised entropy) | ρ = −0.095 | 0.185 | Null |
| Referring domains (the control) | b = −0.220 | 0.008 | Real, and larger |
Within-SERP comparison with keyword fixed effects, controlling for referring-domain count, standard errors clustered on keyword. A negative number means the page ranks better. The bottom row is a control variable included specifically to prove the method can detect a real effect — it did, in every model.
In positions: going from zero token-matching referring domains to seven moves a page from about position 10 to about position 7.
Correction, 9 September 2026 — this result is weaker than I said
A follow-up run put a brand-size control into this model — the number of anchor rows a page has, which tracks how large and well-linked the site is — and the stemmed-match result attenuates. On the same 299 pages the coefficient moves from b = −0.252, p = 0.011 to b = −0.206, p = 0.037, an 18% reduction. It still clears conventional significance and the interval still excludes zero. It no longer clears the study’s own Bonferroni threshold of 0.0125, which is the bar I held every other measure to above.
In positions, going from zero to seven token-matching domains now reads as roughly position 10 to position 7.4, not 7.0. The mechanism is not mysterious: pages with more anchor rows have more chances to contain a matching one, and within a SERP the two correlate at 0.40. That control was not in the original model because the original model did not have it. The pre-registration, the code and the numbers for this re-analysis are in the study bundle linked at the end.
Nothing above is deleted. This note is here so you can see what changed and why.
The mistake this study nearly made, and you might
Run it the obvious way and you publish the opposite of the truth
Split every search result at its median and compare the pages with more footer links against the pages with fewer. The footer-heavy pages rank better — a clean-looking result at p = 0.047. You could write that up as “footer links help” and nobody would immediately catch you.
It is entirely fake. Chrome-link count and referring-domain count move together at r = 0.68, because big sites accumulate more of every kind of link. Control for how many domains link to the page and the effect collapses to p = 0.35. The finding was never about footers. It was about being a big site.
That is the argument for fixed effects and control variables in one paragraph, and it is why I keep a variable in every model whose only job is to prove the instrument works. If a method cannot rediscover something already known to be true, its other results do not mean anything either.
What I am not claiming
- This is a snapshot, not an experiment. I did not change anybody’s anchors and watch what happened. It says pages with more topical anchors tend to rank better; it does not say that building them will move you.
- It only speaks for pages that carry their own links. 254 of the 460 ranking URLs I pulled never had enough page-level backlinks to score. Those pages rank on domain strength, and this study has nothing to say about them.
- Stop optimising the phrase, not the relevance. The result is that precision does not pay, not that anchor text is irrelevant. Those are different sentences.
- It is 46 keywords. Enough to detect an effect of this size, not enough to close the subject.
Why I am showing you the receipts
The whole thing cost 12.87 USD across 507 API calls, and every figure came from saved responses rather than a dashboard total. The frame was written down before the data was pulled, including which measures counted as primary, so I could not go shopping for a result afterwards. I found one defect in my own pipeline mid-run, fixed it, re-pulled the affected pages and published the corrected numbers.
None of that makes the answer right. It makes it checkable, which is the part most anchor-text advice is missing. If you want the full working — every measure, the pre-registered frame, the sensitivity tests and the two questions that came back undetermined rather than answered — the complete technical report is at the end of this post.
And if you are the sort of person who wants to argue with a method rather than a conclusion, come argue with mine. That is what it is published for.
The full technical report follows. Every measure, the pre-registered frame, the sensitivity tests, the practitioner quotes in full, and the two questions that came back undetermined rather than answered. It is long on purpose — the argument for trusting the number above is that you can check it.
The answer
Read this first
Exact wording does not predict rank. Topical overlap does. The distance between those two sentences is the whole finding.
Three measures were tested against rank across 303 URLs and 43 keyword clusters, and only one of them moved. Anchor diversity came back at ρ = −0.095, p = 0.185. Exact match in its literal form — the anchor string equals the query string — came back at ρ = −0.127, p = 0.087, interval brushing zero. Exact match in its stemmed form reached ρ = −0.182, and under the fixed-effects regression with standard errors clustered on keyword it is b = −0.257, SE 0.093, p = 0.0081, 95% CI [−0.444, −0.071]. That clears the Bonferroni threshold of 0.0125 and the interval excludes zero.
The two exact-match measures differ in exactly one way, and that difference is the result. The literal measure requires the anchor to be the query, character for character. The stemmed measure drops stopwords, ignores word order, and compares token sets — so best project management software, project management software and software for project management all count as the same anchor. The strict version finds nothing. The loose version finds something. What is being rewarded is not a matched phrase. It is whether the anchors are about the thing.
That reading also reconciles the two nulls rather than sitting awkwardly beside them. Diversity measures how spread a profile is, which is orthogonal to aboutness — a page can be described twenty different ways, all of them on topic, or all of them off it. Literal match measures precision, which is a stricter thing than topicality and evidently not the thing being counted. Aboutness is the measure that sits between them, and it is the one that moves.
The calibration row holds throughout. Referring-domain count — included in every model for the sole purpose of checking that the instrument can detect a real effect of this kind — comes back at ρ = −0.241, b = −0.220, p = 0.0079. It is still the larger effect. Nothing here displaces link volume as the primary driver; the stemmed-match coefficient sits alongside it, not above it.
Four measures, four intervals, one zero line
Primary result · ranks 1–20
Each bar is a 95% confidence interval on the within-SERP partial correlation with log₂(rank); the dot is the point estimate. Negative means the variable is associated with ranking better. The shaded band is the smallest effect this design can detect at 80% power. Two intervals clear zero: stemmed token match, and the referring-domain count that is there to prove the instrument works. Literal exact match reaches the edge and stops; anchor diversity does not get close.
b −0.257
p = 0.0081 cluster-robust. Clears Bonferroni α = 0.0125
ρ −0.127
p = 0.087. The pre-registered primary, and it is null
ρ −0.095
p = 0.185. Null at 1–10 and null again at 1–20
Holds
log(ref domains) b −0.220, p = 0.0079. Still the larger effect
303
204 at ranks 1–10, 99 at ranks 11–20, 43 clusters
0.17
Was 0.22 on the top-10 sample alone
The one sentence from this page that survives scrutiny
Among pages that carry a real link profile of their own, how many referring domains describe them in the query’s own vocabulary tracks with where they sit — and how exactly those anchors reproduce the phrase does not.
Three ways that sentence breaks in the retelling. 1. “Go build token-match anchors.” This is a cross-section, not an experiment. Nobody’s anchors were changed and nothing was watched afterwards, so the page cannot tell you what happens when you intervene. 2. “Anchor text beats links.” It does not, and not in a close race — referring-domain count carries the larger coefficient in the same model on the same pages. 3. “Anchor text matters after all.” Only the loose measure moved. The exact-phrase version, which is the one most people mean and most tools count, came back null at p = 0.087. And all three of those readings quietly drop the boundary: 254 of the original 460 ranking URLs never had enough page-level links to score, and pages ranking on domain strength alone sit outside this frame completely.
Anchor diversity lands on zero, in every model tried
Four diversity measures · n = 206 at ranks 1–10
The left panel is what most people would compute — every URL pooled, keyword ignored. The right panel is the comparison that actually controls for keyword difficulty: each SERP’s own cloud centred on the origin, so only rank-1-versus-rank-10 within the same keyword contributes. When these two disagree, the disagreement is the finding. They agree, and the cloud is round.
Every scored URL, with keyword difficulty removed
Within-SERP · keyword fixed effect
Both axes centred within each keyword, which is algebraically identical to a keyword fixed effect. Everything that varies between keywords has been removed; only within-SERP rank variation remains.
| Model | Measure | Raw r | Partial ρ ctrl log(ref domains) |
95% CI | Reading |
|---|---|---|---|---|---|
| Pooled | Anchor entropy | +0.007 | −0.082 | [−0.217, +0.055] | Null |
| Pooled | Exact-match domains | −0.099 | −0.030 | [−0.166, +0.108] | Null |
| Within-SERP | Anchor entropy | +0.097 | +0.000 | [−0.154, +0.155] | Primary test. Null, and powered |
| Within-SERP | Exact-match domains | −0.168 | −0.079 | [−0.231, +0.077] | Primary test. Null, and powered |
Note the raw within-SERP figure for exact match: −0.168, nominally significant at p = 0.016. It survives exactly until referring-domain count is controlled for, at which point it falls to −0.079 and the interval swallows zero. Pages with more exact-match anchors have more links; more links is the thing that predicts rank. That is what the control variable is there to catch, and it caught it.
The head-to-head is also a null. The pre-registered rule was that diversity beats exact match only if their intervals do not overlap. They overlap almost entirely. Neither wins.
| Fixed-effects regression | Coefficient | Cluster-robust SE | t | p | 95% CI |
|---|---|---|---|---|---|
| Anchor entropy | +0.062 | 0.611 | +0.10 | 0.919 | [−1.169, +1.293] |
| Exact match, log(1+domains) | −0.114 | 0.128 | −0.89 | 0.380 | [−0.372, +0.145] |
| log(1+referring domains) — calibration | −0.199 | 0.076 | −2.64 | 0.012 | [−0.351, −0.047] |
The calibration row, and why it is the most important number here
Outcome is log₂(rank), so a negative coefficient means better ranking. Standard errors are clustered on keyword across 45 clusters — enough that the cluster-robust asymptotics are no longer a stretch. The last row moved from p = 0.201 to p = 0.012 when the sample went from 43 URLs to 206. Nothing else about the design changed: same filter, same measures, same exclusions, same estimator. That is what a working instrument looks like. The two rows above it did not move, and now their stillness means something.
| Sensitivity strip — all exploratory | Within-SERP partial ρ | p | Note |
|---|---|---|---|
| H / ln(N) — normalised entropy (pre-registered primary) | +0.000 | 0.995 | Shape-sensitive, volume-corrected |
| Exact match, literal (pre-registered primary) | −0.079 | 0.261 | Zero tuning parameters |
| Exact match, stemmed token set | −0.188 | 0.007 | Order-insensitive, stopwords dropped |
The variant that moves — and what happened when it was tested properly
On the top-10 sample alone, stemmed exact match reached ρ = −0.188, p = 0.007. Three things stopped it being called a result at that point. One: it was not the pre-registered measure — literal equality was chosen in advance precisely because it has no tuning dial. Two: seven sensitivity measures were run, so the Bonferroni threshold was 0.0071, and it landed on the line rather than past it. Three, decisively at the time: re-fit with standard errors clustered on keyword it weakened to b = −0.243, p = 0.038, interval reaching −0.014.
It was therefore written up as a lead requiring a fresh pre-registered test, and then that test was run. Section 3 pre-registered stemmed match as a named hypothesis before any band-extension data was pulled, on the stated rule that it would either strengthen or wash out. On the widened sample it is b = −0.257, p = 0.0081 under the same clustered estimator that had weakened it — now clearing a Bonferroni threshold of 0.0125. The table above is the top-10-only sample and is left as it was recorded. The pooled figures supersede it, and the honest caveat is that this is the same 46 keywords with more of each SERP, not an independent replication.
Robustness on the entropy estimator. Normalised entropy runs high on small profiles by construction — mean 0.84 where N < 30 against 0.73 where N ≥ 30. Restricting the primary test to the 138 URLs with N ≥ 30, where the estimator is stable, gives ρ = +0.061, p = 0.478. Still nothing, and now nothing from the well-measured half of the sample.
Subgroups do not rescue it either. Run within each niche with 15 or more scored URLs, the diversity correlation ranges from +0.28 in B2B SaaS to −0.33 in Finance, and not one of six reaches significance. Two of the largest point in opposite directions, which is the signature of noise rather than a hidden vertical-specific effect. They are reported here so that nobody has to go looking for them later.
A top-ten-only sample was structurally blind to this
Ranks 11–20 · pre-registered before collection
Everything above the previous version of this page measured was conditioned on winning. All 206 URLs already ranked top ten — they had beaten millions of candidate pages to be there. Conditioning on the outcome is conditioning on a collider, and it does something specific: it induces negative correlation between the causes of the outcome. A page sitting at rank 4 with thin, off-topic anchors has to be compensating somewhere else, or it would not be in the set at all. That compensation structure drags every predictor toward zero inside the winners’ circle.
Which means the original null was ambiguous in a way the design could not see. It was consistent with “anchor structure does nothing”, and equally consistent with “anchor structure decides whether you enter the consideration set, and then stops mattering for ordering within it.” Those are very different claims and the first one was published. The way to separate them is a comparison group that did not win.
That group turned out to be already paid for. The 46 SERP calls in the main run were issued at depth: 20 and returned 16 to 20 organic results each. The pipeline stored the top ten and discarded the rest, but the raw responses were saved to disk, so positions 11 onward were sitting there at zero additional SERP cost. Re-running the study’s own filter over them — the domain blocklist, then one URL per registrable host, then sequential renumbering — reproduced the published top ten exactly for all 46 keywords, which is what makes the extension a like-for-like comparison, and yielded 319 URLs at ranks 11–20.
Band 11–20 clears the pre-screen at 46% against 55% for the top ten, and survives the anchor filters at 67% of those against 81%. Both gaps point the same way: pages on the second screen carry thinner page-level link profiles. That is unsurprising, and it is also the reason the band-only tests below are underpowered.
| Ordering test, by sample | n | Clusters | Diversity | Exact match, literal |
Exact match, stemmed |
Calibration log(ref domains) |
|---|---|---|---|---|---|---|
| Ranks 1–10 only | 202 | 41 | +0.000 p 0.995 |
−0.079 p 0.266 |
−0.188 p 0.0075 |
−0.264 p 0.00015 · detects |
| Ranks 11–20 only | 90 | 27 | +0.065 p 0.540 |
−0.109 p 0.308 |
−0.117 p 0.271 |
−0.008 p 0.939 · fails |
| Pooled, ranks 1–20 | 303 | 43 | −0.095 p 0.099 |
−0.127 p 0.027 |
−0.182 p 0.0015 |
−0.233 p 0.00004 · detects |
Read the calibration column first, because it decides which rows are allowed to mean anything. It detects on the top ten and it detects pooled. Within band 11–20 alone it does not — ρ = −0.008, p = 0.94, at n = 90. That whole row is therefore uninformative rather than null, and it is printed only because removing it would be choosing which rows to show after seeing them.
What the extension actually bought was power on the pooled model, not a second independent sample. Stemmed match was already at ρ −0.188 in the top ten. Pooled it is −0.182 — the same magnitude, a smaller p-value, and it now survives the clustered standard errors that had previously halved it. That is the signature of an underpowered true effect gaining n, not of a new effect appearing.
| Pooled 1–20, keyword fixed effects, cluster-robust SE | Coefficient | SE | t | p | 95% CI |
|---|---|---|---|---|---|
| Exact match, stemmed — log(1 + domains) | −0.2573 | 0.0925 | −2.78 | 0.0081 | [−0.4440, −0.0706] |
| log(1 + referring domains) — calibration | −0.1573 | 0.0692 | −2.27 | 0.0282 | [−0.2968, −0.0177] |
| Anchor diversity (all-measures model) | −0.7497 | 0.5566 | −1.35 | 0.1852 | [−1.8730, +0.3735] |
| Exact match, literal (all-measures model) | −0.1896 | 0.1082 | −1.75 | 0.0872 | [−0.4080, +0.0289] |
How big is it, in the only units anyone cares about
The outcome is log₂(rank) and the predictor is log(1 + domains), so a coefficient of −0.257 means: going from zero stemmed-match referring domains to seven multiplies rank position by 2−0.535 = 0.69. Position 10 becomes roughly position 7. That is a real effect size rather than an academic one — and it is still smaller than the referring-domain effect measured in the same model, which is the point that should survive any summary of this page.
| Qualification test — does the measure separate band 1–10 from 11–20? | Coefficient | p | Verdict |
|---|---|---|---|
| Anchor diversity | +0.302 | 0.211 | Undetermined |
| Exact match, stemmed | +0.043 | 0.208 | Undetermined |
| log(1 + referring domains) — calibration | +0.047 | 0.155 | Fails |
The qualification question is still open, and this is why
The hypothesis the extension was built for — that anchor structure decides whether a page enters the top ten rather than where it sits inside it — came back undetermined, not null. Collapsing rank into a binary top-ten indicator throws away most of the information in the outcome, and the linear-probability model could not even recover the referring-domain effect it was calibrated against (p = 0.155). When the calibration row fails, no other row in the model is admissible. Answering the qualification question properly needs pages that did not rank at all, which no SERP endpoint returns; it needs a sampling frame built some other way.
Link placement cannot be measured from ranking pages alone
Reasonable surfer · no new API calls
The anchors endpoint returns referring_links_semantic_locations on every row already paid for — a breakdown of where each link sits in the referring page’s HTML. That makes a structural hypothesis testable for free: Google’s reasonable surfer model values a link by the modelled probability someone clicks it, and position on the page is the strongest published component of that model. In-content links should be worth more than nav and footer links.
The test cannot be run, and the reason is the finding. The share of a page’s classified backlinks sitting in main content has a median of 0.998 and an interquartile range of 0.990 to 1.000. Half the sample is pure in-content. There is a real tail — one page reaches 94.6% chrome — but it is fourteen pages out of 193, and those fourteen are large brands whose ranking is overdetermined. Regressing rank on main-content share returns ρ = +0.017, p = 0.81, and a null on a predictor that is 97% constant says nothing about the hypothesis.
The obvious objection was tested and rejected. If the 45% unclassified bucket were quietly absorbing chrome links — pages built without HTML5 sectioning elements — the 97% would be a parser artefact rather than a property of the web. It is not: r(unclassified share, chrome share) = +0.023, p = 0.75, and pages with zero chrome links have fewer unclassified links, not more. Truncation cuts the same way — 35 of 301 anchor pulls hit the 1,000-row cap, and because rows are ordered by referring-domain count, truncation drops chrome-heavy anchors preferentially. The true chrome share is a little above 4.34%, not below it.
Nor does the band extension rescue it. Main-content share is 0.9981 at the median for ranks 1–10 and 0.9968 for ranks 11–20. Welch’s t reports p = 0.0495 on the means; Mann-Whitney reports p = 0.37 on the ranks. When a mean difference is significant and the rank-based test is not, the mean is being moved by outliers, and here it is two of them.
The trap this section exists to document
Split each SERP at its own median chrome share and the chrome-heavy pages rank better — mean log₂(rank) 1.75 against 2.09, p = 0.047. Taken at face value that reads as “footer links help”, which inverts the hypothesis. It is entirely an artefact: chrome-link count correlates with referring-domain count at r = 0.68, because large sites collect more of every kind of link. Control for domain count and it collapses to p = 0.35. Run this analysis without fixed effects and without a control variable and you publish a confident finding that is backwards.
What can be said, and it is a composition claim not a value claim
Across 233,845 classified backlinks pointing at pages that rank in the top twenty, 4.34% sit in nav, footer, header or sidebar. That is a statement about how the link graph is shaped, not about what a footer link is worth. Nothing here shows that boilerplate links are worthless — the sample contains only pages that already rank, so there is no comparison group of pages that failed, and the pages whose profiles are overwhelmingly chrome rank at positions 2, 4 and 7. The confidence interval on the effect, [−0.224, +0.092], rules out a large effect and is comfortable with a modest real one.
Where this lands next to the people who build links for a living
Verbatim, from the Unscripted archive
A correlation is worth more when you can hold it next to people who do the work and see whether it matches what they already believe. Every quote below is verbatim from an Unscripted interview. Some of them corroborate the result. One of them is a practitioner asserting something this study could not measure at all, which is the more useful kind of disagreement.
No one’s going to tell us to what degree that patent is still in effect, how much impact it has, or whether it connects with something else. So how do we follow that chain from academic side to this is how it impacts? If you know how to test it and you know how to be skeptical, then it’s still valuable. Still highly valuable.
Chris is describing the exact gap this page fills. He also names the failure mode it guards against: “I see a lot of people will often over-index on these things. It’s kind of like, well, I’ve proved this patent proves this to be true, therefore I’m going to keep doing it.”
It’s not some stupid third-party metric like DA or DR or trust flow or anything else that matters about whether a link is valuable or not. What matters is whether it’s relevant.
This is the closest thing to independent corroboration in the archive, and it predates the measurement. Bradley places roughly 90% of his tier-one links as brand or compound anchors — deliberately not exact-match — on the argument that relevance compounds across the content, the site and that site’s own backlink profile. His second line puts it in the study’s own terms: “Our job should be about creating associations, strengthening those associations, and forcing the models and the algorithms to recognize those associations.” Association is what the stemmed measure counts.
Anchor text and surrounding supporting content are definitely important too. But those come secondary to that initial review to make sure that this is even a domain that’s worth going after.
A three-tier test that ranks anchor text third, behind whether the site is a legitimate business and whether its authority and traffic are healthy. The regression agrees with the ordering: referring-domain count carries the larger coefficient in every model on this page.
He had a banana hammock page and he bought 400 comment spam links and the darn thing had nothing else — it had never been optimized in any other way, shape, or form — but it ranked better because it had anchor text links related to banana hammocks for men.
Related to, not matching. That is the finding, told as a story years before it was measured, and it is a cleaner natural experiment than anything in the sample: a page with no other optimisation, moved by topically-worded links alone. Paul Baterina independently reports the same shape — a client’s ~500 comment-spam anchor links that “actually worked.”
There are potentially some links — links that are in white colour on white background somewhere in the footer. These are probably not counted anymore as strongly as in the past.
Malte wrote his undergraduate thesis on PageRank, and he is confident some variation of it still runs: “We cannot know if the exact algorithm from this random surfer paper is being used. But some variation of it is definitely used still today.” His footer claim is the practitioner consensus. This study cannot confirm or deny it — 4.34% of the classified links pointing at ranking pages sit in chrome at all, and main-content share has a median of 0.998, so there is no spread to correlate against. The disagreement is not about the answer. It is that one of us has data and it says the question is unanswerable from ranking pages.
In theory, if a robot acted completely random or reasonable for a task, it could. But the reality is that a lot of crawling on the internet is done for very specific purposes. I would say it is more an approximation of user behaviour. Artificial crawling by bots has very different characteristics from human crawling.
Asked directly whether reasonable surfer applies to robots. If the crawler is not a surfer, a placement-weighted model is describing a reader who increasingly is not the one doing the reading — which is a better reason to be sceptical of the placement hypothesis than any number on this page.
Whoever has the bigger footprint and the more people talking about them across the web in a positive way — that is the ultimate be-all and end-all of winning and dominating the SERPs.
Footprint, in plain speech, is referring-domain count. It is the effect that survived every model here, and it is larger than the anchor effect. Adrian Nikolov says the same thing in four words: “Authority is the new PageRank.”
There was a full page before I got to the data saying, this is correlation, not causality, this is what it means. Don’t look at it as this is baked into any algorithm. We’re just looking at patterns here.
Ben ran his own correlation study across 100 industries and 1,100 personas, and led with the caveat rather than burying it. The same caveat governs this page: nothing here says that changing a link profile changes a ranking. It is a cross-section, not an experiment.
The signal five guests raised that this study never tested
Distance to seed comes up independently from Alejandro Meyerhans, Paul Baterina, Steven Schneider, Nick Eubanks and Matt Brooks — the idea that a link’s worth depends on how far the linking domain sits from a trusted seed set inside a vertical. Nick Eubanks names it as a large factor in a ranking shift he investigated; Steven Schneider’s team concluded the same thing about a medical-niche move. Nothing on this page measures it. The models here control for how many domains link to a page, not for what those domains are close to. If the anchor effect found here is partly a proxy for something else, that is the most likely candidate, and it is measurable with the same endpoint.
Which keywords, and why those twelve niches
Pre-registered frame
Forty-six keywords across twelve niches, every one of them pre-registered in a frame file — niche, intent, head or mid-tail, and a stated reason each — before the first SERP call. Allocation was weighted by measured yield rather than spread evenly, because usable-URL yield, not SERP count, is what buys statistical power.
Scored-URL yield by niche, all 46 SERPs
460 ranking URLs
Solid bar is the share of ranking URLs that produced a usable entropy score. The pale bar behind it is the share that cleared the 10-referring-domain pre-screen — the gap between them is what the spam filter and the two-distinct-anchor requirement removed after the anchors were paid for. Marketing/SEO was added on the expectation that it would be the densest link economy in the frame, and it was.
Two niches were added. Marketing / SEO because it is the densest content-marketing link economy on the open web and the closest available analogue to the vertical Jeremy’s internal ρ 0.33 came from — the most informative place to look if a real effect exists. It returned the highest yield of any niche at 77%. Tech / developer because docs, tutorials and comparison pages accrue deep page-level links from forums and repos, in an anchor economy where nobody is optimising anything. It returned 43%, below expectation.
The thin niches were kept, not dropped. Travel and local service were the thinnest verticals in the frame and were kept at reduced weight rather than dropped. Dropping them would have raised the yield number by making the frame agree with itself; keeping them means “which niches have unscoreable page-level profiles” stays a measured finding. Travel came back at 4 of 30 across three keywords, home improvement 6 of 30, hobby 7 of 30. The pattern held.
Three keywords were deliberate probes rather than sample. mortgage refinance rates was placed inside high-yield Finance as a low-yield page type — a rate table, rarely linked at page level — to separate vertical effects from page-type effects; it returned 7 of 10, so the prediction was wrong and page type is not the whole story. how much water should i drink a day was included as a negative control on exact match, because anchors to question pages are question-shaped rather than keyword-shaped. power of attorney form tests the same idea on a utility page.
| Niche | Early yield | New keywords | Full-run yield | Why it was weighted this way |
|---|---|---|---|---|
| Marketing / SEO | — | 3 | 23/30 · 77% | New. Densest link economy available and the nearest analogue to the internal study’s vertical. |
| B2B SaaS | 9/10 | 5 | 38/60 · 63% | Highest early yield, so the largest allocation. Replicated across a second head term rather than resting on one. |
| Finance | 6/10 | 4 | 27/50 · 54% | Where practitioners most believe exact match works. Both intent poles sampled. |
| Education | 7/10 | 4 | 27/50 · 54% | Second-best early yield; enormous referring-domain counts stress the control variable. |
| Health / medical | 6/10 | 4 | 24/50 · 48% | YMYL editorial anchors, plus the exact-match negative control. |
| Legal | 7/10 | 3 | 19/40 · 48% | The adversarial case for the spam filter, now at two independent head terms. |
| Tech / developer | — | 3 | 13/30 · 43% | New. Descriptive, unoptimised anchors as a contrast with Finance and Legal. Underperformed expectation. |
| Ecommerce / product | 3/10 | 3 | 12/40 · 30% | Retested with more editorially covered terms to see whether the thinness was the vertical or the modifier. It was the vertical. |
| Local service | 1/10 | 1 | 6/20 · 30% | Minimum weight. The regime where entropy is least stable, kept so the instability is measured. |
| Hobby / enthusiast | 2/10 | 2 | 7/30 · 23% | Earned rather than built links give the most natural anchor distribution in the frame. |
| Home improvement | 2/10 | 2 | 6/30 · 20% | Kept to confirm whether DIY publisher pages sit reliably below the floor or merely near it. Reliably below. |
| Travel | 0/10 | 2 | 4/30 · 13% | Worst yield in the frame, retained on purpose. Dropping it would have made the yield estimate self-fulfilling. |
What the frame excludes, and why that matters
Jeremy’s own 25 properties and contentguppy.com were excluded by domain filter before any URL entered the frame; none appeared in these 46 SERPs, and the filter is logged either way. That exclusion is what makes this an independent test rather than a rerun of the internal study, which measured internal links across sites sharing templates, hosting and heavy interlinking. One URL per registrable domain per SERP was also enforced, so no single site can contribute twice to the same cluster.
Why 254 of 460 URLs could not be scored
The binding constraint, confirmed at scale
Across 460 top-ten URLs, and a frame deliberately weighted toward link-dense niches, usable yield came to 44.8%. Weighting the frame toward the densest link economies barely moved it: page-level backlink profiles for ordinary ranking results are thin almost everywhere, and choosing better niches buys less than it looks like it should.
The pre-screen paid for itself. One bulk_referring_domains call priced all 460 URLs for under eight cents and let this run skip 156 of its 360 anchor calls, every one of which would have returned nothing scoreable — about $4.40 of avoided spend. Validated before it was trusted: every scored URL in the study had 10 or more referring domains, so the gate cannot discard a URL that would have survived. It is lossless by construction, since the scored sample size N can never exceed a page’s referring-domain count.
Four keyword clusters produced a single scored URL each — commercial cleaning services, air fryer vs convection oven, how to install laminate flooring, beginner guitar chords. A cluster with one member contributes no within-SERP contrast, so those four sit in the descriptive counts but carry no weight in the primary estimate. One cluster, things to do in lisbon, still produced nothing at all across both runs.
The interpretation consequence stands and is now larger. Restricting to URLs with ten-plus referring domains narrows the question to “among pages that have a real link profile of their own, does anchor structure explain position?” More than half the ranking web is outside that frame — those pages rank on domain strength, not on links to the page. That is a defensible study, but it is a narrower one than “does anchor text matter”, and the headline null should never be quoted without it.
Every measure, fixed before collection
Applied as written
The outcome, and why within-SERP
The outcome is log₂(rank) within the SERP. Raw position is non-linear — the distance from 3 to 5 is not the distance from 43 to 45 — and a log scale puts those on a comparable ratio footing.
Inference. Standard errors clustered on keyword. Forty-five clusters is a real cluster count — the small-sample correction is still applied, and the intervals do not need an apology.
Diversity
Unique ÷ total was rejected as primary because it falls mechanically as a profile grows — repeat phrasings accumulate faster than novel ones — so it would return a correlation with size dressed up as a correlation with diversity. It is reported in the sensitivity strip regardless, and it moved the answer by 0.04.
String normalisation, fixed in advance: NFKC, lowercase, collapse internal whitespace, strip leading and trailing punctuation and quotes.
Exact match
Literal equality was chosen over stemming and containment for one decisive reason: it has zero tuning parameters. There is no dial to turn once the results are in — which is exactly why the stemmed variant’s nominal significance is being reported as a lead rather than a finding.
Observed spread: 71 of 206 URLs carry at least one literal exact-match domain, the mean is 4.7 and the maximum 122. The variable has real variance — the null is not an artefact of an empty column.
The spam filter, and what it removed
Applied in order, before either measure, entirely from fields on the anchors endpoint itself.
S2 WEIGHT = referring_main_domains − referring_main_domains_nofollow
DROP WHERE weight <= 0 // kills nofollow-only and UGC-only anchors
S3 Weighting by DOMAIN, not by link, is the sitewide defence in itself
S4 BUCKET non-descriptive: naked URLs, bare domains, brand token, empty
S5 REQUIRE N >= 10 and >= 2 distinct anchors, else exclude the URL
S3 is the rule that matters most. contentguppy.com carries 92 backlinks on the single anchor http://contentguppy.com/ from one referring domain. Weighted by link it would dominate an entropy score outright; weighted by distinct domain it contributes exactly 1, and then S4 removes it as a naked URL. Both defences are load-bearing, and both fire on that same row.
The attrition ladder, as promised
32,774 raw anchor rows entered the filter across 283 URLs. Publishing this is what makes “spam removed” auditable instead of a claim, and the filter behaves the same across every niche in the frame.
- 31,002 rows — 94.6% survive the spam-score threshold. Attrition here is modest, which is itself informative: DataForSEO’s own score is not doing much work at a cut of 30.
- Well clear of the tripwire. The pre-registered abort was set at retaining under 10% of rows. At 75.2% the filter is doing real work without gutting the sample, so the null cannot be blamed on over-filtering.
What this changes, and what would move it
One effect found · two questions still open
Where the study lands
- Stop optimising the exact phrase. The thing that tracks with rank is topical overlap. Anchors matching the query’s token set — word order and stopwords ignored — predict position at p = 0.008 with the interval clear of zero. Anchors matching the query string exactly do not, at p = 0.087. If you are going to influence anchors at all, the target is whether they are recognisably about the page’s subject, not whether they reproduce a phrase.
- Anchor diversity remains a null, now twice over. ρ +0.000 on the top ten, −0.095 pooled across 1–20, p = 0.185. There is no version of the diversity measure — entropy, Hill number, Simpson, unique-over-total — that finds anything. Treat it as settled at this resolution.
- Referring domains still dominate. In the same model, on the same sample, link volume carries the larger coefficient. Nothing on this page displaces it. The stemmed-match effect is a second-order adjustment on top of a first-order driver, and any summary that reverses that ordering is misreading the table.
- The honest limit on the new finding is that it is not an independent replication. The extension added ranks 11–20 of the same 46 keywords. It widened the outcome range and bought real power, and it was pre-registered before collection, but it re-used the frame. A clean confirmation means a fresh keyword set with stemmed match named as the primary measure in advance. That is roughly 40 new SERPs and $6, and it is the single highest-value next spend.
- Two questions came back undetermined and should not be quoted as nulls. Whether anchor structure governs qualification rather than ordering — the binary model could not recover its own calibration row. And whether link placement matters — the measure is 97% constant, so there is nothing to correlate. Both need a different sampling frame, not more of this one.
- The whole thing cost 12.87 USD across 507 billed API calls. Every figure on this page is the sum of
tasks[].costfrom saved responses, not an account-level total — the key is shared with a production service, so the dashboard number for the day is higher and none of that difference belongs here. - The frame boundary is unchanged and still the first caveat. 254 of the original 460 ranking URLs never had enough page-level links to score. Everything here applies to pages that carry their own link profile, and to nothing else.
