Most keyword research stops at a list sorted by volume. This SOP turns that list into a relevance score Google itself supplied, and then turns the score into a revenue number you can put in a proposal. It distills the process Jeremy Rivera (founder, SEO Arcade) described as a guest on the Unscripted SEO Interview Podcast with host Mark A Preston — the manual sheet-stacking routine that eventually became the SEO Arcade tool. As he put it: “And then lightning struck my brain and I said, that’s the signal. That’s how I can tell that Google sees the overlap.” It belongs to our library of repeatable SEO SOPs.
It’s free inside The Vault — create a free SEO Arcade account and grab the printable SOP plus every other playbook, deck and eBook in there.
Open The Vault →Objective
Produce two ranked keyword lists and one revenue forecast from a single seed set, using co-ranking overlap as the relevance score. You take a handful of seed keywords, pull the pages already ranking for them, export everything those pages rank for, and stack it all in one sheet. The duplicate count is the score: a phrase that ten out of ten ranking URLs also rank for is a phrase Google has already told you is connected to your core term. The rare phrases become your long-tail list. Then you human-check the list, strip the terms you are legally or contractually forbidden to chase, rate each survivor by distance to conversion, and run the surviving list through volume → organic CTR → conversion rate → sale rate → revenue per sale. The deliverable is a tiered, directional proposal — not a single fake number.
Why co-ranking overlap beats a volume sort
The argument is that you do not need to model semantics yourself, because the SERP already contains the model. “I have a shard of data that Google says these pages are all connected to this term semantically, right? And they also co-rank for these other terms and phrases,” Jeremy said. “So what a fantastic way — instead of using machine learning and trying to figure out semantics on the back end of how are these terms connected, I can say that I know Google connects” them. His example was a local service term: “a good example is Cookeville pest control. You know, you have Google pest control, you have pest control service, pest control company, bug spraying. The terms and phrases that they co-rank for when you rank for one term gives you a little bit of a keyword universe for that particular service.”
The overlap count is what converts that universe into a priority order: “If for this particular phrase 10 out of 10 URLs all co-rank for one of those phrases, that is a higher priority, that is more relevant to the core phrase, versus one of these domains is really strong and it ranks for peanut butter chips because they have an article. You know, it’s a one-off, or it’s a deep cut, or it’s a longtail term.”
Key Steps
- Pick one to five seed keywords, not one. Jeremy started with a single seed and revised: “you have a single seed keyword. Actually, I bumped it up to five.” The reason is that near-synonyms return genuinely different keyword universes — “‘natural HVAC’ and ‘natural air conditioning,’ ‘natural heating’ and ‘natural cooling’ — they’re going to have a different matrix, a different universe of keywords.” Choose seeds that a customer would plausibly type, and make them cover the different words your market uses for the same thing.
- Take the top 10 ranking URLs for each seed, then export the top 100 keywords each of those URLs ranks for. This is the raw material. Jeremy’s manual version: “I’d go to Ahrefs, I’d plug in the URL, see where I’m ranking, and then grab one particular term for one particular page, and look at those top 10 URLs, go and download a CSV of the top 100 keywords that those pages rank for.” Any rank-data source with a per-URL keyword export works; the tool matters less than doing all ten URLs for every seed rather than the convenient three.
- Stack every export into one sheet and count the duplicates — the duplicate count is the score. “Then spend the next 30 minutes in Google Sheets copying, pasting, copying, pasting, copying, pasting from those downloaded CSVs until I had a document. And I realized, you know what, okay, now I’ve got duplicates, because a lot of these five pages all had the same term.” Do not de-duplicate. Add a count column. Sort descending, and the phrases appearing on 8, 9 and 10 of the ranking URLs float to the top as your high-relevance, high-priority list. Expect roughly 500 to 1,000 rows: “it usually is somewhere between 500 and a thousand keywords that you’re looking at.”
- Flip the sort and read the bottom of the list as a second deliverable. “So you could flip it on its head and say, okay, what are the terms and phrases that just one of our competitors really does well for? And boom, there you go, longtail research. You can find that phrase that only one of your competitors is going after, and create unique relevant content to target it.” A count of one means one of two things: a strong domain’s unrelated deep cut, which you discard, or an unclaimed opening nobody in the set has covered, which is the cheapest content you will build all quarter. Reading them apart is a human job, which is the next step.
- Do the human pass on the SERP — this step is mandatory, not optional. “And this is the thing: you still have a human process. There is no tool that’s going to be able to look at this list and tell that, you know — you have to go to the SERP, you need to look at that term, plug it in and see, how is Google recognizing that term or phrase?” You are hunting two things: false positives, where a high-volume term belongs to another business or a regional centre with the same name, and split intent. Jeremy’s example: “It’s like Nashville air conditioning — the traffic is higher than any other phrase, because half of the people want to buy an air conditioner, half of the people want air conditioning service and somebody to do it for them, and you’ve got to split that traffic in two intuitively.” Clear out “the ones that are not correct, or false positives, or you’re not targeting those” before anything downstream touches the sheet.
- Build the negative keyword list and get the client to vet it before you build anything. This is the step the play is really built around: “you’ve got to get your keyword list vetted by your client. You’ve got to find out your negative SEO keywords. In PPC, half of the job is figuring out what not to bid on, and SEOs on the regular miss this step.” Three sources of red lines: legal history (“we can’t say [X] because there was a lawsuit”), YMYL claim restrictions (“if you’re in the Your Money or Your Life category and you’re working with any sort of supplement, there are red lines of terms and phrases left and right, of claims that you cannot make”), and brand phrases that are simply off limits. Each banned phrase gets one of two dispositions: cross it off the optimisation list entirely, or “rank negative” — use it in a context that shows up for the term while stating you are not that term.
- Add a “distance to conversion” column and rate every surviving keyword. “I suggest in my sheet, you know, have an empty column for you to rate: how close to conversion is this keyword? Because there are people who are searching for a term or phrase, and yeah, that page ranks for it and it’s relevant, but they’re not really ever going to buy it. Or, hey, that’s my money term.” A simple three- or five-point scale is enough (inferred execution detail — Jeremy specified a rating column, not a particular scale). This column is what stops the forecast in step 8 from being driven by a high-volume phrase that never converts.
- Run the surviving list through the five-stage forecast. Search volume → organic CTR → conversion rate → sale rate → revenue per sale. Take monthly volume from your keyword source, apply a position-based organic click-through rate to turn volume into traffic (“out of the distribution for that particular type of phrase, you know, 25% go to the number one listing, you know, 23 go to the second”), multiply traffic by the site’s conversion rate to get leads, multiply leads by the lead-to-sale rate to get sales, and multiply sales by revenue per sale. “I know how many sales I have, I multiply that by the revenue for that product, boom, there’s an estimate. If I rank at the top of page one for that term, it could equal this much revenue at the end. There’s a funnel that goes from left to right, but you can math it out for a singular page or service.” For the CTR curve, use the Advanced Web Ranking organic CTR study — Jeremy’s reference of choice, “a fantastic resource, because it changes by whether it’s commercial intent or whether it’s image intent.” Do not skip the lead stage: “Even on e-commerce, it’s you add it to the cart and then you buy it… there is always a lead stage.”
- Tier the result into a layered proposal instead of one headline number. Nobody ranks first for everything — “you never rank number one for all of the terms that you rank for” — so present the forecast in coverage bands. Jeremy’s structure: “if I was in the top three for half of these terms with this conversion rate, the sale rate represents this much revenue. And then you take it down by half, say, if I ranked a quarter of these words in the top three, it could be this much.” The output is explicitly a sales document: “that kind of gives you a directional, layered goal and a proposal — really a proposal — of, hey, this is the organic search potential of your market for this service or product.”
Worked example: the space cooler that was never allowed to be a swamp cooler
This is what step 6 looks like when it goes right and wrong in the same engagement. The client was “obsessed with their brand phrase. They wanted to be a ‘space cooler’ instead of a space heater. They didn’t want to be known as a swamp cooler, or a portable evaporative cooler, even though that’s what people were searching for, that’s what the product literally is.” The red line was explicit and it was hard: “do not call it a portable evaporative cooler, and if you call it a swamp cooler we will fire you.”
The ceiling arrived exactly where the forecast said it would. “A year goes by, we’re number one for ‘space cooler,’ and they’re getting this much traffic.” Number one for an invented category is still capped by the demand for that category — “I told them on the outset, there’s only so much traffic that’s going to show up for this.” The fix was the “rank negative” disposition, not abandoning the brand phrase: “Why don’t I put on the homepage, ‘not your mama’s swamp cooler,’ and ‘looking for a portable evaporative cooler,’ and ‘want something cooler’? And we tripled their traffic in the first three months once we’d added those terms and phrases.” The banned phrase went back on the page in a context that honoured the red line.
Cautionary Notes
- Do not build this forecast on Google Search Console query data. Jeremy’s own numbers from his own site: “out of the 140 clicks to my article that’s literally about how accurate is Google Search Console, they gave me five clicks that I know what those queries are… 149 people visited, clicked to it. It gave me details for five of those clicks. So that’s 3% of the data that’s accurate.” The API and BigQuery escape hatches barely move it: “It took me half a day, and I’ve done this for 16 years… and it still only gave me — now I’m up to 12 clicks for that article.” Even the best case he has seen is partial: “I’ve never worked on a project where there was more than 70% of details.” His verdict: “No statistician can work with 3% of the data.” Use GSC to sanity-check what you already rank for, never as the volume input to a revenue model.
- A false positive at the top of the sort will hijack the whole plan. High-volume terms float up first, and that is exactly where the impostors are — another business with the same name, a regional centre, a split-intent phrase. Skipping step 5 means the biggest number in your proposal is the least real one.
- Skipping the client vetting call is how you get fired at month twelve, not month one. The red lines are invisible in every keyword tool. Nothing in Ahrefs or Semrush knows about the lawsuit, the FTC claim restriction, or the phrase the founder has banned. Get the negative list in writing before content is commissioned.
- The forecast is directional, and you must say so. Jeremy’s own framing of the tool is “In a way, yes. In a sandbox type of way. So it’s directional.” Present bands and assumptions, never a single confident figure. Every stage — CTR curve, conversion rate, sale rate — carries its own error, and they multiply.
- Co-ranking overlap does not cover topical or top-of-funnel content. Stated plainly: “It’s not perfect, because obviously I didn’t talk about topical, and there’s a huge universe of topical phrases, ways to think about things, problems that people have that are in this floating nexus around a service. And you wouldn’t find them on the same page often… TOFU, MOFU, BOFU — only down at the bottom does it overlap and it actually turn into money.” This SOP maps the bottom of the funnel. Run a separate clustering process for the top.
- Volume-first thinking is the failure mode this whole process exists to prevent. “Especially if you’re doing something very regional, there’s a lot of false positives in the raw data if you just use an SEO tool like Semrush or Ahrefs and you’re just looking at, oh, what’s the highest traffic potential, I should do that.”
- Do not turn the overlap list into ten more copies of what already ranks. A high overlap count tells you a phrase is relevant, not that you should rewrite the pages that rank for it. “You’ll go to the SERP and there’ll be 10 different versions of nearly exactly the same thing.” Use the list to pick targets, then differentiate: “I rank for this, they rank for that — I want to rank for what neither of us have touched.”
Tips for Efficiency
- Budget the manual version honestly, then decide whether to automate. Done by hand this is a real chunk of a day: “I was doing this manually and it took, you know, three hours for me to do it for a particular page.” If you will run it more than a few times, the exports are API-shaped — Jeremy’s automated version pulls the top 100 keywords for each of the ten URLs from DataForSEO and does the overlap maths on the way back.
- Keep the overlap count as a visible column forever, not a sort you throw away. It is the single number that separates this sheet from every other keyword export, and it stays useful for prioritising content long after the proposal is signed.
- Do the SERP checks in one batch, top-down. The highest-volume rows are both the ones most likely to be false positives and the ones that most distort the forecast. Check the top of the list first and you catch the expensive errors early.
- Ask the negative-keyword questions on the kickoff call, in plain language. “Is there anything we are not allowed to say?” gets you legal history, claim restrictions and brand rules in one pass, before anyone commissions a page that has to be pulled down.
- Reach for the “rank negative” move before you accept a traffic ceiling. When a phrase is banned as a self-description, it may still be usable as a comparison on the page. That is the whole space-cooler recovery, and it tripled traffic in three months without breaking the client’s red line.
- Store the CTR curve, the conversion rate and the sale rate as named assumptions at the top of the sheet. Every forecast is a chain of multiplications; if the inputs are buried in formulas, nobody can challenge them and you cannot re-run the model when the real conversion rate comes in.
- Remember the sheet is not the strategy. “You can drown in data. You can have oodles and oodles, you can have spreadsheets and spreadsheets of data, but it’s like being a captain of a ship and you’ve got all of your charts but you never plot a course.”
Sources & Relevant Episodes
- Jeremy Rivera, SEO Arcade — SEO practitioner since 2007, founder of the SEO ROI forecasting tool SEO Arcade, and at the time of recording an SEO product lead at CopyPress. He described this as his own go-to keyword research process, built by hand in Google Sheets long before it became software, and framed the overlap count as a signal Google hands you for free: “that’s the signal. That’s how I can tell that Google sees the overlap.” Every quote in this SOP is his, from the episode below.
- Read the full interview: Forecasting SEO Revenue Before You Rank: Jeremy Rivera — Unscripted SEO Interview Podcast, hosted by Mark A Preston, recorded 28 October 2023.
- Watch it: Jeremy Rivera on the Unscripted SEO Interview Podcast.
- The CTR reference used in step 8: Advanced Web Ranking — Google Organic CTR study.
- Related SOP: SEO SOP: Find the Traffic You’re Leaving on the Table with the Expected-CTR Model — the companion play for the click-through-rate stage of the forecast.
- More episodes: Unscripted SEO — 150+ interviews with SEO practitioners.
