Clusters keywords by SERP result overlap rather than semantic similarity, then rules on which pages to build, merge or drop.
/plugin marketplace add mkhalid1/locul-skills
Then /plugin to install SEO, which includes this skill.
This skill has not been measured. It has not been run against a control, and nothing below should be read as a claim about what it produces on real keyword data.
What it carries is a specific, checkable method rather than general clustering advice. It groups keywords by measuring the actual overlap between the pages already ranking for each one, using a stated working default, four or more shared URLs out of the top ten organic results, roughly a forty per cent overlap by a smaller-set calculation, rather than by how similar the keyword phrases look. It builds clusters as connected components on that overlap graph, which is a different and more defensible operation than pairwise grouping alone, and it names the specific risk that comes with it: transitive chaining, where two keywords end up in the same cluster only because each overlaps with a third, without ever overlapping with each other directly. It states, with the source, that Google has denied a duplicate content penalty since 2008 and documents the real mechanism as clustering and consolidation, which changes what "fix this cannibalisation" should actually mean inside the decision rule, rather than treating consolidation as damage control against a penalty that does not exist. And it treats Search Console's average position correctly, as a blended, impression-weighted figure across different SERP element types, rather than a number a strategist can safely rank keywords by.
This is not for a site with five or six pages and an obviously distinct topic for each, where the clustering question barely arises. It is also overkill for a single new article going onto an already well-mapped site, and it assumes access to keyword-level SERP result data that this skill does not fetch on its own.
Measuring one skill honestly costs about twenty model sessions: five runs with it, five without, on real material, each output graded alone by a session that is not told the other arm exists, against a rubric written by somebody who never saw the skill. We have not spent that on this one yet, so it ships labelled rather than ships silently.
How it would be measured. Tier A for the mechanical layer: five invented keyword lists of 30 to 60 terms each, paired with a fabricated top-ten result-URL set per keyword, checking whether the overlap procedure groups keywords to match a hand-built ground-truth clustering at the stated threshold, whether the decision rule's 'cannot tell' branch fires on a deliberately mixed-evidence case, and whether the folklore claims it touches, the duplicate content penalty, keyword density, LSI keywords, and average position, are stated correctly rather than repeated. Whether the resulting page briefs are what an experienced information architect would actually choose to build needs blind pairwise preference judging (Tier B) on top of that mechanical layer.
A fair test needs fabricated SERP result sets, not just a keyword list, because the whole method depends on which URLs actually rank for each term. That means building fixtures: keyword lists paired with invented top-ten result sets, plus a hand-built ground-truth clustering to grade the overlap procedure against, since no public dataset states which keywords "should" share a page. The mechanical layer, the overlap threshold and the decision rule's branches, is gradable that way. Whether the resulting page map is what an experienced information architect would actually choose to build is a harder, preference-shaped question, and it is also the honest reason this skill might not add much on its own: a capable model asked to group a keyword list already produces something plausible-looking, and the open question a fair test would answer is whether measuring SERP overlap actually changes that grouping, or just relabels the same one.
The rule that decides pass or fail was written down before any run was executed and it does not move afterwards. It is in the method note on the hub, along with the full results table including every skill that was tested and cut.
Stated plainly, because a skill that claims everything is useful for nothing.
name and description in the file's frontmatter, so you can also invoke it by name.Keyword Insights (keywordinsights.ai) is the closest direct alternative: a paid tool built specifically to cluster keywords from live SERP data rather than semantic similarity, with an adjustable overlap setting. Pick it over this file when the list runs into the thousands and pulling and comparing that many result sets by hand stops being realistic.
Ahrefs Keywords Explorer's Parent Topic grouping does a related job, assigning keywords to the page that already ranks best for the highest-volume term in a loose group. It is faster to use but is not the same computation this file describes: it does not check overlap between every pair, so it can miss a genuine split inside what it treats as one topic. Pick it when speed matters more than precision on the merge and split calls.
For a list under a couple of hundred keywords, a competent SEO strategist with a spreadsheet, a search engine open in a private window, and an afternoon can run this exact method by hand, and often should. The procedure below does not require software, only patience and a consistent way of recording what actually shows up in the results.
--- name: keyword-cluster-map-builder description: Takes a raw keyword list, exported from a keyword research tool or pulled from Search Console's query report, and turns it into a page map, which keywords belong on the same page, which need a page of their own, and which existing pages should be merged into another or retired. It clusters keywords by measuring how much the pages already ranking for them actually overlap, rather than by how similar the keyword phrases look, and it names the cases where the evidence is too mixed to decide. This skill should be used when a keyword list exists and nobody has yet decided which pages to build, merge or drop from it, before any brief or article gets written. --- # Keyword cluster map builder ## The claim, and why grouping by word similarity fails Two keywords belong on the same page when the pages that actually rank for them are substantially the same set. They belong on different pages when that set is substantially different, regardless of how close or how far apart the words themselves look. That is the entire claim this skill is built on, and it runs against how almost every keyword tool actually groups terms, which is by string similarity or embedding distance: how alike the words look or sound, not what a search engine has already decided to show for them. The reason that default is the wrong unit is straightforward once stated. A search engine does not rank pages by matching word shapes; it ranks by an assessment of intent, and two queries that look almost identical can carry different intent while two that look nothing alike can carry the same one. "How to write a will" and "will template" read as obviously related, but the first commonly returns explainer articles and legal-process guides, while the second commonly returns downloadable documents and paid drafting services: different jobs, different formats, different pages. "Shift scheduling software" and "rota planning software" look like different vocabularies entirely, one American, one British, and a semantic pass has no reason to merge them. But if the underlying product category and market are the same, the pages ranking for both are very often close to identical, because the search engine has already done the intent-matching work. Semantic grouping keeps those two apart. Results-page evidence puts them together. The cost of getting this wrong runs in both directions. Split a shared intent into two pages and both compete for the same ranking opportunity, diluting signals either one could have accumulated alone. Merge two different intents into one page and it tries to serve two searchers at once, satisfying neither. Neither mistake shows up by reading the keywords. Both show up by reading the search results. This is a build skill. It produces the map: which pages to write, which to leave alone, which to merge, which not to build at all. It does not audit an existing site's content quality, and it does not write the pages themselves once their jobs are set. ## What the map needs before you start Four inputs, and being honest about which ones this skill cannot supply itself matters more than any step in the method: 1. **The keyword list**, deduplicated of exact repeats but not yet trimmed of anything that merely looks similar to something else on the list. Trimming on appearance is the mistake this method exists to avoid, so every keyword survives unless a human is confident it is literally the same string as another, not merely the same idea. 2. **The organic results ranking for each keyword.** The top ten, or top twenty on a thin or noisy SERP, pulled from a rank tracker, a SERP-scraping tool, or a manual search recorded by hand for a small list. Exclude paid results: Search Console's documentation is explicit that ads do not occupy a search position at all. Exclude blended feature boxes too, since a shared "People also ask" box is weak evidence next to two shared ranking URLs. 3. **Search volume for every keyword.** This skill has no way to produce that number. It has to come from Google Keyword Planner, a third-party keyword tool, or Search Console impression counts used as a rough proxy where the site already gets some visibility. 4. **A list of the site's own existing pages** and whatever they currently rank for, so the merge, expand and drop decisions later have something real to check against. Partial data does not stop the process. It feeds the "cannot tell" branch of the decision rule below, and those keywords get held rather than forced into a wrong answer. ## The method, in order Order matters here specifically. Normalise before computing overlap, or near-duplicate spellings and word-order variants inflate the graph with nodes that were never independent evidence. Compute overlap before clustering, or the grouping ends up anchored on the same word-similarity judgement this method exists to replace. Cluster before writing each page's job, or the job gets decided from assumption rather than from what is actually ranking. Check the map against existing pages last, not first, or the process quietly optimises for minimal disruption to what already exists rather than for what the evidence actually supports. **Phase 1: normalise the list.** Deduplicate only exact and near-exact duplicates, plurals, obvious misspellings, and reordered versions of the same phrase that a search engine treats as one query. Leave everything else, including pairs that look related, untouched. The next phase is where relatedness gets decided, not this one. **Phase 2: pull the organic results.** Collect every keyword's top-ten result set from the same location, the same device type, and as close to the same session as practical. A set pulled from a desktop search in one country and another pulled from mobile in a different one will differ for reasons that have nothing to do with topical relationship, and that difference will corrupt every overlap calculation built on it. While collecting, note the SERP's format mix too, a run of listicle titles, a run of product or category pages, a run of long single-answer articles, since this becomes evidence later even when the raw URL overlap alone is ambiguous. **Phase 3: compute pairwise overlap.** For every pair of keywords in the batch, calculate the overlap as the count of shared URLs divided by the smaller of the two result-set sizes, not by their union. The smaller-set denominator avoids penalising a pair where one keyword's results are more fragmented across more domains than the other's, which happens often between a broad head term and a narrower long-tail variant of the same intent. State the default threshold plainly: four or more shared URLs out of ten, roughly forty per cent by this method, counts as the same intent. This number is a working default for this method, not a published search engine constant; no search engine documents an official same-page threshold, so treat it as adjustable. Lower it, to roughly thirty per cent, on thin SERPs where ten genuinely distinct organic URLs rarely exist once repeats from one dominant domain are excluded. Raise it, to roughly fifty per cent, on commercial terms where every competitor's homepage and category pages compete for the same slots and a lower bar would merge products that are not actually the same search. **Phase 4: build clusters as connected components, not as pairwise groups.** Treat every keyword as a node and draw an edge between two keywords whenever their overlap crosses the threshold. A cluster is a connected component of that graph: keyword A and keyword C can land in the same cluster even if they never crossed the threshold with each other directly, provided both crossed it with keyword B. Name the risk this creates explicitly, because it is the most common way this phase goes wrong: transitive chaining, where a broad, non-representative keyword in the middle links two genuinely different intents into one cluster. Check any cluster with more than four or five members by hand. If the highest-overlap pair and the lowest-overlap pair inside it share almost nothing directly with each other, the cluster has probably chained through that middle keyword and should be split at the weak link rather than accepted whole. **Phase 5: read each cluster's page job off its own SERP, not off an assumption.** Look at what format dominates the top five shared results. A run of numbered or listicle titles points toward a listicle. A run of tool or software homepages points toward a commercial landing or comparison page. A run of long single-answer articles points toward one definitional or how-to piece. A mix of retailer product pages points toward a buying guide or comparison table. Write the job as one sentence: who is searching, what they are trying to do, and which format the evidence favours. That sentence is the handoff to whichever skill takes a single brief through to a finished draft; this skill's job stops at the keyword-to-page mapping and the sentence describing what each page is for. **Phase 6: check the map against the site's existing pages.** For each cluster, look for an existing page that already ranks for any keyword inside it. Three outcomes follow: no existing page, build new; one page roughly matching the cluster's job, keep and expand rather than start a new one; two or more existing pages splitting the cluster, a consolidation candidate. State the reason for consolidating correctly, because published advice usually gets it backwards: it is not to avoid a duplicate content penalty, since Google has denied that one directly since 2008. The documented mechanism is signal dilution, ranking signals such as links splitting across near-duplicate URLs instead of accumulating on one. The documented fix is a redirect from the weaker page to the stronger one, since redirects and canonical annotations are the two signals Google's documentation calls strong, not a side-by-side rewrite of both pages. **Phase 7: decide what not to build.** A cluster clearing every earlier phase is still not automatically a page. Drop or defer a cluster when its combined volume sits below whatever floor makes production cost worthwhile for this site, a threshold this skill will not invent on the site's behalf; when Phase 6 already found a page doing the job, so the honest fix is expanding it rather than adding competition; or when the cluster's evidence is genuinely mixed, covered next. This list of what not to build is often the single most useful line in the deliverable, since it is the one thing a stakeholder pushing for more content will not produce unprompted. ## The decision rule Apply this to any pair of keywords, or to a cluster being checked against an existing page. **Branch A, same page.** Overlap sits at or above the threshold, and the SERP composition is consistent between the two, a similar mix of formats with no sharp swing from one shopping-heavy result set to a definition-heavy one. Merge into one cluster with one job statement. **Branch B, same page but flagged.** Overlap sits at or above the threshold, but the format mix diverges sharply between the two keywords. Check the actual URLs shared, not just the domains, before finalising. A single large publisher can rank for both a shopping query and an unrelated review query on the same domain, which inflates domain-level overlap without the two keywords sharing real intent. Merge only once the shared URLs themselves, not just the shared sites, hold up. **Branch C, same page on stronger evidence than the raw number.** Overlap sits below the threshold, but an existing page on the site already ranks for both keywords without the site cannibalising itself, showing two of its own URLs for the same query. Weight this above the general-population overlap figure. A page already succeeding at both is stronger evidence of shared intent than a SERP built from competitors who have never tried serving both with one page. **Branch D, separate pages.** Overlap sits below the threshold and no existing page already serves both. Write two job statements and treat them as independent. **Branch E, cannot tell.** This is the branch a two-way yes-or-no rule always misses, and it fires for three distinct reasons rather than one vague feeling of uncertainty. First, thin evidence: one or both result sets returned fewer than eight distinct organic URLs because a single domain occupies several slots, leaving too little independent evidence to compute a meaningful ratio. Second, collection drift: the two result sets came from different locations, devices or sessions, so any apparent difference between them may be an artefact of how they were gathered rather than a real difference in intent. Third, threshold-edge cases: the overlap sits within roughly ten percentage points of the cutoff in either direction, close enough that a working default set for convenience should not be asked to carry the decision alone. In any of these three cases, do not force Branch A or Branch D. Hold the pair in a staging list with the specific reason attached, then either recollect the data properly, same location, same device, a fuller result set, or route it to a person for a direct side-by-side read of both SERPs. ## Worked example An invented shift-scheduling software company has a twelve-keyword list drawn from a keyword tool export: shift scheduling software, rota planning software, employee scheduling app, free shift schedule template, how to create a work schedule, weekly rota template excel, staff rota app, work schedule maker, employee scheduling software comparison, best shift scheduling app, rota vs schedule, and shift swap app. Pulling the top ten organic results for each, from the same country and device, produces this pattern. "Shift scheduling software" and "rota planning software" share seven of ten URLs, well above the forty per cent default, despite sharing no words: same vendors, same category pages, same comparison articles. "Employee scheduling app", "staff rota app" and "work schedule maker" form a second tight group, six or more shared URLs across every pair, all app-store listings and vendor homepages. "Employee scheduling software comparison" and "best shift scheduling app" overlap heavily with that group too, at five and six shared URLs, both showing listicle-style comparison articles rather than plain homepages, so under Phase 5 that sub-cluster's job reads as a comparison page. "Free shift schedule template" and "weekly rota template excel" share eight URLs with each other, all download pages and template galleries, and almost nothing with the software cluster: a second, separate page. "How to create a work schedule" returns long how-to articles and shares only two URLs with the template cluster and one with the software cluster, so it sits in its own third cluster. "Rota vs schedule" returns only six distinct organic URLs after excluding three repeats from one dictionary-style domain, below the eight-URL floor for a trustworthy ratio, so it lands in Branch E on thin evidence. "Shift swap app" shares four URLs with the app cluster and three with the software cluster, close enough to the threshold on both sides to also land in Branch E, this time on the threshold-edge reason. The verdict: build one commercial comparison page covering the software and app terms as a single cluster of eight keywords, build one separate template-download page for the two template terms, keep "how to create a work schedule" as its own how-to page since no existing content covers it, and hold "rota vs schedule" and "shift swap app" in a staging list pending better data rather than assigning either one now. ## Failure modes **Semantic clumping.** Grouping the list by a keyword tool's default topic clusters or embedding similarity, then treating that grouping as final without ever checking what actually ranks. It looks identical to correct output until a genuinely different-intent pair like "will template" and "how to write a will" gets merged, or a same-intent pair like the US and UK scheduling terms gets left apart. **Threshold theatre.** Treating the forty per cent default as a documented constant rather than a working default this method sets for itself, then defending a borderline call by citing the number as though it were a search engine rule. No search engine publishes one, and stating otherwise undermines the file's own credibility on every other claim in it. **Location or device drift.** Pulling one keyword's results from a desktop search in one market and another's from mobile in a different one, then reading the resulting difference as evidence about intent rather than about collection. The fix is procedural, not statistical: collect the whole batch under matching conditions before computing anything. **The penalty myth driving the wrong fix.** Believing pages must be merged to avoid a Google duplicate content penalty, when Google has denied that penalty exists since 2008 and names the real risk as signal dilution instead. This produces two opposite errors: panicked consolidation of pages that are not actually competing, and refusal to write a genuinely distinct second page out of fear of a penalty that was never real. **Average-position false prioritisation.** Ranking clusters by Search Console's average position column without correcting for what it measures: an impression-weighted average across blended SERP elements, not a rank the property holds. A cluster's average position can look strong purely because its weakest impressions never registered, not because anything is actually ranking well. **Single-day snapshot with no expiry.** Building the entire map from one SERP pull and treating it as permanent. Search results move, and a map built in one week can misdescribe the following month's evidence, particularly for competitive commercial terms. The map needs a stated recheck point, not an assumed shelf life. **Volume laundering.** Inventing a plausible-sounding search volume for a keyword instead of sourcing one, then letting that invented number decide whether a cluster clears the production-cost floor in Phase 7. An invented number dressed up as data is worse than an honest gap, because it looks decided when it was guessed. **Orphan cluster.** Finishing the map with a cluster whose page has a clear job and no plan for how anything on the site will link to it. A page can answer the right query and still go undiscovered internally if nothing on the site points to it, which is exactly the boundary where this skill's job ends and an internal linking pass has to begin. ## What this skill does not do It does not fetch live SERP results itself. Every overlap calculation in this method depends on result data supplied from a rank tracker, a scraping tool, or manual collection, not generated by this skill. It does not produce search volume figures. Every number used to prioritise or to drop a cluster has to be sourced externally and brought in as an input, not invented to fill a gap. It does not write the article, landing page or product copy once a cluster's job is set. It produces the keyword-to-page map and the one-sentence brief for each page, and stops there. It does not design the internal linking between the pages it recommends. A map that says which pages should exist says nothing about how a visitor or a crawler actually reaches them. It does not evaluate a site's crawl budget, technical indexing health or CMS redirect capability. A consolidation call it recommends is a content decision, and still needs a technical check before a redirect or removal goes live. It does not diagnose why an existing, already-mapped page has started losing traffic. That is a different investigation with its own method, not a rerun of the clustering procedure on the same keyword list.
These skills all ask your assistant to check things against your actual codebase, your actual schema, your actual design system. Locul keeps that context current on its own, from the files you already have, on your machine. Mac and Windows, free to start.