Keyword Research
Keyword Research with Claude Code: From Seed Terms to Page Maps
Build a defensible keyword research workflow in Claude Code using real data, intent classification, SERP overlap, clustering, and page mapping.
Keyword research with Claude Code works best as data engineering plus editorial judgment. Claude Code can combine exports, normalize messy phrases, classify intent, compare result-page overlap, build clusters, and map those clusters to a website. It cannot produce dependable search volume or competition data from its language model memory. If the workflow begins with invented numbers, every prioritization step that follows is false precision.
The goal is not a giant keyword list. The goal is a defensible page map: which audience need deserves a page, what that page must accomplish, which existing URL should own it, and how success will be measured. This guide builds that map step by step.
Begin with the market, not a tool export
Write a one-page market brief before collecting keywords. Define the product, ideal customer, buying triggers, alternatives, geography, language, sales motion, and conversion action. Add the vocabulary customers use in calls, tickets, reviews, communities, and internal site search. This context helps distinguish a commercially relevant query from a popular but useless one.
For a Claude Code SEO product, “Claude Code” is related but extremely broad. A visitor seeking installation help may not want an SEO workflow. Terms such as “Claude Code technical SEO audit,” “Claude Code keyword research,” and “SEO skill for Claude Code” express narrower tasks and different stages of awareness. The site needs to decide which of those needs it can satisfy better than the current results.
Document negative scope too. If the product only supports Claude Code, queries specifically seeking a ChatGPT plugin or WordPress SaaS may be poor targets even when their wording seems adjacent.
Collect seeds from multiple evidence sources
A seed list should represent users, the product, and the search market. No single source covers all three.
| Source | What it reveals | Useful fields | Common bias |
|---|---|---|---|
| Search Console | Queries already associated with the site | query, page, clicks, impressions, position, country, device | Limited to existing visibility |
| Analytics and conversions | Landing pages that create business value | page, channel, engagement, lead or revenue | Often hides the exact query |
| Sales and support language | Problems and objections in the customer's words | phrase, persona, stage, frequency | Sample may skew toward current customers |
| Product taxonomy | What the business actually offers | feature, use case, audience, integration | Internal naming may not match search language |
| Competitor pages | Topics the market publishes and links to | title, heading, URL type, positioning | Competitors can target the wrong things |
| Keyword platform | Estimated demand and SERP metrics | volume, trend, difficulty, CPC, results | Estimates vary by provider and market |
| Live search results | Current intent and content format | ranking URLs, result types, freshness | Results vary by locale, device, and time |
Export source labels with every phrase. A row should never become detached from where it came from. Add locale and collection date because a US-English dataset from August is not interchangeable with a global export from January.
When privacy matters, remove names, emails, account IDs, and full support transcripts. Extract the problem language needed for research and retain the original material in its approved system.
Create a clean keyword table
Put all seeds into a common schema. A practical table includes:
keyword
normalized_keyword
source
country
language
device
volume
trend
difficulty
cpc
current_page
current_clicks
current_impressions
intent
funnel_stage
cluster
target_page
decision
notesKeep unknown values empty. Do not convert missing volume into zero because zero is a factual value. Store provider metrics in separate columns if you combine tools; “difficulty” is not standardized across platforms.
Ask Claude Code to normalize Unicode, whitespace, obvious casing differences, and repeated punctuation. Preserve the original query in another column. Do not stem phrases so aggressively that meaning disappears: “SEO audit tool” and “SEO auditor job” share tokens but not intent.
Create deterministic deduplication rules first, then review fuzzy matches. Exact normalized duplicates can merge safely while preserving all source labels. Near duplicates should remain separate until intent and result overlap are examined.
Expand seeds without producing noise
Use controlled expansion patterns tied to the product and customer journey. Useful modifiers include:
- Task: audit, research, cluster, optimize, write, monitor, fix.
- Problem: duplicate content, missing schema, slow pages, content decay, keyword cannibalization.
- Audience: developer, founder, agency, in-house SEO, content team.
- Comparison: alternatives, versus, best, free, paid, template, plugin, skill.
- Platform: Claude Code, Next.js, Shopify, WordPress, Search Console.
- Stage: what is, how to, checklist, workflow, tool, pricing, review.
Expansion is a candidate-generation step, not validation. Claude Code can combine relevant entities and modifiers, but candidates need a source signal or a live result review before they enter the roadmap. Discard awkward phrases that no real user would type or that simply permute the same words.
Question mining is valuable when it reveals decision barriers. “Can Claude Code crawl JavaScript sites?” suggests a capability concern. “How do I validate a canonical tag?” suggests a task. Each may become a subsection in a larger guide rather than a separate page.
Classify search intent with evidence
Intent labels should describe the outcome visible in the result page, not only the grammar of the query. Use a compact taxonomy:
- Learn: understand a concept or problem.
- Do: complete a specific task or follow a process.
- Compare: evaluate approaches, products, or alternatives.
- Buy: select, price, install, or obtain a solution.
- Navigate: reach a known brand, product, or page.
Add a content-format label such as guide, product page, category, comparison, template, video, tool, or documentation. Search results are the best available evidence for what format currently satisfies the query. If product pages dominate, a generic educational article may struggle to match the same job.
Claude Code can assign preliminary intent from the phrase and known results. Require a confidence field and manually review valuable or ambiguous clusters. A query such as “Claude SEO” may be a broad concept, a brand lookup, or a product search depending on the result set.
Use SERP overlap to build clusters
Semantic similarity helps generate candidate groups; ranking-page overlap helps validate them. For each priority query, collect a consistent number of organic results from the same market and time window. Normalize ranking URLs by removing tracking parameters and standardizing host and path rules.
Calculate overlap between two result sets. Jaccard similarity is a simple measure:
overlap = shared ranking URLs / unique ranking URLs across both queriesIf “Claude Code SEO audit” and “technical SEO audit with Claude Code” share many strong results, they may belong to one page. If “Claude Code SEO pricing” returns product pages while “how to use Claude Code for SEO” returns tutorials, they deserve separate targets.
Do not use one universal threshold. Narrow result sets, local packs, video-heavy results, and brand queries need judgment. Record overlap as evidence alongside semantic similarity and intent, then name the cluster for the user job rather than the highest-volume phrase.
| Relationship | SERP pattern | Recommended decision |
|---|---|---|
| Same intent | High overlap and same dominant format | One primary page with natural variants |
| Related subtask | Moderate overlap; one query is narrower | One guide with a substantial subsection, or a child guide if depth warrants it |
| Different intent | Low overlap or different formats | Separate pages with distinct purpose |
| Ambiguous | Volatile or mixed results | Delay, test, or choose the audience's clearest job |
| Irrelevant | Results serve another market | Exclude even if tools show volume |
Map clusters to existing pages first
Before planning new content, inventory indexable pages with title, H1, canonical, organic queries, conversions, internal inlinks, and backlinks. Assign every approved cluster to the best existing candidate where one exists.
Use four decisions:
- Keep: the existing page already owns the intent and needs minor refinement.
- Improve: the URL is right but lacks coverage, clarity, evidence, or experience.
- Consolidate: multiple pages compete for the same intent and should become one stronger resource.
- Create: no suitable URL exists and the cluster supports a distinct valuable job.
A fifth decision, exclude, protects the strategy from irrelevant opportunities. It is a sign of focus, not lost traffic.
When several existing pages receive impressions for the same query, investigate cannibalization rather than assuming it. Pages can alternate because of seasonality, localization, weak signals, or genuinely mixed intent. Compare their purposes and result-page overlap. Consolidate only when one page can satisfy the jobs without harming a distinct audience need.
Score opportunities without pretending the math is truth
A prioritization model makes assumptions visible. It does not eliminate judgment. Score a cluster on business fit, intent fit, evidence of demand, current authority, achievable differentiation, conversion potential, effort, and risk.
One transparent formula is:
opportunity =
(business_fit * 0.25) +
(intent_fit * 0.20) +
(demand_evidence * 0.15) +
(conversion_potential * 0.15) +
(differentiation * 0.15) +
(existing_authority * 0.10)
priority = opportunity / max(effort * risk, 1)Use a consistent scale and write the rationale. If the team disagrees with an output, change the input or weighting openly rather than manually overriding a hidden score. Review high-value low-volume topics; they can be commercially important in specialized software markets.
Estimated volume should not dominate the model. Search Console impressions, customer frequency, paid-search conversion, and sales evidence can establish demand even when third-party tools show little data.
Create a five-page topic cluster without cannibalization
For a Claude Code SEO product, a coherent first cluster could look like this:
| Page | Primary job | Funnel role | Internal links |
|---|---|---|---|
| Claude Code for SEO | Understand the full operating workflow | Broad discovery | Links to all focused guides and product |
| Technical SEO audit with Claude Code | Diagnose and fix crawl/index issues | Task and solution evaluation | Links to core guide and product audit features |
| Keyword research with Claude Code | Build clusters and page maps | Task and process | Links to content workflow and core guide |
| Generative engine optimization guide | Understand AI-search readiness | Concept and strategy | Links to technical and content guides |
| Claude Code content writing workflow | Produce and verify useful articles | Task and solution evaluation | Links to keyword research and product |
The topics share an audience but own different jobs. Their titles, introductions, headings, examples, and calls to action should preserve that distinction. Internal links explain the relationship rather than repeating the same anchor everywhere.
Turn each cluster into a content brief
A keyword is not a brief. Build a document that helps an expert create the best answer for a specific audience. Include:
- Primary audience and situation.
- Job to be done and search intent.
- Primary query and natural secondary language.
- Existing target URL or proposed slug.
- Dominant result formats and gaps.
- Unique angle and first-party evidence available.
- Required sections, questions, tables, examples, and limitations.
- Primary sources to verify.
- Internal links in and out.
- Conversion action that follows naturally.
- Acceptance criteria and measurement plan.
Ask what the page can contribute that is not commodity knowledge. It may be a tested workflow, a downloadable template, a real failure analysis, original data, annotated screenshots, or a clear decision framework. Google's guidance for AI search emphasizes useful, unique, non-commodity content rather than special tricks: Google's AI optimization guide.
Use Claude Code to process the dataset reproducibly
Keep scripts and transformations in the project so the research can be rerun. A useful pipeline is:
01-import -> 02-normalize -> 03-enrich -> 04-intent
-> 05-serp-overlap -> 06-cluster -> 07-page-map -> 08-briefsEach step should read an immutable input and write a new dated output. Generate a quality report containing row counts, duplicates, missing metrics, invalid locales, clusters without targets, and pages assigned conflicting primary intents. This makes silent data loss visible.
Use version control for scripts, schemas, decisions, and briefs, but not secret keys or restricted provider data that the license forbids committing. Keep a data dictionary explaining every column and metric source.
Validate the research manually
Review the highest-priority clusters in the live search market. Open the ranking pages and record what they genuinely provide, not only their headings. Check result features, freshness, brands, content type, and whether the query appears settled or mixed.
Then interview the business. Product, sales, support, and subject-matter experts can expose false assumptions: a high-volume feature may be discontinued, a low-volume integration may close valuable deals, or a proposed tutorial may attract users the product cannot serve.
Run a cannibalization review before publishing. Compare the new brief with every existing page assigned related terms. State the unique intent in one sentence. If the sentences are indistinguishable, revise the scope or consolidate.
Measure by page and cluster
After publishing or improving a page, annotate the date and monitor its query set, not only the exact primary keyword. Track impressions, clicks, click-through rate, position distribution, conversions, and internal-link discovery. Compare equivalent periods and account for brand campaigns or seasonality.
At the cluster level, measure whether the site gains coverage across the intended journey. A broad guide may assist discovery while a technical guide earns fewer visits but more product exploration. Both can be successful.
Do not rewrite a page every time daily position changes. Wait for enough crawl and impression data, inspect which queries moved, and change the page only when evidence identifies a mismatch or missing value.
A reusable Claude Code request
Build a keyword-to-page map from the supplied evidence.
Inputs:
- market brief: research/market.md
- Search Console export: research/gsc.csv
- keyword provider export: research/provider.csv
- current URL inventory: research/pages.csv
- SERP results: research/serps/
Rules:
- never invent missing volume or difficulty
- preserve source, country, language, and collection date
- use live SERP overlap to validate priority clusters
- prefer improving an existing URL over creating a duplicate
- assign keep, improve, consolidate, create, or exclude
Deliver:
- cleaned master dataset
- intent and cluster QA report
- prioritized page map with rationale
- briefs for approved create/improve decisionsReview the results before asking for articles. Research and production are separate quality gates.
What good keyword research produces
Good research leaves the team with fewer ambiguities, not merely more rows. Every important cluster has a source, audience, intent, decision, target, rationale, and measurement plan. Existing pages receive attention before new inventory expands. Closely related topics have explicit boundaries and useful internal links.
Claude Code makes the process faster and more reproducible by handling joins, transformations, comparisons, and documentation. The durable advantage still comes from real customer language, current search evidence, a differentiated product, and editorial decisions that choose what not to publish.