---
name: research-keywords
description: >-
  The discovery layer for Search: builds a comprehensive, de-duplicated keyword list with volume, competition, and
  CPC data by expanding business-derived seeds through the planner tools (keyword and URL-seeded ideas, thematic
  grouping, forecasts), mining the account's own converting search terms, and layering in AI brainstorming that is
  always volume-validated before inclusion. Output is a master list with intent labels, ready for clustering. Use
  it for new campaign builds, new product/service expansion, relaunches after an offer change, or client
  onboarding. It does not group keywords into ad groups (cluster-and-map-keywords), assign match types
  (build-search-campaign-structure), or harvest single queries from live traffic on an ongoing basis
  (promote-search-terms-to-keywords). Third-party SEO tools and autocomplete scraping are optional human add-ons.
---
# Keyword Research and Discovery

## Purpose
Everything downstream, clusters, ad groups, ads, budgets, inherits the quality of the initial keyword list. This
skill produces that list systematically: seeds grounded in what the business actually sells, expansion through
multiple independent sources so no single tool's blind spots become the account's blind spots, real queries mined
from account history, and a cleanup pass that leaves every keyword carrying volume, cost, and intent metadata.

## When to run
- A new Search campaign is being planned from zero.
- The account is expanding into a new product or service area.
- A relaunch follows a major offer or business-model change (old keyword assumptions are void).
- A new client is being onboarded and needs their keyword foundation built.

## When NOT to run
- The list exists and needs organizing → `cluster-and-map-keywords`.
- Ongoing harvesting of individual winners from live traffic → `promote-search-terms-to-keywords`.
- The question is match types or structure → `build-search-campaign-structure`.
- A quick top-up of one existing theme, a full research pass is overkill; mine the term report and move on.

## Prerequisites
- Business context: products/services (ranked by priority), target audience, USPs, conversion goals.
- Website URL(s), plus 2–3 competitor names or URLs.
- Google Ads account access (planner-backed tools require it).

## Procedure
1. **Document the business frame.** HUMAN STEP (interview the user if not on file): what is sold, to whom, what
   differentiates it, what counts as a conversion, and which offerings matter most. Every later
   relevance judgment leans on this, do not skip to tooling without it.
2. **Build 10–30 seed keywords** from three angles: (a) offering names plus the plain-language variants customers
   actually use (not internal jargon); (b) the site's own vocabulary, titles, headings, page themes; (c) how
   competitors describe the same things, including phrasings the business itself never uses. Record each seed
   with its source and the offering it maps to. Gate: every priority offering has at least 2 seeds.
3. **Expand seeds through the planner.** Call `gads_keyword_ideas` in batches of 3–5 related seeds per call,
   scoped to target geography and language; collect volume, competition, and CPC per idea; drop zero-volume rows.
   Then run the URL-seeded mode of `gads_keyword_ideas` against the business site AND each competitor site, keyword seeds surface what you already suspected, page seeds surface what you missed. One expansion set per
   priority offering, minimum.
4. **Group and gap-check with themes.** Run `gads_keyword_themes` over the growing pool to see the thematic
   shape of the space; themes with thin coverage indicate a seed gap, loop back to step 3 with new seeds for
   those themes.
5. **Layer external discovery.** HUMAN STEP (outside VigilDog, optional but valuable): autocomplete and
   related-search capture in an incognito session for each seed; third-party SEO suites for competitor paid/
   organic terms and gap reports; long-tail generators for question/preposition variants. Import anything
   relevant with its source marked.
6. **Brainstorm with the model, validate with the planner.** Generate additional candidates from the business
   frame: synonyms, problem-phrasings, question forms, competitor-adjacent terms. These are hypotheses, not
   keywords, every AI-suggested term must show real volume via `gads_keyword_ideas` before it enters the master
   list. Plausible-but-unsearched terms are the known failure mode of this source.
7. **Mine the account's own history** (skip for brand-new accounts). `gads_get_search_terms_report` over the last
   90 days: terms with 1+ conversions or 50+ impressions that are not already keywords and not close variants of
   existing ones join the list, tagged with their conversion data, these are the only entries with proof
   attached, and they get priority weighting in clustering.
8. **Merge and de-duplicate.** One row per keyword, lowercase-normalized, source column preserved. Remove exact
   duplicates, then near-duplicates (singular/plural, trivial reorderings), keeping the highest-volume form.
9. **Complete the metrics.** Every surviving keyword needs volume, CPC estimate, and competition level; backfill
   gaps via `gads_keyword_ideas` (historical metrics mode) and use `gads_keyword_forecast` on the shortlist to
   sanity-check expected clicks and cost at realistic bids, an early economics read before anything is built.
10. **Label intent and finish.** Tag each keyword transactional (buy/order/pricing signals), commercial
    (best/vs/review), informational (how/what/guide), or navigational (brand names). Cut remaining zero-volume
    rows (except conversion-proven term-report entries) and any stragglers that are obviously irrelevant, filing
    the latter on a rejects sheet, which seeds the negative lists downstream. Deliver the master sheet: keyword,
    volume, CPC, competition, source, intent, plus the rejects sheet. This is a read-only skill, no Google Ads
    writes, so no approval gates apply; if any write becomes necessary, consult `gads_policy_guardrail` first.

## Decision rules
- **Seed count 10–30, every priority offering ≥2 seeds**, below that, expansion inherits the gaps.
- **Multiple independent sources are mandatory,** because each source has a systematic blind spot: planners
  under-surface long-tail, autocomplete over-surfaces questions, AI invents. Overlap between sources is a
  quality signal, not waste.
- **AI terms enter only after volume validation.** No exceptions.
- **Term-report mining threshold:** conversions ≥1 or impressions ≥50, not already covered, not a close variant.
- **De-dup keeps the highest-volume form** of near-duplicate sets.
- **Intent labeling is mandatory before handoff:** transactional and commercial terms are the Search payload;
  informational terms survive only flagged, for the conversion-path scrutiny that clustering applies.
- **Zero-volume keywords are dead weight** unless account history proves they convert.

## Common failure modes
- **Single-tool research.** A planner-only list misses the phrasings real users type; the multi-source rule is
  the whole method.
- **Jargon seeds.** Seeding with internal product-speak expands into more jargon; seed with customer language.
- **Unvalidated AI keywords** padding the list with terms nobody searches, they cluster fine and then deliver
  zero impressions forever.
- **Skipping the term-report mine on an existing account,** ignoring the only keywords with attached proof.
- **Deferring de-duplication** to clustering, where triplicate near-variants slow every judgment.
- **Unflagged informational terms** slipping through to spend budget answering questions with no conversion path.
- **Treating this as one-and-done:** research decays; new offerings and market language shifts warrant a re-run,
  and live harvesting continues via the promotion skill in between.

- **Merging without source tags,** losing the ability to weight proof-backed term-report entries above
  speculative AI suggestions during clustering.
- **Letting competitor phrasing dominate.** Competitor sites are a vocabulary source, not a strategy source;
  terms describing what THEY sell still need the relevance frame from step 1.

## What done looks like
- Business frame on file; 10–30 seeds with ≥2 per priority offering.
- Planner expansion done in both keyword-seeded and URL-seeded modes, own site plus 2–3 competitors.
- Account history mined (where history exists) with conversion data carried on each entry.
- One de-duplicated master sheet: keyword, volume, CPC, competition, source, intent, plus the rejects sheet.
- Zero unvalidated AI terms and zero unlabeled rows in the final list.

## Handing off
The master sheet goes to `cluster-and-map-keywords` unmodified; the rejects sheet travels alongside it and is
applied during `build-search-campaign-structure` as the seed of the shared irrelevant-terms list. If clustering
later exposes a thin theme, this skill re-runs scoped to that theme rather than from scratch.

## Related skills
- Run after: `cluster-and-map-keywords` (consumes the master list), then `build-search-campaign-structure`.
- Parallel lifecycle: `promote-search-terms-to-keywords` (continuous discovery from live traffic),
  `run-n-gram-analysis` (the rejects sheet's spiritual successor once traffic flows).
- Economics context: `calculate-and-validate-unit-economics`, CPC viability judgments in clustering need it.
