---
name: rsa-testing-with-the-iteration-loop
description: >-
  The ongoing, hypothesis-driven RSA testing system: enforce a shared slot template across a cluster of
  same-intent ad groups, pool asset-level data across the cluster to beat the small-sample problem, classify
  every asset on conversion efficiency versus how often it actually serves, then cut the drags, boost the
  underexposed winners, and deploy the next round of variants. Run it on a 2-8 week cycle once an ad group has
  a healthy, testing-ready RSA. It is a continuous optimization loop, not a one-time fix: use write-compelling-rsas
  to build the initial ad, improve-ad-relevance or improve-expected-ctr for below-average quality components,
  and run-a-creative-testing-cycle when coordinating testing across PMax, Demand Gen, and video too.
---
# Systematic RSA Asset Testing Across Ad-Group Clusters

## Purpose
Single ad groups rarely generate enough impressions to prove anything about individual headlines. This skill
solves that by templating: give every same-intent ad group the identical asset in each slot (only the
query-matching slot varies), so the same headline accumulates data across twenty ad groups instead of one.
Assets are then judged on business outcomes per impression, not click-through rate, not the platform's
strength rating, and the loop repeats indefinitely.

## When to run
- An RSA built to the testing-ready spec (7–8 headlines, 2–3 descriptions) has served ~2 weeks and is stable.
- Ad relevance and expected CTR are average or better (a broken foundation makes test results meaningless).
- On the recurring cadence: every 2–8 weeks per cluster, indefinitely, mature accounts keep cycling.

## When NOT to run
- No deployed RSA yet, or a bloated one needing a rebuild, `write-compelling-rsas`.
- Ad relevance rated below average, `improve-ad-relevance` first; testing on top of a relevance problem
  wastes cycles.
- Expected CTR rated below average, `improve-expected-ctr` first.
- You need cross-format coordination (PMax assets, Demand Gen, video in the same cadence), `run-a-creative-testing-cycle` wraps this skill for the Search slice.

## Prerequisites
- A defined cluster: the set of ad groups similar enough in intent to share identical creative outside the
  query-echo slot. One RSA per ad group within it.
- A chosen primary metric: conversions per impression for lead gen, revenue per impression for ecommerce,
  profit per impression where margins are tracked.
- A learning log (any persistent doc/sheet) recording hypotheses, results, and verdicts across cycles.

## Procedure
1. **Define or confirm the cluster.** List candidate ad groups and group by shared intent (generic product
   terms, category terms, competitor terms are typical splits). Start with the highest-volume cluster.
2. **Enforce the template.** Fix a slot map: each headline slot is permanently assigned one persuasion
   angle (query echo, core value, differentiation, credibility, risk removal, action prompt, plus a doubled
   slot for the lead angle). Within the cluster every ad group carries the identical asset per slot; only the
   query-echo slot adapts. If ad groups deviate, bring them onto the template with `gads_update_ad`, consult `gads_policy_guardrail` before the session's first write; every edit previews first (validate_only
   default) and applies only after explicit user approval.
3. **Let data accumulate** for the cycle length (2–8 weeks by volume; see Decision rules).
4. **Extract asset-level data.** Call `gads_get_rsa_asset_performance` for each ad group in the cluster, and
   `gads_run_gaql_query` for asset-level impressions, clicks, conversions, conversion value, and the ads'
   total impressions over the window.
5. **Compute the two axes per asset.** (a) The primary efficiency metric, conversions (or revenue/profit)
   divided by the asset's impressions. (b) Serve share, the asset's impressions divided by its ad's total
   impressions, i.e. how often the system chooses to show it. Tag each asset by its persuasion angle (tag
   the doubled slot by its actual angle, not its position) and aggregate both metrics per angle across the
   whole cluster.
6. **Classify each asset into the 2×2.** Efficiency above vs below the cluster average, crossed with serve
   share high (>~25%) vs low (<~15%); assets between the serve-share bounds stay unclassified this cycle.
   The four cells: proven winners (high/high), underexposed winners (efficient but rarely shown),
   high-exposure drags (shown constantly, converting poorly), and duds (poor and rarely shown).
7. **Act by cell, in priority order** (all edits via `gads_update_ad`, preview → user approval → commit):
   1. **High-exposure drags first**, they are actively taxing the whole cluster. Remove or replace them
      immediately.
   2. **Underexposed winners**, force more exposure: pin to a position where justified, or echo the
      theme in an additional description.
   3. **Proven winners**, leave untouched; record them in the learning log.
   4. **Duds**, swap out at convenience, never urgently.
8. **Generate the next round of variants.** For each freed slot, write a hypothesis first, "changing X
   should improve the primary metric because Z", sourced from review mining, sales objections, competitor
   gaps, or the previous cycle's log. Then call `gads_generate_rsa_copy` for candidates. A human approves
   every generated asset before deployment, no AI copy ships unreviewed.
9. **Deploy across the cluster** with `gads_update_ad` (preview → approve → commit), keeping the
   template invariant: same new asset in the same slot in every ad group. Cap changes at 2–3 new assets
   per slot per cycle.
10. **Sanity check** with `gads_get_ad_strength`, completeness feedback only; the rating is not a KPI.
11. **Log and schedule.** Record hypothesis, data table, verdict (confirmed / rejected / inconclusive),
    learnings, and the next hypothesis. Set the next cycle date.

## Decision rules
- **Primary metric, never CTR.** High-CTR assets can attract clicks that never convert; efficiency per
  impression measures the business outcome. CTR gains are a byproduct, not the target.
- **Combination math caps asset count.** n headlines yield n·(n−1)·(n−2) ordered three-headline
  combinations (8 → 336; 15 → 2,730). Hold RSAs at 7–8 headlines or per-combination data starves.
- **Serve-share thresholds:** above ~25% counts as high exposure, below ~15% as low; tune with asset
  count and volume.
- **Minimum evidence:** aggregate ~1,000+ impressions per asset across the cluster before classifying;
  under that, extend the cycle or widen the cluster.
- **Cycle length by volume:** 2–4 weeks for high-volume clusters, up to 8 for modest ones; chronically
  data-poor clusters should merge with adjacent ones or stretch to 8–12 weeks.
- **Test breadth follows maturity:** first discover which persuasion angles win (angle vs angle), then
  variations within winning angles, and only in mature clusters test phrasing tweaks of proven assets.
  Testing word choices before concepts is premature optimization.
- **Traffic warmth steers hypothesis priority:** colder clusters test pain-framing variants first; hotter
  clusters test credibility and risk-removal variants first.
- **Isolation:** keep bids, budgets, audiences, and landing pages stable during a cycle, or attribute nothing.

## Common failure modes
- **Template drift.** Ad groups accumulate one-off edits, slots stop matching, and cross-group aggregation
  becomes invalid. Re-audit the template every cycle before analyzing.
- **Judging assets in a single ad group.** The sample never reaches significance; aggregation across the
  cluster is the entire point.
- **Chasing the strength rating.** Padding headline count to please the meter reintroduces the data-poverty
  problem the 7–8 cap solves.
- **Ignoring high-exposure drags.** The most-served weak asset does more damage than any dud; it is
  always the first cut.
- **Random testing.** Variants without a written hypothesis produce results nobody can interpret; the
  hypothesis line in the log is mandatory.
- **Learning amnesia.** Failed tests get re-run a quarter later because nobody logged them. The log is part
  of the deliverable each cycle.
- **Confounded cycles.** A bidding migration or landing-page redesign landed mid-cycle and every asset
  delta became unattributable. Check the account change history before analyzing; if a confound landed,
  extend the window rather than trusting the numbers.

## Related skills
- `write-compelling-rsas`, builds the testing-ready RSA this loop consumes.
- `improve-ad-relevance`, `improve-expected-ctr`, prerequisite repairs when quality components are
  below average.
- `run-a-creative-testing-cycle`, the cross-format coordinator that schedules this loop alongside PMax,
  Demand Gen, and video testing.
- `review-and-optimize-ad-extensions`, parallel maintenance for the extension layer.
