---
name: run-a-creative-testing-cycle
description: >-
  The monthly cross-format creative testing coordinator: writes explicit hypotheses per format, snapshots
  baselines, deploys variants for Search RSAs, PMax asset groups, Demand Gen, display, and video inside one
  synchronized window, freezes other variables during the test, then analyzes with per-format duration and
  significance gates before promoting winners and replacing losers. Use it when multiple ad formats are live and
  their testing needs one shared cadence and learning log. For the Search-only deep methodology use
  rsa-testing-with-the-iteration-loop; for composing the initial RSA use write-compelling-rsas; for PMax
  campaign-level work (budgets, structure) use the PMax optimization cycle skills rather than this creative
  layer.
---
# Cross-Format Creative Testing Cycle

## Purpose
When Search, PMax, Demand Gen, and video each test on their own schedule, results contaminate one another
and learnings never transfer. This skill runs all formats through one synchronized monthly cycle:
hypothesis → variant → controlled window → analysis → deployment → documented learning, so a winning
angle discovered in one format becomes next cycle's hypothesis in the others.

## When to run
- Monthly, once at least one campaign per active format has been live 4+ weeks.
- The previous cycle's changes have had 2+ weeks to accumulate data.
- Conversion tracking is verified and stable (test results on broken tracking are noise).

## When NOT to run
- Only Search is live and the work is RSA asset testing, `rsa-testing-with-the-iteration-loop` alone is
  simpler and sufficient.
- Initial creative has not been deployed for a format, you cannot test what does not exist; build first
  (`write-compelling-rsas` for Search, the PMax/Demand Gen launch skills for those formats).
- The problem is campaign structure, budgets, or bidding, that belongs to the per-format optimization
  cycles, not the creative layer.

## Prerequisites
- Asset-level data access for every active format.
- The learning log and hypothesis backlog from prior cycles (or empty templates on the first run).
- Human creative capacity lined up: new images and videos must be produced by people before the deploy
  window.
- Agreement with the user that bids, budgets (beyond ±10%), audiences, and landing pages stay frozen for
  the test window.

## Procedure
1. **Review the backlog and write hypotheses.** Reread the prior learning log; pull deferred items forward.
   For each active format write 1–3 hypotheses in the form "if we change X in this format, metric Y should
   move in direction Z because [reason]", sourced from customer reviews, sales objections, competitor
   gaps, seasonality, or a win in another format. Never more hypotheses than the format's volume can
   answer in one cycle.
2. **Snapshot baselines (reads only).** Over the last 14–30 days record per-asset/per-creative impressions,
   clicks, conversions, and the primary efficiency metric:
   - Search: `gads_get_rsa_asset_performance` per cluster.
   - PMax: `gads_get_asset_group`, `gads_list_assets`, and `gads_get_pmax_channel_performance` for
     channel mix context.
   - Demand Gen / display / video: `gads_run_gaql_query` on the relevant ad and asset stats.
3. **Produce the variants.**
   - Text: `gads_generate_rsa_copy` drafts headline/description candidates against each hypothesis; a
     human approves every AI-drafted asset before it ships.
   - **HUMAN STEP (outside VigilDog): image and video production**, new visuals are made by people.
     Practical guardrails to pass to the producer: feed-context video under ~30 seconds; images in both
     1200×628 and 1200×1200; isolate one element per video variant (change the opening hook OR the
     length OR the style, keeping the rest constant).
   - Upload finished media with `gads_upload_image_asset` and `gads_upload_youtube_video_asset`.
4. **Deploy inside one window.** Consult `gads_policy_guardrail` before the session's first write, then ship
   all variants within a ~2-day window so every test shares the same market conditions. Every write
   previews first (validate_only default) and applies only after explicit user approval:
   - Search RSAs: `gads_update_ad`, respecting the cluster template from
     `rsa-testing-with-the-iteration-loop`, one slot under test per cluster.
   - PMax: `gads_add_asset_group_asset` for new assets, `gads_remove_asset_group_asset` for the 1–2
     weakest per asset type, never more than 2 swaps per asset type per cycle, or measurement clarity
     dies.
   - Tag every variant with a cycle ID (format–date–number) and record launch dates.
5. **Hold the freeze.** For the duration: no bid-strategy changes, budget moves capped at ±10%, no
   audience edits, no landing-page changes, no creative edits outside this cycle. Any violation gets logged
   as contamination.
6. **Monitor weekly.** Confirm variants are approved and serving (`gads_run_gaql_query`, `gads_list_assets`);
   check for unplanned changes; log external events (promotions, seasonality, competitor moves). Kill rule:
   a variant running 50%+ worse than its control after a full week of data gets paused and logged as a
   rejected hypothesis. Do not end tests early because something "looks like" a winner.
7. **Analyze at the gate, not before.** When a format hits its duration and data minimums (Decision rules),
   pull closing data with the same tools as step 2. Judge on conversions- or revenue-per-impression. For
   Demand Gen and video with view-through conversions, use blended conversions = click-through
   conversions + view-through × 0.3–0.5 (0.3 conservative, 0.5 only when validation supports it).
   Awareness-goal campaigns are judged on CPM, view rate, and frequency instead, efficiency-per-
   impression does not apply there.
8. **Check significance before declaring anything.** Rough conversion requirements per variant: ~100–200
   for effects of 20%+, ~300–500 for 10–20%, ~1,000+ for 5–10%; differences under 5% are inconclusive at
   most volumes. Results short of roughly 80% confidence are "directional": extend 2 weeks, or deploy
   tentatively with continued monitoring, per the low-volume ladder below.
9. **Deploy winners, replace losers** (previews + user approval on all writes): promote winning RSA assets
   and cut the high-serving weak ones via `gads_update_ad`; keep strong PMax assets and replace
   underperformers via the asset-group tools; scale winning Demand Gen/video creative and pause losers.
   Replacements must test a *different* angle, not a rewording of the loser.
10. **Document and look across formats.** Complete a log entry per hypothesis (ID, dates, hypothesis, data
    table, verdict, learnings, next hypothesis). Then scan for cross-format patterns, an angle winning in
    two formats becomes a priority hypothesis everywhere; format-specific style preferences get noted so
    creative is tailored, not blindly copied.
11. **Schedule the next cycle.** Standard: 4 weeks after deployment. Accelerate to ~3 weeks after major
    wins; extend 2 weeks for inconclusive tests. Refresh the backlog to 3+ ready hypotheses, and verify
    Demand Gen creative age, refresh anything older than ~60 days to preempt fatigue.

## Decision rules
- **Duration and data minimums before judging:** for Search RSAs and PMax, wait 2–4 weeks and
  until each asset has accumulated at least ~1,000 aggregated impressions; Demand Gen needs 3–4
  weeks (its learning phase included) plus roughly 30 conversions in the ad group; display asset
  tests follow the same 3–4 week / ~1,000-impression bar; video is slowest, plan on 4–8 weeks
  with several hundred views per creative and a double-digit conversion count before calling it.
- **Hypothesis load caps:** Search 2–3 per cluster; PMax 1–2 per asset group; Demand Gen 2–3 per ad
  group; display and video 1–2 (production cost and low conversion volume make more unmeasurable).
- **Early-kill rule:** ≥50% worse than control after one full week → pause and log; everything milder waits
  for the gate.
- **Strength ratings and completeness meters are not performance signals** in any format; ignore them for
  win/lose calls.
- **Low-volume ladder:** at 4+ weeks nearing but not reaching confidence → extend 2 weeks; at 6+ weeks
  still short → classify directional, deploy tentatively, keep monitoring; at 8+ weeks with thin data → stop
  asset-level testing in that campaign, consolidate variants, and test bigger contrasts at ad or format level;
  chronically short → pool data across ad groups/campaigns or test format-level changes instead of subtle
  variants.
- **Demand Gen coverage rule:** keep at least one video and one image live per ad group (video feeds the
  video surfaces, image feeds the feed/mail surfaces); 3–5 variants per format gives the system room to
  optimize.
- **Test order by maturity:** discover the winning *format* first (video vs image vs carousel, style A vs B),
  then refine assets inside the winning format. Broad before narrow.

## Common failure modes
- **Unhypothesized testing.** Variants shipped "to see what happens" return data nobody can act on; the
  if/then/because line is the price of entry.
- **Mid-test meddling.** A bid change or landing-page tweak during the window makes results
  unattributable; the freeze list exists for this.
- **Over-swapping PMax assets.** Replacing half an asset group at once destroys the comparison; the
  2-per-type cap is hard.
- **Declaring winners off tiny deltas.** A 7% gap at 60 conversions is noise; hold the significance ladder.
- **Learnings never written down.** Winners deploy, the log stays empty, and next quarter repeats a failed
  test. The cycle is not done until the log is.
- **Evergreen assumptions in feed formats.** Demand Gen creative fatigues; the ~60-day refresh check is
  part of every cycle close.
- **Testing five video variants on a channel producing eight conversions a month.** Match test ambition to
  conversion volume or move the test up a level.

## Related skills
- `rsa-testing-with-the-iteration-loop`, the Search-format engine this cycle schedules and consumes.
- `write-compelling-rsas`, builds the Search creative foundation before any testing.
- `review-and-optimize-ad-extensions`, parallel monthly maintenance on the extension layer.
- PMax / Demand Gen / video optimization cycle skills, campaign-level counterparts; route structural
  findings (asset-group splits, budget shifts) there rather than acting in the creative cycle.
