What to measure
| Stage | Metrics | What it answers |
|---|---|---|
| Attention | Viewability, thumb-stop, first-three-second hold, completion | Did the asset earn attention? |
| Message | Recall, recognition, comprehension, brand linkage | Did people understand and link the idea to the brand? |
| Interaction | Saves, shares, qualified comments, product exploration, assistant task completion | Did the audience do something meaningful? |
| Commerce | Conversion, margin, CAC, ROAS, repeat purchase | Did the campaign create economic value? |
| Incrementality | Holdout, geo test, conversion lift, media mix, matched market | Did the campaign cause change beyond what would have happened? |
| Trust | Disclosure comprehension, sentiment, complaints, opt-outs, correction rate | Did synthetic or personalized treatment damage credibility? |
| Creative system | Time to first viable route, cost per approved asset, rework rate, rights exceptions | Did AI improve the workflow rather than just output volume? |
| Long-term brand | Distinctive asset recognition, brand consideration, price resilience, mental availability | Did the campaign strengthen memory, not only short-term clicks? |
A minimum test design
- Define the business outcome before generating assets.
- Preserve a human-created or existing control.
- Test one meaningful variable at a time where possible.
- Use a holdout, geo split, randomized audience split, or conversion-lift study.
- Report spend, dates, audience, sample size, confidence interval, and exclusions.
- Inspect subgroup performance for bias and uneven delivery.
- Compare short-term efficiency with trust, fatigue, and repeat behavior.
Do not make “AI-generated” the independent variable if the AI version also has a different offer, edit length, audience, media placement, or budget. That measures a bundle of changes, not the creative method.
How to interpret evidence
- Vendor case study: useful for learning what the vendor says worked; weak as general proof.
- Agency case study: richer on creative process; performance claims remain self-reported.
- Platform experiment: may have large scale but can reflect a product-specific model, audience, and optimization objective.
- Independent experiment: stronger for causality, but usually studies a bounded task rather than a whole campaign.
- Brand lift study: useful for memory and association; still needs method and sample transparency.
- ROAS: useful for an operating decision but can over-credit the platform and ignore margin, incrementality, and long-term effects.
The current evidence is strongest for workflow acceleration and bounded assistance, not for the universal claim that synthetic creative is more effective than human creative.
Advertising-specific evidence from 2025–2026
Four newer studies help put the case-study claims in proportion:
| Study | Design | Finding | What it does not prove |
|---|---|---|---|
| Youth vaping-awareness ads, JAMA Network Open (2025) | Randomized online study; 614 Australians aged 16–25; 25 youth-codesigned AI ads vs. 25 existing agency/health ads | AI ads were noninferior and slightly better on four of five perceived-message measures; labeling the ads as AI-made had no significant association with those outcomes | That AI ads increase sales, or that a public-health sample generalizes to commercial categories |
| Synthetic personalized event ads, Sport Business & Management (2025) | Experiment with 175 women comparing human, synthetic, and synthetic-with-disclaimer event promotions | Synthetic ads scored lower on perceived quality, realism, attitudes, and interest; a disclaimer brought conative responses closer to the human condition | That disclosure always harms performance, or that one event category predicts all advertising |
| Meta AdLlama, arXiv preprint (2025) | Platform A/B test across roughly 35,000 advertisers and 640,000 generated variations over 10 weeks | A reinforcement-learning ad-text system reported a 6.7% advertiser-level CTR improvement over a supervised imitation baseline and 18.5% more variations | AI vs. human creative effectiveness; the comparison is between two automated text systems and is not yet peer-reviewed |
| Coca-Cola synthetic-ad backlash, Journal of Retailing and Consumer Services (2026) | Qualitative analysis of 7,822 YouTube/Reddit comments on the AI remake | Negative reactions clustered around visual execution, authenticity, creativity, labor displacement, corporate motives, and human connection | Population-level sentiment, sales impact, or a general condemnation of all synthetic advertising |
Sources: JAMA Network Open, Sport Business & Management, Meta AdLlama preprint, and Coca-Cola backlash analysis.
The practical conclusion is modest but useful: AI can improve perceived effectiveness when it is applied to a bounded, well-designed communication task, and it can improve production throughput. Synthetic realism is not automatically persuasive; in some settings it lowers trust or quality perceptions. A responsible test therefore measures efficiency and audience response separately and includes disclosure comprehension and sentiment as first-class outcomes.
Follow the evidence
This chapter comes from the September research notebook. Linked sources and qualifications remain attached to the claims.
Open all sources and footnotes →