AD FIELD NOTESAI advertising · the work behind the output

Research library / September 2026 notebook

Measure more than output

Test design, causal claims, business outcomes and the limits of platform case studies.

What to measure

StageMetricsWhat it answers
AttentionViewability, thumb-stop, first-three-second hold, completionDid the asset earn attention?
MessageRecall, recognition, comprehension, brand linkageDid people understand and link the idea to the brand?
InteractionSaves, shares, qualified comments, product exploration, assistant task completionDid the audience do something meaningful?
CommerceConversion, margin, CAC, ROAS, repeat purchaseDid the campaign create economic value?
IncrementalityHoldout, geo test, conversion lift, media mix, matched marketDid the campaign cause change beyond what would have happened?
TrustDisclosure comprehension, sentiment, complaints, opt-outs, correction rateDid synthetic or personalized treatment damage credibility?
Creative systemTime to first viable route, cost per approved asset, rework rate, rights exceptionsDid AI improve the workflow rather than just output volume?
Long-term brandDistinctive asset recognition, brand consideration, price resilience, mental availabilityDid the campaign strengthen memory, not only short-term clicks?

A minimum test design

  1. Define the business outcome before generating assets.
  2. Preserve a human-created or existing control.
  3. Test one meaningful variable at a time where possible.
  4. Use a holdout, geo split, randomized audience split, or conversion-lift study.
  5. Report spend, dates, audience, sample size, confidence interval, and exclusions.
  6. Inspect subgroup performance for bias and uneven delivery.
  7. Compare short-term efficiency with trust, fatigue, and repeat behavior.

Do not make “AI-generated” the independent variable if the AI version also has a different offer, edit length, audience, media placement, or budget. That measures a bundle of changes, not the creative method.

How to interpret evidence

  • Vendor case study: useful for learning what the vendor says worked; weak as general proof.
  • Agency case study: richer on creative process; performance claims remain self-reported.
  • Platform experiment: may have large scale but can reflect a product-specific model, audience, and optimization objective.
  • Independent experiment: stronger for causality, but usually studies a bounded task rather than a whole campaign.
  • Brand lift study: useful for memory and association; still needs method and sample transparency.
  • ROAS: useful for an operating decision but can over-credit the platform and ignore margin, incrementality, and long-term effects.

The current evidence is strongest for workflow acceleration and bounded assistance, not for the universal claim that synthetic creative is more effective than human creative.

Advertising-specific evidence from 2025–2026

Four newer studies help put the case-study claims in proportion:

StudyDesignFindingWhat it does not prove
Youth vaping-awareness ads, JAMA Network Open (2025)Randomized online study; 614 Australians aged 16–25; 25 youth-codesigned AI ads vs. 25 existing agency/health adsAI ads were noninferior and slightly better on four of five perceived-message measures; labeling the ads as AI-made had no significant association with those outcomesThat AI ads increase sales, or that a public-health sample generalizes to commercial categories
Synthetic personalized event ads, Sport Business & Management (2025)Experiment with 175 women comparing human, synthetic, and synthetic-with-disclaimer event promotionsSynthetic ads scored lower on perceived quality, realism, attitudes, and interest; a disclaimer brought conative responses closer to the human conditionThat disclosure always harms performance, or that one event category predicts all advertising
Meta AdLlama, arXiv preprint (2025)Platform A/B test across roughly 35,000 advertisers and 640,000 generated variations over 10 weeksA reinforcement-learning ad-text system reported a 6.7% advertiser-level CTR improvement over a supervised imitation baseline and 18.5% more variationsAI vs. human creative effectiveness; the comparison is between two automated text systems and is not yet peer-reviewed
Coca-Cola synthetic-ad backlash, Journal of Retailing and Consumer Services (2026)Qualitative analysis of 7,822 YouTube/Reddit comments on the AI remakeNegative reactions clustered around visual execution, authenticity, creativity, labor displacement, corporate motives, and human connectionPopulation-level sentiment, sales impact, or a general condemnation of all synthetic advertising

Sources: JAMA Network Open, Sport Business & Management, Meta AdLlama preprint, and Coca-Cola backlash analysis.

The practical conclusion is modest but useful: AI can improve perceived effectiveness when it is applied to a bounded, well-designed communication task, and it can improve production throughput. Synthetic realism is not automatically persuasive; in some settings it lowers trust or quality perceptions. A responsible test therefore measures efficiency and audience response separately and includes disclosure comprehension and sentiment as first-class outcomes.

Follow the evidence

This chapter comes from the September research notebook. Linked sources and qualifications remain attached to the claims.

Open all sources and footnotes →

The weekly field note

One useful note for the next brief.

Campaigns worth studying, the work behind them, and questions to bring into your next creative review. Delivered weekly.

How we handle your email