AI ad creative in 2026: what's actually verifiable
Most of the numbers in this category can't be traced to anyone who measured them. Here's what survived the check — and what the tools really change.

We build with these tools daily, so this isn't a sceptic's takedown. But we went looking for the primary sources behind the performance claims that dominate this topic, and the results were bad enough to be worth publishing on their own.
What we could verify
Google publishes real figures for asset generation in Performance Max. Advertisers using asset generation when creating a campaign are 63% more likely to publish with Good or Excellent Ad Strength; those improving Ad Strength to Excellent see 6% more conversions on average; and campaigns including at least one video see 12% more total conversions.[1]
Two caveats that almost nobody attaches to them. First, that announcement is dated February 2024 — if you've seen the 63% presented as a 2026 statistic, it's two years older than advertised. Second, and more important: Ad Strength is Google's own measure of how completely you've filled in your assets. Generating more assets raising your asset-completeness score is close to tautological. Only the 6% and 12% figures speak to conversions at all.
Google's own product documentation is refreshingly plain about the limits, noting that assets created by generative AI aren't guaranteed to pass ads policy review.[2] That's a useful sentence to keep in mind when someone offers you a fully automated pipeline.
What we could not verify
The Meta-side statistics that circulate in this category — a 12% higher click-through rate from Advantage+ creative, 22% higher ROAS, a 4% lower cost per result from standard enhancements — we could not confirm any of them on a Meta-owned page. The "4% lower cost per result" line in particular appears word-for-word across many vendor blogs, which is the signature of a single unsourced claim being copied rather than several parties measuring the same thing.
It's entirely possible Meta has published these somewhere we couldn't reach. But we're not going to repeat a number we can't stand behind, and neither should the agency pitching you.
The test, in one line
For any AI-performance statistic: who measured it, over how many accounts, at what spend, and against what control? If the answer is a vendor blog citing another vendor blog, you don't have evidence — you have marketing about marketing.
What AI genuinely changes
Strip out the unverifiable numbers and something more useful is left, which is a change in the shape of the constraint rather than in the quality ceiling.
It moves the bottleneck from production to judgement
Producing forty concepts used to be the hard part. It isn't any more. Deciding which four deserve spend — and being right often enough to matter — is now the entire game. Anyone selling you volume without telling you how the selection happens is selling you the easy half.
It makes the replacement rate achievable for the first time
The production arithmetic that broke traditional retainers — needing five new creatives a week, permanently, just to hold position — is genuinely tractable now. That's the real unlock, and it's a throughput story, not a quality story.
It does not make the average asset better
This is where our own data is uncomfortable. We scored fifteen of our own launch concepts with the same model and rubric we use on client work. The spread ran 45 to 70, and when we checked which sub-score actually moved the overall, clarity correlated at r=+0.97 and offer at r=+0.91, while hook came in at +0.37 and distinctiveness at −0.11.
The generated visuals were consistently competent — hook scores clustered in an eight-point band across all fifteen. What varied enormously was whether the ad said something specific: clarity ranged across 31 points, offer across 35. Our worst concept scored 45 with a perfectly good hook and a clarity score of 40 — striking image, copy that meant nothing. It's still published with its score on our lab page.
That's the honest read on AI creative in 2026: it reliably clears the visual bar and it does nothing at all for the thing that actually separates a winner from a loser. If the strategy is thin, the tools will render thin strategy faster and in higher resolution.
Where the humans have to stay
Three places, in our build, and we'd argue in anyone's:
- The claim. What the ad asserts, and whether it's true. No model should be deciding what your brand promises a customer.
- The kill. Something has to say no, and it can't be the same thing that made the asset. We score everything 0–100 before a human sees it, and a senior human who didn't make the asset signs off what ships.
- The account. Nothing touches live spend without approval, and every change gets a log entry that can be reversed. "The AI changed something" should never be a sentence anyone has to say about their own ad account.
The category is going to keep producing impressive-sounding percentages. Ask where they came from. The ones we could source are above, correctly dated; the ones we couldn't aren't in this post.
↳ Scored creative and a written verdict on your account, with the workings shown — including where the evidence wasn't strong enough to be sure.
Sources
- Google, Gemini models are coming to Performance Max — 22 February 2024
- Google Ads Help, About asset generation in Performance Max — Product documentation — undated
↳ Every figure above is either sourced here or is our own measured data, stated as ours
Keep reading
Want this run againstyour account?
The free Crank Audit is the same analysis pointed at your brand — senior-grade, workings shown, in your inbox in 24 hours.