Start a conversation Contact
← Research

Give AI the creative queue, not the campaign budget

Generative AI makes advertising variation cheap and can improve click-through. Its commercial value appears only when a controlled test follows each creative through to conversion and the outcome the business keeps.

A performance-marketing director can now turn one campaign brief into dozens of finished images before the creative team would previously have reviewed three. The proposed business case counts production hours removed and clicks gained.

The opportunity is larger than cheaper production, and less certain. More variants let the team search a wider set of messages, product contexts and audience treatments. That search creates value only when the organisation can identify which creative caused more of the outcome it actually buys advertising to produce.

Give AI the creative queue, not the campaign budget. Let it create more eligible variants, then make brand review and a controlled commercial test decide which ones receive further spend. A click is useful evidence that an ad earned attention. It is not evidence that the attention was worth buying.

The production gain is real

A 2026 Journal of Marketing study shows what a well-bounded visual generation workflow can do. The researchers trained an open-source image model using consumer ratings of attention, interest, desire and action, then adapted it to the visual identity of a new car. They compared generated Polestar 3 adverts with conventional Polestar material through surveys and live platform tests. Heitmann and colleagues, 2026

Across two Meta A/B tests, 28,775 impressions produced 496 clicks. The generated adverts recorded a combined click-through rate of 2.04 per cent, compared with 1.37 per cent for the conventional adverts included in the tests. The generated work had been selected through a process that used audience feedback, brand images and stated communication objectives. This was not a test of typing a short prompt into an unconfigured image service. Heitmann and colleagues, 2026

The result supports a production decision. Generative AI can make additional, testable creative that is competitive on an observable attention measure. It does not establish that the additional clicks became purchases, qualified leads or profitable customers. The authors themselves describe click-through as a poor performance measure on its own because some clicks never convert and some advertising creates effects without a click.

Another 2026 randomised field experiment sharpens that limit. More than 150 AI-generated and human-created video adverts were compared across the funnel. The generated adverts earned higher click-through, held attention longer and drew fewer negative reactions, but were less likely to convert. The public conference record gives no effect sizes or campaign-level detail, so it is a directional caution rather than a forecast. Li and Li, 2026

Where an AI creative programme can mistake activity for commercial value Fig. 01
What becomes easier to produce What still has to be established
More finished images and videos from one brief More eligible concepts that remain distinct after brand and factual review
More impressions and clicks Incremental purchases, qualified leads or another declared business outcome
A lower effort per creative A shorter learning cycle or more useful tests with the same team
A platform-selected winning advert A controlled comparison whose delivery rules did not choose the winner first

The first column can be reported on the day the system arrives. The second is the investment case. Keeping them separate prevents an efficient content factory from being mistaken for an effective growth system.

Change the unit of work from asset to experiment

Before AI, the creative team receives a brief, develops a small set of concepts, produces the approved assets and hands them to the media team. Limited production capacity forces human judgement to decide which ideas deserve a test.

After AI, the scarce step moves. The system can propose many executions, but the team must decide whether they represent different commercial hypotheses or the same composition with superficial changes. It must reject factual errors, unlicensed material and work that weakens the brand. The media team then needs a test that gives eligible variants a credible chance to succeed.

That changes the deliverable. It is no longer a folder of generated assets. It is a queue of experiments. Each entry carries the audience, proposition, creative hypothesis, source material, model version, approval decision and the commercial outcome against which it will be judged.

A separate 2026 working paper examined more than 16 billion display-ad impressions across nearly fifty product categories. In a narrower set of 1,186 within-campaign comparisons, AI-generated imagery performed no differently from human-made imagery on average. Generated images did better when people did not perceive them as looking AI-made, and lost that advantage as perceived artificiality rose. The study is quasi-experimental and its conversion analysis was underpowered, but it shows why source labels alone do not decide an asset’s performance. Visual treatment, product category and audience response do. Exner and colleagues, 2026

Use those variables to create hypotheses rather than universal design rules. A colour treatment, face, setting or degree of polish can work differently by brand and audience. The organisation’s own controlled exposure is the evidence that can settle its decision.

Build a funnel the test cannot skip

Choose the primary commercial outcome before generation begins. For a retailer it may be completed purchases read beside returns. For an enterprise campaign it may be qualified opportunities that reach a defined sales stage. For a subscription service it may be activated customers who remain after an agreed observation period. Click-through and viewing time remain diagnostic measures. They explain where the funnel changed, but they do not replace its end.

Randomisation also needs to survive the advertising platform. If an optimisation system quickly directs most impressions towards the creative it predicts will earn clicks, exposure is partly a result of the same proxy the test is supposed to examine. Use the platform’s controlled experiment where available, fix the allocation long enough to support comparison and record actual exposure by variant and audience.

An AI creative workflow that can support a growth decision Fig. 02
  1. 01 Brief Name the audience, proposition, constraint and business outcome before generation.
  2. 02 Generate Produce genuinely different hypotheses from approved product and brand material.
  3. 03 Qualify Reject factual, rights, brand and duplication failures before buying exposure.
  4. 04 Test Randomise eligible variants under recorded delivery and attribution rules.
  5. 05 Promote Increase exposure only when the declared commercial outcome supports it.

Measure the cost of obtaining the result as well as the result. Record staff review time, rejection rate, revision rate, number of materially distinct hypotheses and the interval from brief to a decision. If the system generates ten times as many assets and the team rejects nearly all of them, production has not become cheap. It has moved into review.

Read conversion with post-conversion quality. Orders that return, leads that never qualify and subscriptions that cancel quickly can make an advert look successful by moving friction beyond the measurement window. The definition of a useful conversion belongs in the brief.

What the Inference Institute can help decide

The useful engagement starts with the campaign workflow and measurement design, not an image-model demonstration. It defines what the organisation wants more of, which brand and source constraints every variant must respect, how hypotheses remain distinct, where human approval sits and how the platform will expose the assets without invalidating the comparison.

The result is a testable operating specification for a marketing director and their data team. It names the eligible use case, the evidence recorded for each asset, the metric hierarchy, the attribution window and the result that expands, narrows or stops the programme. A model comparison can then serve that design instead of becoming the design.

What this does not tell you

The live tests in the Journal of Marketing study concerned one car brand, a small set of banner adverts and click-through rather than purchase. The platform limited each experiment to five variants, delivery was not under complete researcher control and some people could have encountered both concurrent tests. Survey studies elsewhere in the paper cannot substitute for behaviour.

The video-ad finding is a conference paper represented publicly by its abstract. Without the full methods and estimates, it supports caution about the funnel but not a numerical expectation. The display-ad paper is a working paper using observational platform data arranged into quasi-experiments. Its large raw population does not remove the limits of that design.

None of the evidence says generated creative will improve another campaign. It says there is enough promise, variation and metric disagreement to justify a controlled test. The performance-marketing director’s next decision is which campaign has a commercial outcome that can be observed beyond the click. Give AI room to enlarge the creative search there, and let that outcome decide which work survives.

Filed under · Method · Digital advertising · Generative AI · Experimentation Inference Institute · 11 Sept 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.