What the research says about creative effectiveness
A global sample of marketers was asked which television advertisements were more sales effective. They averaged 51% accuracy. Against validated sales outcomes, that is close to a coin toss.
The finding predates the current argument about AI and creative work, and it complicates both sides of it.
The full citations are in the Creative Effectiveness & Why Intuition Fails research index. These are peer-reviewed papers testing execution against in-market sales rather than against recall or liking.
Expert intuition is not calibrated
Hartnett, Kennedy, Sharp and Greenacre (2016) ran the prediction study in the Journal of Marketing Behavior. Marketers assessed advertisements whose commercial outcomes were already known, and their intuitive predictions landed at 51%.
Multivariate analysis found small improvements for people with category experience and for those in marketing or insights roles. Small, not decisive.
The result does not say creative judgment is worthless. It says confidence in creative judgment is poorly calibrated against sales, which is a different claim and a more actionable one. The person most sure in the review is not reliably the person most correct.
This matters before AI enters the conversation. A great deal of creative process is built on the assumption that a senior reviewer can identify what will work.
Creativity is two dimensions, and both are specifiable
Smith, MacKenzie, Yang, Buchholz and Darley (2007) modeled creativity in Marketing Science across a sequence of studies, and found perceived creativity is produced by the interaction of divergence and relevance.
Divergence is how far the work departs from what the category leads people to expect. Relevance is whether that departure connects to the audience and the product. Neither dimension alone produces the effect.
That is a useful decomposition, because both halves can be requested in a brief and reviewed separately. "Make it more creative" is not an instruction anyone can act on. "This is divergent but the divergence has nothing to do with the product" is.
| Divergence | Relevance | Result |
|---|---|---|
| High | High | The interaction the research associates with creative effect |
| High | Low | Noticeable, unconnected to the brand or category |
| Low | High | Accurate, unremarkable, easy to ignore |
| Low | Low | Wallpaper |
Execution features can be tested against sales
Hartnett and colleagues (2016) also coded 158 creative variables across 312 television advertisements carrying commercially validated short-term sales outcomes, extending earlier work on creative devices.
The method is the contribution. Most creative testing measures recall or liking because those are cheap to collect, and then treats them as stand-ins for commercial effect. Testing coded execution features directly against sales removes a substitution that weakens most creative evidence without announcing itself.
Williams, Hartnett and Trinh (2023) continued the line in the International Journal of Market Research, applying modern analysis to isolating creative drivers. The through-line across this work is that method moves the field, and opinion does not.
What this changes when AI produces the work
The intuition finding cuts against the standard reassurance about AI creative. The safeguard usually offered is that a human reviews the output. These results say the reviewer is operating near chance on the question of what will sell.
It cuts equally against the opposite claim. An AI model has no privileged access to commercial effect either. What it learned from is largely the same body of admired work whose sales performance experts cannot predict, so a model producing confident creative judgments is reproducing an uncalibrated signal at speed.
Neither party in the review is calibrated. That is the actual situation.
The reasonable response is not to distrust the tool or the person. It is to notice that volume raises the value of measurement, and to hold creative arguments to the standard the research uses. What is the divergence here, is it relevant, and what outcome measure supports the claim that this works?
Copper Sun is built for the part of this that is tractable: modules run a stated process so the reasoning behind a concept is visible and arguable, rather than arriving as a finished opinion. A brief that specifies divergence and relevance separately produces work you can actually review.
The rest is measurement, and no tool removes that obligation.
Frequently Asked Questions
Can experienced marketers tell which ads will sell?
Not reliably. Hartnett, Kennedy, Sharp and Greenacre (2016) found predictions averaging 51% accuracy against validated sales outcomes. Category experience and insights roles helped slightly. The finding argues for testing rather than for deferring to the most senior view in the room, since confidence and accuracy come apart.
What makes advertising creative in measurable terms?
Smith and colleagues (2007) found perceived creativity comes from the interaction of divergence, how far the work departs from expectation, and relevance to the audience and category. Neither alone produces the effect. Both can be specified in a brief and reviewed as separate questions.
Does this research say creative awards do not predict effectiveness?
These papers do not test awards directly. They establish that in-market sales are the measure worth using and that recall and liking are weaker substitutes. Any claim about creative performance should name the outcome it was validated against.
Should AI produce creative concepts at all, given this?
The research does not settle that, and it undercuts the confident answer in both directions. A human reviewer is close to chance at predicting sales effect, and a model trained largely on admired work has no better signal. What follows is a process argument: specify divergence and relevance, keep the reasoning visible, and measure. The related discipline on consistency across volume is in what the research says about how brands actually grow.