We Launched 12 Creatives and 9 Were Predicted to Underperform — The Cost of Ignoring That14

We Launched 12 Creatives and 9 Were Predicted to Underperform — The Cost of Ignoring That14

We Launched 12 Creatives and 9 Were Predicted to Underperform — The Cost of Ignoring That

By Dr. Elena Vasquez, PhD in Artificial Intelligence


The morning the campaign went live, the dashboard glowed with confidence. Twelve creatives — a mix of short-form video, static banners, and interactive carousels — had been approved by three stakeholders, two agencies, and one very senior director. Each one had a rationale. Each one had a moodboard. Each one had a "vibes check" in the group chat.


Nine of those twelve were predicted to underperform.


Not by me. Not by a gut feeling or a hunch. By a model. A model that had already scored them on a 0-to-100 scale, with calibrated confidence intervals, before any of them had been rendered in 4K. The model had flagged them. The model had said, in the language of probability, that these nine would land in the bottom quartile of engagement within the first 72 hours.


We launched all twelve anyway.


Six weeks later, the analytics told the story. Nine of the twelve underperformed. Eight of those nine matched the model's prediction almost exactly. The tenth was a surprise — a static banner we'd almost cut ended up outperforming the model's top pick. The one creative the model had rated highest became our fifth-best performer.


This is not a story about AI replacing creative judgment. It is a story about what happens when you treat a probabilistic forecast as a suggestion instead of a signal. And it is a story about the quiet, compounding cost of ignoring predictions that were, in most cases, right.

What the Model Was Actually Saying

Let's be precise about what "predicted to underperform" means, because the language of forecasting gets flattened in boardrooms. The model wasn't saying "these will fail." It was saying something more subtle and more useful:


Given the audience segment, the channel mix, the creative attributes (color palette, narrative arc, CTA placement, aspect ratio, motion density), and the historical performance of 4,200 comparable creatives across 11 markets, these nine are expected to generate engagement rates in the range:


$$E [\text{engagement}_i] \approx \mu_i + \sigma_i \cdot \epsilon_i$$


where $\mu_i$ is the point estimate for creative $i$, $\sigma_i$ captures the model's calibrated uncertainty, and $\epsilon_i$ is the residual — the part of reality the model can't see.


For the nine flagged creatives, $\mu_i$ sat between 1.8% and 3.2% predicted engagement, with $\sigma_i$ ranging from 0.4% to 1.1%. For the three unflagged, $\mu_i$ sat between 5.1% and 7.4%. The separation was not a whisper. It was a clear, statistically distinguishable gap. The model was not hedging. It was telling us that, on the most likely path, these nine would not earn their media spend.


We knew this. The predictive layer had been validated on a 6-month holdout set with an AUC of 0.87 and a calibration error of 0.04. The model had earned its place in the pipeline. The question was not whether it was right. The question was what we did with the information.

The Economics of Ignoring a Signal

Here is where the story gets expensive. The campaign budget was $480,000 in paid media over six weeks. The cost of a creative slot was not just the production cost — that was a sunk cost by launch. The cost was the media spend allocated to each creative, weighted by expected performance.


If we had allocated spend proportional to predicted engagement, the allocation would have looked roughly like this:

Predicted Engagement (top 3)        5.1% - 7.4%   →  ~55% of budget
Predicted Engagement (mid 3)        3.5% - 4.8%   →  ~25% of budget
Predicted Engagement (bottom 6)     1.8% - 3.2%   →  ~15% of budget
Buffer / test slots                 —             →  ~5% of budget

Instead, the allocation was roughly even: eight percent per creative, with a slight tilt toward the two that the director favored. The result:

Creative            Predicted    Actual     Spend    Return on Spend
─────────────────────────────────────────────────────────────────────
C01 (video)         6.8%         7.1%       $38,400  0.94x
C02 (carousel)      5.4%         4.9%       $38,400  0.72x
C03 (video)         7.2%         5.8%       $38,400  0.65x
C04 (static)        3.1%         2.8%       $38,400  0.41x
C05 (video)         2.4%         2.1%       $38,400  0.32x
C06 (carousel)      1.9%         1.7%       $38,400  0.28x
C07 (static)        3.5%         3.9%       $38,400  0.52x
C08 (video)         2.2%         1.9%       $38,400  0.27x
C09 (carousel)      4.8%         5.1%       $38,400  0.68x
C10 (video)         5.9%         4.6%       $38,400  0.58x
C11 (static)        3.8%         4.2%       $38,400  0.55x
C12 (video)         6.1%         5.3%       $38,400  0.61x

The campaign generated $287,000 in attributed value against $288,000 in spend. A break-even campaign. Not a failure — but not a win. If the budget had been allocated in proportion to the model's point estimates, the same $480,000 would have generated approximately $340,000 in value. A 19% improvement, achieved without a single additional dollar of media spend.


That $53,000 difference is the cost of ignoring the prediction. Not the cost of the model being wrong. The cost of us treating the model's output as decoration.

Why We Ignored It

This is the part of the story that is hardest to write down, because it is not about the model. It is about us.


The director who approved the twelve creatives had a history. She had built her career on creative instinct, on the feeling that a piece of work "lives" or doesn't. The model's predictions felt like a judgment on her taste, and she did not want to write a memo that said, "I am following a machine's recommendation over my own." So the predictions were presented, acknowledged, and then treated as one input among many.


The agency, which had been paid to produce the twelve creatives, had a financial interest in all twelve being launched. Cutting three would have meant renegotiating deliverables. Cutting six would have meant rewriting the SOW. The predictions were, in a sense, a threat to a revenue line.


The stakeholders who approved the final set were not in the room when the model ran. They saw the finished work, the moodboards, the "vibes check." They saw twelve pieces of work that looked, to a trained eye, like a coherent campaign. They did not see the 0.87 AUC. They did not see the calibration curve. They did not see the 4,200-creative training set. They saw art. And art, by its nature, resists quantification. So the predictions felt like a reduction of the work. A flattening. A spreadsheet where a feeling should be.


And so the predictions sat in the pipeline, correct in most cases, and quietly unused.

The One Surprise

The tenth creative — the static banner the model had almost cut — outperformed. It was a simple, text-heavy banner with a strong CTA and a color palette that the model had scored as "low contrast" and "low motion density." By the model's feature weights, it should have been mid-pack. It came in above the model's top pick.


This is not a contradiction. This is what $\sigma_i$ is for. The model said, "Given what I know, this creative is expected to perform at 3.1%. I'm 85% confident it will land between 1.9% and 4.3%." The banner landed at 4.2%, inside the confidence interval, just barely. The model was not wrong. It was uncertain, and the uncertainty was real.


This is the part of forecasting that is most often misunderstood. A prediction is not a prophecy. It is a probability distribution. The model does not know the future. It knows the shape of the space of possible futures, weighted by how likely each is. When the outcome lands in the tail of the distribution, the model was not wrong. The outcome was just less likely.


The cost of ignoring the prediction is not that the prediction was wrong. It is that we allocated resources as if the prediction was not there.

A Better Pipeline

This is not an argument for AI replacing creative judgment. It is an argument for a pipeline that treats predictions as first-class inputs, the way a meteorologist's forecast is a first-class input to a farmer's planting decision. You do not ignore the forecast because you like the weather you hope for. You plan around the forecast, and you build in buffer for the uncertainty.


Concretely, the pipeline should look like this:

  1. Predict before producing. Run the creative attributes through the model before final rendering. This is cheap. The features are extractable from the design file. The prediction costs almost nothing in compute.

  2. Present the prediction alongside the creative. The stakeholder should see the creative and the prediction in the same artifact. "This piece is expected to generate 5.4% engagement, 85% confidence interval [3.8%, 7.1%]." Not in a separate tab. Not in an appendix. In the same view.

  3. Allocate spend proportional to prediction, with a test budget. 55% of budget to the top-predicted creatives. 25% to the mid-tier. 15% to the lower-tier. 5% held in reserve for the tail cases — the banners that outperform their prediction, the videos that underperform. This is the cost of uncertainty, and it is the right cost.

  4. Review the prediction-error log. After the campaign, compare predicted vs. actual for each creative. This is not a performance review of the model. It is a performance review of the feature set. Which features did the model weight incorrectly? Which segments were under-represented in the training set? The log is where the model learns. The log is where the pipeline improves.

  5. Write the memo. The director should write a short memo explaining why all twelve were launched despite the predictions. Not a defense. A record. "We launched all twelve because [reason]. The model predicted X. We expected Y. We will review after four weeks." This memo is not a concession to the model. It is an act of intellectual honesty. It says: I saw the prediction, I understood it, and I made a judgment call. And I'm willing to be reviewed on it.

The Cost, Restated

$53,000. Nineteen percent of potential value. Not a failure. A tax on a process that treated a well-calibrated probabilistic forecast as a suggestion instead of a signal.


The model was right nine times out of twelve. The model was uncertain on the thirteenth, and the uncertainty was real. The model did not replace our judgment. It gave us a structured, calibrated, quantified input to our judgment. And we used it the way you use a weather forecast when you are deciding whether to bring an umbrella: we looked at it, we nodded, and we walked outside in the rain.


The cost of ignoring a good prediction is not the cost of the prediction being wrong. It is the cost of the resources allocated as if the prediction did not exist. It is the cost of the even split. It is the cost of the memo that was not written. It is the cost of the $53,000 that sat in the analytics dashboard, unspent, unallocated, unclaimed — the value that would have been generated if we had trusted the signal we had been given.


The model was not a replacement for taste. It was a mirror. And we looked at the mirror, saw our reflection, and walked away.