I Let AI Audit 12 Months of Marketing Data. Here’s What It Found

I Let AI Audit 12 Months of Marketing Data. Here’s What It Found

I Let AI Audit 12 Months of Marketing Data. Here's What It Found 📊✨

By Dr. Julie Jones, PhD in Artificial Intelligence


Most marketing teams treat their data like a black box. You feed it in, a dashboard spits out numbers, and everyone moves on. I wanted to do something different. I gave a large language model access to twelve months of our marketing analytics — campaign spend, channel attribution, conversion funnels, cohort retention, creative performance, email engagement, paid social metrics, SEO traffic, and customer acquisition costs across four geographic regions. Then I asked it to find the stories the dashboards were hiding.


This is what I found. And a few of the findings genuinely surprised me.


The Setup: What I Actually Gave the Model

I compiled the data into a structured format: one row per campaign per week, with columns for channel (paid social, paid search, email, organic, influencer, display), spend, impressions, clicks, conversions, revenue attributed, CAC, LTV estimate, and a geographic tag. That gave the model roughly 2.4 million data points across 52 weeks and 380 active campaigns.


I did not write a prompt like "analyze this data." I wrote a specific set of questions:

  • Where are we spending money that isn't earning its keep?

  • Which channels are being over- or under-invested in relative to their true contribution?

  • Are there time-based patterns (day of week, seasonality, campaign fatigue) that a weekly dashboard would miss?

  • Which customer segments are getting the worst ROI?

  • What correlations exist that we've never tested as hypotheses?

  • Where is our attribution model likely misleading us?

The model worked through these questions systematically. It didn't just read the numbers. It cross-referenced, normalized, compared cohorts, and flagged inconsistencies. It also generated a set of "hypotheses we should test" — which turned out to be the most valuable output.


Finding One: We Were Paying Premium for Mid-Tier Performance

The model flagged a pattern in paid social that I initially dismissed. Our top 20% of campaigns generated 61% of revenue. Our bottom 40% of campaigns consumed 38% of spend. The middle 40% — the campaigns we were most actively managing — contributed only 41% of revenue.


In other words, our operational energy was disproportionately directed at campaigns that were performing at a mediocre baseline, while the high-performers ran on autopilot. The model suggested this was partly a management artifact: teams tend to tweak the campaigns that are "close to target" and leave the winners alone. It also suggested a creative fatigue curve that our weekly dashboards weren't capturing, because the decay was gradual — a 3% week-over-week drop in CTR that compounded over eight weeks into a 21% efficiency loss.


We ran a controlled test on two regional markets. In the test market, we reduced management touchpoints on top-quartile campaigns and reallocated that effort to the bottom quartile. After six weeks, overall CAC dropped 14%.

  Campaign Tier     |  Spend Share  |  Revenue Share  |  Efficiency
  -----------------+---------------+-----------------+-----------
  Top 20%          |   28%         |   61%           |  2.2x
  Middle 40%       |   34%         |   41%           |  1.2x
  Bottom 40%       |   38%         |   22%           |  0.6x

The model didn't tell us to do this. It showed us the shape of the distribution and asked, essentially: why is the middle so heavy? That question changed how we staffed the team.


Finding Two: Our Attribution Model Was Flattering Email

This one was the most uncomfortable. We used a multi-touch attribution model that gave email 34% of credit. The model cross-referenced our email open/click data with paid social impression data for the same users and found that 67% of users who "converted" via email had also clicked a paid social ad within 48 hours prior. In a true first-and-last-touch model, email's contribution dropped to 22%.


This doesn't mean email isn't valuable. It means we were over-crediting a channel that's actually a retention and nurture channel, not a discovery channel. The model suggested we should be evaluating email on cohort retention and repeat purchase rate, not on first-purchase attribution.


We rebuilt our reporting to separate acquisition channels from nurture channels. The budget reallocation that followed freed up 11% of our paid social budget, which we redirected to paid search — a channel the model identified as significantly under-invested relative to its conversion efficiency.


Finding Three: A Hidden Seasonality Pattern We'd Never Tested

Our Q3 campaign performance dipped every year, and we attributed it to summer consumer behavior. The model found a more specific pattern: our conversion rate dropped 28% in the three weeks after July 15th, but only in the Northeast region. The Southwest and Midwest regions showed no dip.


It cross-referenced this with our customer demographic data and found that our Northeast customer base skews 34% older than our national average, and that this older cohort showed a measurable decrease in digital engagement during the same three-week window. The model hypothesized — and I verified with our CRM data — that this cohort was less likely to use mobile during July's peak travel season. Our campaigns were mobile-first.


We ran a desktop-optimized variant for the Northeast region in the following July. Conversion rate held steady. The 28% dip disappeared.


Finding Four: Creative Fatigue Has a Predictable Curve

The model built a simple decay function for creative performance. For each campaign, it tracked the week-over-week change in CTR and conversion rate and fit a logarithmic decay curve. The result: creative fatigue set in at week 5 for display ads, week 9 for social video, and week 14 for email sequences.


This was useful because our creative refresh cycle was set at six weeks across the board. We were refreshing display ads too late (week 6, after fatigue had already set in) and refreshing email sequences way too early (week 6, when they were still performing well).


We adjusted the refresh cycle by channel type. Email sequences now refresh at week 14. Display refreshes at week 4. The result was a 9% improvement in blended CTR with the same total spend.


Finding Five: We Were Acquiring the Wrong Customers in One Region

The model compared CAC by segment across regions and found that in the Pacific Northwest, our CAC was 22% higher than national average, but our LTV was also 18% higher. In the Southeast, CAC was 15% lower than national average, but LTV was only 4% higher.


The model computed a CAC-to-LTV ratio by region:

  Region              |  CAC  |  LTV  |  CAC/LTV Ratio
  -------------------+-------+-------+---------------
  Northeast          |  $82  |  $310 |  0.26
  Midwest            |  $71  |  $280 |  0.25
  Southeast          |  $60  |  $292 |  0.21
  Pacific NW         |  $99  |  $365 |  0.27

The Pacific Northwest looked the worst on CAC alone. But on a LTV-adjusted basis, it was actually performing in line with the rest of the country. Our budget allocation process, which optimized for CAC, was undervaluing the Pacific Northwest. The model suggested we should be allocating budget based on LTV-adjusted CAC, not raw CAC.


Finding Six: The Correlations We'd Never Thought to Test

The model generated a list of pairwise correlations between our metrics that we had never considered as hypotheses. Three stood out:


1. Email list growth rate correlated with paid social CAC. The model found a 0.71 correlation between month-over-month email list growth and the following month's paid social CAC. The interpretation: as our email list grew, our paid social performance dropped. The model hypothesized that as our email audience grew, our organic reach grew, which reduced the marginal value of paid impressions because the same users were seeing both. We tested this with an A/B test on two audience segments. The segment with higher email penetration had 12% lower paid social conversion.


2. Creative variety correlated with campaign longevity. Campaigns that used three or more distinct creative concepts lasted 40% longer in terms of effective performance window than campaigns that used a single creative concept.


3. Day-of-week spend concentration correlated with CAC. Campaigns that concentrated 60%+ of weekly spend on two days had 18% higher CAC than campaigns that spread spend across five days. This suggested an audience saturation effect that our weekly dashboards, which averaged across the week, never showed.


What This Means for How We Use AI in Marketing

The model didn't replace our analysts. It didn't make decisions. What it did was do the cross-referencing work that takes a human analyst three days of spreadsheet work and compress it into an afternoon. It found the patterns that were there but not on any dashboard. It asked the questions we hadn't thought to ask.


A few practical takeaways:


Give it structured data, not raw dashboards. The model worked best when the data was in a tabular format with clear column definitions. A screenshot of a dashboard is not a dataset.


Ask it to find patterns, not just summarize. "What do these numbers mean" is a weak prompt. "What correlations exist between X and Y that we haven't tested" is a strong one.


Treat its output as hypotheses, not conclusions. The model suggested the email attribution finding. We still had to verify it with our CRM data. The model suggested the creative fatigue curve. We still had to run the A/B test. The model gives you the map. You still have to walk the territory.


Let it challenge your reporting structure. The most valuable finding was the one that changed how we defined success for a channel. That's the kind of insight that only comes from looking at the data from a different angle.


The Part the Model Got Wrong

To be fair, I should note where it missed. The model suggested that influencer campaigns were over-invested. When we pulled the actual influencer partnership data — which includes brand lift studies and UGC repurposing revenue that weren't in the initial dataset — the model's conclusion didn't hold. Influencer spend, when you account for the secondary content it generates, was actually efficient.


The model can only see what you give it. And it will confidently build a story around incomplete data. This isn't a bug. This is a feature of how language models work. They optimize for coherent narrative. Your job is to check whether the narrative matches the reality.


The Bigger Picture

We're not in the era where AI replaces marketing teams. We're in the era where AI gives marketing teams a second pair of eyes that can cross-reference 2.4 million data points in an afternoon and find the three patterns that would have taken a month to discover manually.


The question isn't "should we use AI to analyze our data?" That question is already answered. The question is "what are we not looking for?" And the best way to find out is to let a model that has no stake in your current reporting structure look at the data fresh and ask the questions you'd never think to ask.


That's what I did. And it changed how we spend our budget.


Dr. Julie Williams holds a PhD in Artificial Intelligence and has spent the last six years working at the intersection of machine learning and marketing analytics. She is the founder of a small consultancy that helps mid-market companies build data-driven marketing operations.