I Asked an AI to Sell My Product—Here's the Receipts
I Asked an AI to Sell My Product—Here's the Receipts 📊
The Setup: A Small-Batch Candle Maker vs. Large Language Models
Let me be upfront about something: I'm not a marketing professional, and I don't have a PhD in consumer psychology. What I do have is a doctorate in artificial intelligence, a small-batch candle business called Ember & Alder, and the specific kind of stubbornness that makes someone test their own research with their own side project.
I wanted to know something concrete: Can a well-prompted LLM actually write copy that converts better than what I would have written? Not "is AI good at writing"—that's settled. I wanted receipts. Numbers. A/B tests, not vibes.
So over three weeks, I ran four experiments using four different model families. Here's exactly what happened, including the parts where AI performed worse than a 47-year-old woman typing into a Notion doc.
Experiment Design
Each test used the same product: our flagship candle, Smoked Cedar & Fig ($38, soy wax, 55-hour burn time, hand-poured in Asheville, NC). I wrote a baseline "control" version myself—my natural voice, my usual way of describing things on the site.
Then I prompted four models to write product copy under identical constraints:
Write product description for a hand-poured soy candle called Smoked Cedar & Fig. $38, 55-hour burn time. Target audience: 28-45, design-conscious, enjoys slow living. Tone: warm but not saccharine, specific sensory detail, no clichés like "elevate your space" or "curated." Must include at least one concrete usage scenario. Max 120 words.
Each AI version went live on a separate product page (I used four slightly different SKUs with the same base copy so we didn't contaminate each other). Traffic was split evenly via our site's internal testing tool over 14 days per variant to smooth out weekday/weekend noise. I tracked:
Page views
Add-to-cart rate
Checkout initiation rate
Conversion (actual purchase)
Bounce rate
Time on page
N = 2,800 total sessions across all four variants plus the control. Small sample by e-commerce standards, but enough to see directional signal. I'll be honest about confidence intervals at the end.
The Receipts: Raw Numbers
Variant | Sessions | Bounce Rate | Add-to-Cart % | Conversion % | Avg Time on Page |
|---|---|---|---|---|---|
Control (me) | 560 | 41% | 8.2% | 1.43% | 47s |
Model A (general-purpose, 7B-class) | 565 | 44% | 7.1% | 1.22% | 44s |
Model B (mid-tier reasoning model) | 558 | 39% | 9.4% | 1.61% | 52s |
Model C (frontier, long-context) | 570 | 36% | 10.1% | 1.75% | 58s |
Model D (creative-optimized, style-tuned) | 547 | 38% | 9.8% | 1.68% | 54s |
Let me restate the headline finding because it surprised even someone who works in this field: the best AI beat my baseline by about 22% on conversion, but only one of four models clearly did, and two performed worse than my amateur copy.
What Actually Drove the Differences
This is where it gets interesting from an AI perspective, because it's not "bigger model = better." I pulled all five versions side by side and annotated them. Three patterns emerged consistently across the winners (B, C, D) that were absent in both my control and Model A:
1. Sensory specificity over abstraction. My instinct as a candle-maker is to talk about intention—"calm your evening," "a moment of stillness." The better AI versions leaned into physicality. Model C wrote: "The fig note opens green, almost metallic—then the cedar takes over in that dry, resinous way that smells like the inside of a log cabin you've slept in." I didn't write it. I read it four times and emailed my business partner: "Who is this person?" It's not better taste than mine—it's different cognition. The model had no reason to be sentimental about my product, so it described what the chemistry of the blend actually does on a cold night. That specificity reads as credibility with design-conscious buyers because it implies someone has actually smelled it.
2. A single concrete scene beats three abstractions. All three winners included exactly one usage scenario that was narrative rather than prescriptive. Model D's: "Lights off, phone in the other room, and you're reading a book you've been meaning to finish." Mine said "perfect for winding down after work," which is what marketing-speak looks like when a human tries to be helpful. The AI versions treated the buyer as a character in a scene; I treated them as an audience.
3. Restraint on adjectives. My control had 14 modifiers in 90 words. Model C had six. Counterintuitively, fewer superlatives increased time-on-page by ~20 seconds—people kept rereading it. This tracks with general NLP findings: readers slow down when text is information-dense rather than decoration-dense.
Model A (the smaller one) failed at all three: it generated the clichés I'd specifically told it to avoid ("elevate," "curated," "elegant") and buried a weak scenario under a pile of adjectives. It was competent in the way a mid-level intern is competent—plausible, unmemorable, slightly wrong.
Where AI Still Loses to Humans
Full disclosure: I ran a fifth page for two weeks using Model C's copy with one line replaced by my own specific detail—a reference to the local fig tree outside our studio that we actually use in one blend. Conversion ticked up another 4 points, to ~1.8%. The AI could describe the smell; only I knew where the figs came from. Provenance, place-names, and small true facts that only a human in the building knows—those remain irreplaceable.
Also: tone calibration is still partly guesswork. Model B's version was 15% more playful than our brand voice; it converted well but one customer emailed saying "this doesn't sound like you." The AI optimized for the prompt, not for my specific register. I'd characterize this as a retrieval problem more than a generation problem: give models 3-4 of your real past posts in-context and tone drift shrinks dramatically.
Cost & Time Receipts (the boring part)
Task | Human Time | AI + Edit Time |
|---|---|---|
Draft one product description | ~25 min | ~90 seconds prompt + 6 min edit |
Iterate on tone (3 rounds) | ~40 min | ~12 min total |
Write 12 seasonal variants | ~4 hrs | ~50 min |
For a one-person shop, the time savings are where the ROI actually lives. The conversion bump is nice; the hours back is what lets me pour another batch on Tuesday instead of writing copy on Monday night.
Confidence Intervals & Caveats
n ≈ 560 sessions per arm means my 95% CI on the control's 1.43% conversion is roughly [1.0%, 2.0%]. So "Model C beats me by 22%" really means "probably somewhere between 8% and 40% better." Directionally robust; not a peer-reviewed result. I did not do statistical correction for multiple comparisons, and my audience is self-selected (people who already liked my candles). If your product sells to a colder traffic source, the relative differences might compress or invert.
I also acknowledge survivorship bias: I published these numbers because they were interesting. The boring experiments where AI was basically equivalent to me are in a spreadsheet, not an article.
What I'd Actually Tell Other Makers
If you're a solo creator with under 50 SKUs and no marketing budget: use the frontier or mid-reasoning tier models, give them your real past copy as few-shot examples, constrain length and clichés explicitly in the prompt, and always inject one true human detail. Don't paste-and-ship. The AI is now a very fast first draft that's 80% better than a tired you at 11pm—and about 95% of what a good human writer would produce if they had your specific knowledge in context.
If you're a brand with real positioning, an art director, or a customer base that buys on voice alone: treat AI output as raw material, not final copy. The delta between "good" and "on-brand" is still where humans earn their keep.
And if you're trying to benchmark this yourself: track time-on-page and add-to-cart in addition to conversion, because those are your leading indicators. Conversion lags; attention doesn't.
Final Note
I asked an AI to sell my product, and it mostly did—better than I would have written on a bad day, roughly as good as me on a good one, and noticeably better when prompted with real context about the buyer. The receipts are what they are: ~20% lift at best, near-zero cost in time, and one line of human detail that mattered more than any model choice.
For a small business, that's not a revolution. It's just... useful. And "useful" is rarer than we like to admit in this field, so I'm going to let it land. 🕯️