Message Testing With Buyer Words vs "Professional" Copy: The Ads Got Cheaper.
Message testing guide with a documented Google Ads case: cost per conversion down 49.3% after rebuilding ads in buyer language. Method, prices, and sample sizes included.

At 23:07 on a Tuesday, the client answered a routine check-in question, "How did November go? Any wins?", with this:
"Two new customers: one ordered a carbon footprint assessment for around 800 EUR, the other purchased trees for 38K EUR (!)"
When we started with this client, a German tree-planting and CO2 consulting company, there was no revenue tracking on the site at all. Nobody could say what the marketing brought in, because nothing was measured. So the first move was boring: install the tracking. The second was the one this guide is about: rebuild the pages and ads in their buyers' own words. What happens to your donation, what you get, what it costs, phrased the way customers phrase it, not the way experts write.
Three months after the tracking went live, it read 13,110 EUR a month in organic revenue. In the same period, the Google Ads account cut its cost per conversion by 49.3% to 56.69 EUR and lifted conversions 151.2%. Different channels, same rebuilt message. Screenshots on file.
That's message testing: finding out which words sell while the budget spends, instead of after it's gone.
This guide covers the full method, the real costs the testing industry won't print, and the sample-size truth. The methods we used with that client are now packaged in DomMCP, $9 one time, inside Claude. Start where the client did: with the words buyers already said.
TLDR: What Should You Know About Message Testing?
- Message testing checks which message real buyers respond to before your budget commits to it. It compares promises and angles, not wording polish, and it is not the same thing as A/B testing a finished page.
- The case in this guide, a German climate client: Google Ads conversions up 151.2%, conversion rate at 3.2% (up 66.9%), cost per conversion down 49.3% to 56.69 EUR. Separately, the organic revenue tracking installed at the start of the engagement reached 13,110 EUR a month within three months.
- The method: mine real buyer language first (Sherlock builds this as a Customer DNA report), score your draft against 7 buying triggers (Watson, $9, report in about 10 minutes), pull competitor complaints for differentiation (Moriarty), then let the live account give a directional verdict.
- The whole loop runs inside Claude with DomMCP for a $9 one-time trial: paste a connector link into Claude's settings, about two minutes of setup, one full run of each tool included, money back if it tells you nothing new. No research panel involved, and nothing here requires a statistics background.
What Is Message Testing?
Message testing is the practice of checking which message resonates with real buyers before you commit ad budget, a page redesign, or a launch to it. The output is a decision: this promise leads, that one gets dropped. Message testing helps you evaluate how well each promise lands with your target audience before the spend, not after.
The category converges on five things worth checking: clarity, relevance, credibility, differentiation, and value. For B2B messaging specifically, Gartner defines the market around exactly this job, collecting feedback from target buyers on messaging before it ships (B2B message testing solutions).
Message testing and copy testing get used interchangeably. They aren't the same job:
| Message testing | Copy testing | |
|---|---|---|
| Decides | Which promise or angle leads | How the chosen words execute |
| Typical question | "Pain-led or gain-led?" | "Is this draft clear and credible?" |
| When | Before drafting variants | On a finished draft |
| Validated by | Sales-linked research (ARF) | The same body of research |
Does testing before spending actually predict revenue? The strongest public answer comes from the field's own research body. The Advertising Research Foundation's Copy Research Validity Project, the only public study to validate copy-testing measures against real split-cable sales results, concluded plainly that copy testing works: its strongest measure picked the better-selling ad in 84% of pairings. Across 12,000 to 15,000 interviews, the commercials in the tested pairs were producing real sales differences of 8% to 41%. Choosing the wrong message meant a double-digit sales gap.
Most message testing methods run the same loop: objective, audience, test, analyse, iterate. It's a fine loop. It also skips the two questions that decide whether any of it matters: where does the message come from in the first place, and what does testing cost? Message testing can help with the first: the best message usually already exists in your buyers' own words, and the method below digs it out. On cost, we answered the question for a site that lost roughly 94% of its traffic the expensive way. This guide answers both directly.
Why Does "Professional" Copy Keep Losing?
Polished copy optimises for how the writer sounds; buyers reward how easily they understand. That gap is where most message tests are lost before they begin.
The assumption goes: make it sound more professional and it will convert better. You've probably lived the other side of it. You rewrite the page, it reads better to you, and the numbers don't move. "I keep polishing all of them" is how one marketer put it. Version 9 of the homepage, and still no way to tell which version is right.
The research says polish points the wrong way. In five experiments at Princeton and Stanford, psychologist Daniel Oppenheimer found that needlessly complex wording made authors seem less intelligent, not more, and the effect held whatever the quality of the underlying writing. 86.4% of the students polled admitted deliberately complicating their writing to sound smarter. Readers punished them for it. The mechanism is processing fluency: text that is easier to read is trusted more.
Our climate client's pages were a live demonstration. Written by experts, dense with sector language, and impenetrable to everyone else, including the people paid to advertise them. Daniel, from the client's ads agency, works deep in the client's subject, and he said it straight: "If we didn't even understand it at first glance... no layperson visiting the page for the first time is going to understand it either."
Harvard Business Review named this trap the curse of knowledge: once you know your offer inside out, you can no longer hear your own copy the way a first-time visitor does. Insiders compress rich private knowledge into abstract phrases; outsiders hear only the abstraction. Which is why internal sign-off is not a message test. Everyone in the approval chain has the same curse.
AI drafting makes this worse, not better. A model optimises for fluency, so the draft always reads fine. But fluency is exactly what Oppenheimer showed to be insufficient, and no prompt can supply the words your buyers use if the model never saw them. If your AI copy sounds like everyone else's, the input was missing evidence, not craft. It's not your prompts. It's missing evidence. The conversion rate optimization tools that show you where visitors drop off can't tell you this either; they show the where, never the why.
What Does Message Testing Actually Cost?
No page ranking for "message testing" prints a price, so here is the cost table. Panel platforms quote per test, after a sales conversation. Research networks want a discovery call. Agencies bill by project. Annual contracts for the panels you have already priced up typically run four figures and up, which is why the objection we hear most is some version of: "I can't justify that for one page." (We price the two panel routes line by line in the Wynter alternative comparison and the PickFu alternative comparison.)
The objection is rational. A message test should never cost more than the ad budget it protects. We hold our own work to that rule; the Justin Welsh case study exists because evidence should be cheaper than the mistake it prevents.
Here is the decision ladder, priced and timed:
| Step | What | Cost band | Turnaround | Right when |
|---|---|---|---|---|
| 1 | Watson copy diagnostic | $9 one-time (DomMCP trial) | About 10 minutes | Any page, any budget: find the leak first |
| 2 | Sherlock Customer DNA (buyer-language evidence) | Included once in the $9 trial | Same session | The message needs rebuilding, not polishing |
| 3 | Moriarty competitor teardown | Included once in the $9 trial | Same session | Differentiation is unclear |
| 4 | Panel test | Four figures a year, typical | Under 48 hours to weeks | Budget exists and the stakes justify it |
| 5 | Live A/B test | Free to set up, priced in traffic | Weeks at small-site traffic | Thousands of users per variant available |
DomMCP's $9 one-time trial includes one full run of Watson, Sherlock, and Moriarty inside Claude. On speed: Watson's report is in your hands in about 10 minutes, the fastest panel turnaround we could verify is under 48 hours, and agency research routes run weeks.
How Do You Test Messaging Without a Panel? The 5-Step Method
The highest-value message test happens before the variants exist: in the language source. Panels and market research firms test opinions about drafts. This methodology tests marketing messages against evidence, and the evidence is already written.

- Mine the words. Pull your reviews, sales-call notes, and support threads. Extract what customers care about verbatim, the dreams, objections, fears, frustrations, and roadblocks, in the buyer's own phrasing, not your summary of it. These are the words that resonate with your audience, because your audience wrote them. This is the job Sherlock does automatically: a Customer DNA report built from real buyer language.
- Steal the enemy's complaints. Read your competitors' 1-star and 3-star reviews. Every recurring complaint is pre-tested demand for a different promise. Moriarty runs this teardown from competitor customers' complaints.
- Draft two messages. One built from the mined buyer words. One "professional" version, the control you would have written anyway. The same split works whether you are testing a homepage promise or a product message.
- Score before spending. Diagnose both drafts against the 7 buying triggers and fix the highest-ranked leak first. Watson does this for $9, with the report back in about 10 minutes.
- Let the account judge. Put both versions live as ads or landing pages on a modest budget and watch cost per click, conversion rate, and cost per conversion. Google's auction rewards ad relevance, one of the components of Quality Score, so a directional verdict shows at modest traffic. Directional is the right word: this is evidence for a decision, not a significance certificate.
Before you spend anything, you can score a draft by hand. Ask it seven questions, the same dimensions Watson scores:
- Does the first line name a problem the buyer recognises as theirs?
- Could a stranger say what you're offering after one read?
- Is there a number or proof point a sceptic could check?
- Does it answer the buyer's biggest recorded objection?
- Would your competitor's unhappy customer see a difference here?
- Is the next step obvious and priced?
- Would the buyer recognise their own words anywhere in it?
A no is a leak. Watson runs the same diagnostic automatically, ranks the leaks, and names the first fix. Then you test and optimize from evidence instead of taste.
Why not just generate 20 AI variants and test them all? Because variants from the same evidence-free prompt are the same message in 20 outfits. The buyer-language approach to keywords works for the same reason: the source material differs, so the output differs.
What this method will not do: it will not give you statistical significance, it will not replace a brand tracker, and it is not a substitute for talking to customers when you have access to them. It is the lowest-cost reliable way to stop guessing.
Message Testing Case Study: What Happened When a Real Account Ran It?
A German climate client ran this exact sequence, and both channels moved within three months. Tree-planting and CO2 consulting, marketing efforts running on Google Ads, and at the 2024 baseline: a punishingly high cost per click and five enquiries to date. No organic revenue tracking existed in GA4, so the engagement started by installing it.
The diagnosis matched the curse of knowledge section above. The consulting pages were expert-written and opaque; Daniel's "no layperson will understand this" verdict came from this account. As I put it on the same call, companies deep in their own subject "know 'we offer A, B, C, D, E' but can't get it across".
The intervention was the 5-step method. Buyer language mined from calls and enquiries. Pages rebuilt around the questions buyers actually asked: what happens to a 1,000 EUR donation, what do I get, what does it cost. Transparency where the old pages had abstraction. Campaigns split so plain-language offers could be measured against the old expert framing.
The receipts, channels kept separate:

| Metric | Before | After | Change |
|---|---|---|---|
| CPA, Google Ads (agency's words, Jan 2025) | "we were in the range of 100, 110" | "now partly at 50" | roughly halved |
| Cost per conversion, Google Ads (9 Nov to 9 Dec 2024) | 56.69 EUR | down 49.3% | |
| Conversions, Google Ads (same window) | 112.9 | up 151.2% | |
| Conversion rate, Google Ads | 3.2% | up 66.9% | |
| Organic revenue (GA4, tracking installed Aug 2024) | no tracked data before engagement | 100.6 → 8,040 → 13,110 EUR/month (Sept to Nov) | 13,110 EUR/month within 3 months of tracking |
Internal client data, screenshots on file.
The client's own November summary, verbatim from WhatsApp: "Two new customers: one ordered a carbon footprint assessment for around 800 EUR, the other purchased trees for 38K EUR (!)"
And my own read on the account, mid-engagement: "We spent less but brought in more conversions."
One caveat though. This is one account, in a lead-generation context, with seasonality we can't fully separate out, and the organic revenue series starts at zero partly because the tracking was new, not because the business had never earned a euro. The point of the case is the mechanism: when the message became something a stranger could understand in one read, the business impact showed on every channel we could measure. It is evidence for the method, not a promised multiple. Buyer-worded messaging pairs naturally with bottom-of-funnel keyword targeting, where the searcher is closest to buying.
Which Sample Sizes Do You Actually Need?
The message testing research on page one of Google contradicts itself on sample sizes, so here is the reconciliation. One source says five people. Another says saturation. A third implies you need a stadium. They are all answering different questions.

| Threshold | Source | What it buys you |
|---|---|---|
| 5 people, 5 questions, 5 minutes | WHO rapid message testing | Fast directional feedback on one message |
| 9 to 17 interviews | Hennink and Kaiser, 2022 | Qualitative saturation on a narrow question |
| Thousands to tens of thousands per variant | Kohavi, Deng and Vermeer, KDD 2022 | A statistically rigorous A/B verdict |
The WHO's guidance is refreshingly blunt for a budget method: interview 5 people, ask 5 questions, take no more than 5 minutes. It exists because health agencies need message feedback in days, not quarters.
For the classic route of focus groups and in-depth interviews, the ceiling is lower than most marketers fear. A 2022 systematic review in Social Science and Medicine by Monique Hennink (Emory University) and Bonnie Kaiser (UC San Diego) found qualitative studies typically reached saturation within 9 to 17 in-depth interviews, or 4 to 8 focus groups. Across the 23 studies reviewed, new interviews stopped producing new insight within that narrow range, provided the audience was homogenous and the question specific. A defined buyer profile with a specific question does not need a 500-person panel. It needs the first 10 to 15 real buyer voices, which your reviews and call notes already contain.
The A/B threshold is higher than most marketers hope. A 2022 paper at the ACM's KDD conference by Ron Kohavi (who ran experimentation at Microsoft and Airbnb), Alex Deng and Lukas Vermeer ran the numbers: at a 3.7% baseline conversion rate, detecting a 10% relative lift at standard rigour takes 41,642 users per variant. The same paper dissects a real test that claimed a 337% lift from about 80 visitors per variant. The guidance is thousands of active users per variant at minimum, preferably tens of thousands.
Below those thresholds, running a split test collects noise with confidence attached. Collecting evidence instead, the words buyers already wrote and said, is not the budget compromise. At small-site traffic it is the rational first step, which is why the method in this guide starts there. If your traffic itself is the bottleneck, these search optimisation tips are the companion piece.
Where Should You Start With Message Testing?
Start at the lowest-cost step that produces evidence, and only climb the ladder when the stakes demand it. For most readers of this guide that means the DomMCP $9 one-time trial: one full run each of Watson, Sherlock, and Moriarty, inside Claude, set up by pasting a connector link into Claude's Settings under Connectors. Takes about two minutes, and you need no technical skills.
Which detective first depends on your question. Watson, if the question is "which line is losing the sale": it investigates the page, scores the seven buying triggers, ranks the leaks, and the report is back in about 10 minutes ($9 one-time on its own at getwatson.io, no subscription; the full CRO audit framework shows where that diagnosis sits in the bigger audit). Sherlock, if the question is "what should the message even say": the Customer DNA report builds the full psychological profile of your real buyers from their own words. Moriarty, if the question is "how do we position against the incumbent": the teardown works from your competitor's customers' complaints, the fastest read on brand positioning gaps.
The guarantee, exactly as the product page words it: "Try it without risk. Run all 6 tools once each. If you don't find a single thing worth changing in your marketing, reply to your receipt and get the $9 back."
If you have genuine panel budget and the traffic for step 5, use them. They sit at steps 4 and 5 of the ladder for a reason: they answer narrower questions with more certainty, and they answer them better when the message they test was built from evidence in the first place. This method is the evidence layer under any panel test, a data-driven approach that costs $9 to try, not a religion against panels. For done-for-you work, the studio takes a small number of client engagements.
Ready to Test Your Message for $9? See What DomMCP Finds
Frequently Asked Questions
What is message testing?
Message testing checks which message resonates with real buyers before you commit ad budget, a page redesign, or a launch to it. It evaluates clarity, relevance, credibility, differentiation, and value, and it answers "which promise wins", not "which wording sounds nicer". It is not A/B testing: an A/B test compares finished variants; message testing decides what the variants should say, which is how it helps increase conversion rates before the budget spends.
What is the difference between message testing and copy testing?
Message testing picks the promise; copy testing checks the execution. Message testing compares angles, meaning which dream, pain, or proof leads. Copy testing scores a finished draft. The ARF's validation research found its strongest copy-test measure picked the better-selling ad in 84% of pairings. Neither can rescue a message your buyers never said.
How much does message testing cost?
Panel platforms quote per test after a sales call, and annual contracts typically run four figures and up. DomMCP is a $9 one-time trial with one full run of Watson, Sherlock, and Moriarty inside Claude. Live A/B testing is free to set up but priced in traffic. "Try it without risk. Run all 6 tools once each. If you don't find a single thing worth changing in your marketing, reply to your receipt and get the $9 back."
How do you test brand messaging without a panel?
Mine the language your buyers already produced, then let your live channels judge. Extract dreams, objections, fears, and frustrations verbatim from reviews, calls, and support threads; draft one message in buyer words and one polished control; score both against the 7 buying triggers; run both as live ads. Treat the result as directional evidence, cheaply bought.
How many respondents do you need for message testing?
Fewer than the folklore says, and more than a small site's A/B test can deliver. The WHO's rapid method uses 5 people, 5 questions, 5 minutes. Qualitative research typically saturates at 9 to 17 interviews. A statistically rigorous A/B test needs thousands of users per variant at minimum. Below those thresholds, mined buyer language plus a scored diagnostic gives evidence without the traffic bill.
How do you test AI-written copy?
The same way you test human copy, plus one step: check what evidence the draft was fed. AI drafts read fluent because the model optimises fluency, not persuasion. Feed the prompt real buyer language first, then diagnose the draft against the 7 buying triggers before spending on traffic. If the draft sounds like everyone else's, the input was missing evidence, not craft.
When should you conduct message testing?
Consider testing before money commits to the message, and again when the numbers drift: before a launch, a rebrand, or a new marketing strategy goes live; when your existing messaging stops converting, cost per click rises, or conversion falls without a tracking explanation; when traffic arrives but sales don't. Testing costs least before the ad budget spends. The next best time is now.
What do you need to run DomMCP?
A Claude account and about two minutes. Paste the connector link into Claude's Settings under Connectors; no technical skills needed. The $9 one-time trial then includes one full run of each tool, and Lestrade, the guided intake interview, stays free and unlimited.
Get Updates Like This
Case studies like this one land in the newsletter first, alongside AI marketing that converts and SEO that converts better. You get the receipts as they happen. I won't spam you, I promise.
Sources
- The ARF Copy Research Validity Project, Journal of Advertising Research (split-cable sales validation; 84% pairing accuracy)
- Oppenheimer, Consequences of Erudite Vernacular Utilized Irrespective of Necessity, Applied Cognitive Psychology, 2006 (processing fluency experiments)
- Heath and Heath, The Curse of Knowledge, Harvard Business Review, December 2006
- Kohavi, Deng and Vermeer, A/B Testing Intuition Busters, KDD 2022 (sample-size worked example, 41,642 users per variant)
- Hennink and Kaiser, Sample sizes for saturation in qualitative research, Social Science and Medicine, 2022 (saturation at 9 to 17 interviews)
- WHO, Message testing, the 5-5-5 rapid method
- Gartner, B2B Message Testing Solutions market
