Proof, Not a Pilot
Proof, Not a Pilot
The result, first
In this story, every drop-off is a customer who said yes once, started the process, and then didn't complete the final step. A large recommerce and device resale platform ran a controlled test on roughly 11,300 of exactly those customers.
- Control group, about 4,300 leads: existing marketing messages only. Visit rate 15%. Sales completion rate 6.5%.
- Test group, about 7,000 leads: the same marketing, plus outreach from the omnichannel AI sales agent. Visit rate 16.5%. Sales completion rate 7%.
- Net result: about 9% more visits booked and about 7% more sales completed, both measured against a group that looked identical going in, and a return of roughly 7X on the spend required to run the outreach.
That's the whole result. Everything below is about why it's trustworthy, and what it says about a problem far bigger than one company.
What the "before" actually looked like
Before any AI outreach ran, the company already had a marketing motion in place, WhatsApp, SMS, and RCS messages sent automatically to anyone who dropped off mid-journey.
It mostly didn't work.
Roughly 75% of drop off customers were never recovered by that messaging alone. Not because they'd changed their minds. Because a broadcast message can't do four specific things a live conversation can:
- Confirm the customer still wants to sell
- Answer a question that's actually stopping them
- Capture a preferred visit date and time on the spot
- Rebook a visit that was missed or needs rescheduling
That gap is where the omnichannel AI sales agent went to work, calling each drop off lead directly, confirming intent, and booking the pickup inside the same conversation rather than hoping a text message would do it. The result was just over fifty additional completed sales that the marketing-only baseline had been leaving behind, every one of them revenue the company had already earned by getting the customer that far, and was simply losing at the last step.
Why a control group changes what a number is worth
Here's the distinction most case studies skip. A vendor showing a lift against last quarter is really telling you two things happened at once: time passed and the product got used, and asking you to assume the product caused all of it.
Picture the version of this story without a control group: drop-off recovery improved 9% after AI outreach, measured against the prior month. That claim might be true. It might also be true the prior month was slower, or a pricing change made offers more attractive on its own, or the team simply worked the list harder once someone was watching.
None of that is visible in a single before and after number.
A controlled test only lets one thing vary. Whatever difference shows up between the two groups didn't come from timing or effort, it came from the omnichannel AI sales agent being the only variable that changed. Everything else, the customers, the drop off point, the calendar, stayed identical across both groups.
This is also why the number can afford to be modest. Seven times return on spend is not the largest multiple sitting in this project's own data, other deployments have shown steeper figures. But it's a number earned under harder conditions, tested against a real counterfactual, not just a real result measured against nothing.
[Insert image here: a simple side by side visual, "measured against the past" versus "measured against a control group," to make this distinction visually obvious rather than only argued in text]
What's actually happening industry-wide
Two different kinds of evidence say this problem is bigger than one recommerce platform.
The first is old and still holds up. It also happens to be about exactly what outbound calling actually is at its foundation, reaching someone before they reach you. In 2007, MIT researcher James Oldroyd, working with InsideSales.com, analyzed three years of call data across six companies, more than 15,000 leads and over 100,000 call attempts. The finding: contacting a lead within 5 minutes instead of 30 drops the odds of making contact by roughly 100 times, and the odds of qualifying that lead by about 21 times. A follow up study published in Harvard Business Review in 2011, auditing 2,241 US companies, found the average firm took 42 hours to respond to a new lead at all.
The second is current. Salesforce's 2026 State of Sales report, its seventh edition, found that 94% of sales leaders already using AI agents call them critical for meeting business demands, and that roughly 9 in 10 sales teams are either using agents now or expect to within two years. That's not a prediction anymore. That's most of the market moving at once.
Put those two kinds of evidence next to the result at the top of this piece, and the shape of the argument is hard to avoid. Independent researchers with no reason to sell anything have been measuring this exact gap for almost two decades. Inside that same gap, one controlled test just showed it can be closed, not narrowed, with a return that held up against a real comparison group rather than an assumption. That combination is closer to an obvious move than a judgment call.
Coverage still needs proof, not just a claim
It's easy to claim an AI layer recovers interactions that would otherwise be lost, across a call, a WhatsApp message, or a text. It's harder, and considerably more convincing, to prove it against customers left exactly where they started, and let the gap between two matched groups do the talking instead of the pitch.
That's the standard this result was actually held to.
Sources referenced
- Anonymized controlled A/B test, large recommerce and device resale platform, figures altered from the original deployment per this project's standing anonymization rule.
- Oldroyd, J. (2007). Lead Response Management Study, MIT / InsideSales.com.
- Oldroyd, J., McElheran, K., Elkington, D. (2011). The Short Life of Online Sales Leads, Harvard Business Review.
- Salesforce (2026). State of Sales Report, 7th Edition.