Every coaching vendor has a before-and-after slide. New hires were scoring low, we coached them, now they score high, look at the jump. The slides are impressive and mostly meaningless, because they’re built the one way that guarantees a flattering number whether the coaching worked or not. So rather than show you a chart and ask you to trust it, this piece does something more useful: it walks through exactly how you’d measure whether coaching took new hires to veteran level — honestly enough that you could run it, and catch us if our own numbers didn’t hold up.
Yes, coached new hires can reach or beat veteran-level quality on a specific behaviour — that part is real and repeatable. The catch is proving it. A before-and-after that compares the coached group to its own starting point will show a gain even if the coaching did nothing, because low scores drift upward on their own. The honest version measures the coached cohort against a matched group of new hires who weren’t coached on that behaviour. What survives that comparison is the part you actually caused.
We’re running this as a method rather than a case study on purpose. The specific customer results that would sit here are anonymised and only shared once cleared, and a number you can’t interrogate is exactly the kind of proof this whole approach exists to reject.
Why new-hire ramp is the cleanest thing to measure
Start with why this particular result, new hires reaching veteran level, is worth isolating. Ramp time is expensive and slow. A new agent commonly takes months to reach full productivity, and in that window they’re costing you supervision, making more mistakes, and more likely to leave before they ever pay back the hire. Anything that shortens the ramp has obvious, legible value.
It’s also the cleanest place to run a causal measurement, which is the real reason to use it as the worked example. When you onboard a cohort of new hires, you have a natural comparison group sitting right there: other new starters, same intake, same training, same call types, same starting point. Coach some of them on a specific behaviour and not others, and you’ve built the comparison the method needs without contriving anything. New-hire cohorts hand you a matched baseline for free, which is rarely true elsewhere.
The measurement, step by step
Here’s the method laid out so you could reproduce it. It’s the causal approach the whole series rests on, applied to the specific job of proving a ramp-time gain. It’s the score, coach, re-measure loop run on a single behaviour, for a fresh cohort where the comparison group is built in.
Pick one behaviour and score it across the whole cohort. Not a general “quality” number — a specific, observable behaviour, scored the same way on every call, for every new hire, from day one. You need the starting level for everyone, coached and not.
Split the cohort into coached and untrained-peer. Some new hires get the coaching on that behaviour: voice role-play built from real anonymised calls, targeting the exact gap. A matched group of their peers (same intake, comparable starting scores) doesn’t, on that behaviour. This is the step the flattering slide skips, and it’s the only step that makes the result mean anything.
Re-measure the same behaviour on both groups’ later real calls. Not in a test. On the actual customer calls both groups take over the following weeks. The coached group changes. So does the peer group — carried by ordinary improvement, the call mix settling, and the fact that low early scores drift up on their own regardless of coaching.
The causal gain is the difference between the two changes. Subtract the peer group’s improvement from the coached group’s. What’s left is the part coaching caused, and it’s the only part you should report. If the coached new hires pulled clear of the untrained peers and reached the level of your experienced agents on that behaviour, that’s a real, defensible claim. If both groups rose together, the coaching added nothing on top of ordinary ramp, and you’ve learned that cheaply instead of paying for it.
The same result, measured two ways
This is the part worth sitting with, because it’s the difference between a number that survives scrutiny and one that collapses under the first hard question.
Take a coached cohort whose average score on the target behaviour started low and rose over eight weeks. Measured the flattering way, coached group against its own starting point, you’d report the full rise as your result. It looks excellent. It’s also inflated, because you selected these agents when they were new and scoring low, and low scores contain more bad luck than skill and drift back up on their own. Part of that impressive rise would have happened with no coaching at all.
Now measure it the honest way. The untrained peers rose too, over the same eight weeks, carried by exactly that ordinary drift. Subtract their improvement, and the causal gain is the gap that’s left — smaller than the headline rise, and real. That smaller number is the one you can take to a finance director, because it’s the one that’s still true after they’ve asked how you know.
A vendor showing you the first number is either not measuring carefully or hoping you won’t. The gap between the two is the size of the exaggeration the whole industry quietly reports.
Why our own numbers are smaller than they could be
This is exactly why Future Ready’s stated return is an average 3.7x rather than something splashier. That figure is a causal gain measured against an untrained-peer baseline, not a before-and-after we let drift in our favour. Measured the flattering way, our numbers would be bigger. We don’t measure the flattering way, and the restraint is the point: a number built like this survives your CFO, and a bigger one built the usual way wouldn’t. The discipline that makes the figure smaller is the same discipline that makes it trustworthy.
Honesty about the edges belongs here too. This is a live operation, not a laboratory. Where you can’t cleanly randomise who gets coached, you match peers as closely as the data allows and you say so. A method you’re candid about the limits of is more trustworthy than one sold as airtight — and a case study that shows its own comparison group is more convincing than one that hides it behind a bigger arrow.
What a repeatable proof looks like on your floor
Picture a Nordic contact centre onboarding a new intake, scoring source-language calls from day one, coaching half the cohort on a single behaviour their veterans do well and new hires typically fumble. Eight weeks on, the coached new hires match the veterans on that behaviour and the untrained peers haven’t closed the gap — and because the comparison group is right there in the same intake, the result is causal, not a hopeful reading of a rising line. That’s a claim the operation can defend to its own board and, if it’s a partner, to its client’s.
That’s the whole method, and it’s yours to run. The value isn’t a chart of somebody else’s new hires. It’s a way of measuring your own that produces a number you’d still believe after you’ve tried to break it.
Name the skill your veterans have that your new hires don’t yet. We’ll map the measurement onto your own intake, coached cohort against matched peers, and show you the causal gain you’d be able to defend, not a chart you’d have to take on trust.
