Blog/100 is a scale, not a target

100 is a scale, not a target

Apple will happily give you a perfect sleep score. Garmin almost never will. The difference explains why a quality target of 90 is very good at hiding improvement.


A tall indigo ruler standing among soft cloud shapes, with a single marker high on the scale but clearly below the empty top

My wife uses Apple Watch for health-tracking and sleep scores, while I have a Garmin watch. There's something Garmin users like to joke about: the elusive 100 sleep score. Apple will happily give you a perfect night's sleep from time to time. Getting a 100 from Garmin is apparently enough of an achievement that people post screenshots of it online.

They're both measuring sleep on a scale from 0 to 100, so how can a perfect score be fairly normal on one and almost mythical on the other?

The answer, of course, is that 100 doesn't have any intrinsic meaning. The two companies designed their scales differently. Apple can decide that a sufficiently good night's sleep deserves full marks, while Garmin can reserve the very top of its scale for an unusually exceptional combination of sleep duration, recovery and whatever else goes into its algorithm.

Neither approach is necessarily wrong. But imagine taking the Garmin score and setting yourself a nightly target of 100 because, well, that's what the scale goes up to. You haven't set an ambitious sleep goal. You've misunderstood the ruler.

I think we sometimes do exactly this with quality scores.

What does 100 actually mean?

CallCoach scores behaviours such as needs analysis, objection handling, closing and tone of voice from 0 to 100. Our clients define most of what they want measured, while we've designed the scoring rubrics so that the resulting behavioural scores are roughly normally distributed.

That last part matters. A score of 100 doesn't mean "the rep did everything they were supposed to do", in the way that 100% on a compliance checklist might mean every required step was completed. It sits at the extreme end of a distribution. A 50-ish score describes something fairly ordinary, and as you move towards either end you're looking at increasingly unusual examples of that behaviour.

There are good reasons for doing this. If every competent example of objection handling scored 100, we'd lose the ability to distinguish good objection handling from genuinely exceptional objection handling. The top of the scale gives us somewhere to put those differences.

It also means that getting close to 100 on several behaviours at once should be extremely rare. And now that we've analysed enough calls, we can see just how rare.

Half a million calls later

We looked at more than half a million calls over twelve months, across many companies and industries. A call's overall score is the average of the behaviours assessed on it, and the distribution is remarkably consistent.

The median call scores about 62. Nine out of ten calls sit somewhere between 40 and 79, and only one in ten gets above roughly 74. By the time you reach 90, you're looking at something like one call in 250.

And 100? We looked at every call over that year where at least three behaviours were scored. Not one got there.

The best call of the month never reaches the top of the scale

The highest scoring call in each of twelve months, against the middle call of the same month.

50 60 70 80 90 100 everything the scorecard asks for the best call of the month the middle call of the month Aug 2025 Jan 2026 Jul 2026 call score out of 100, by month

The best call in any given month generally landed somewhere in the mid 90s. There were technically a tiny number of 100s in the data, but all of them came from calls where only one or two behaviours had been assessed. Once you're asking somebody to be exceptional at several different parts of a conversation at the same time, perfection disappears.

I don't think that's a problem with the scoring. It's what I'd expect from a scale designed this way. In fact, I'd be slightly worried if lots of calls were hitting 100, because we'd have compressed the best calls against the top of the scale and lost some of our ability to tell them apart.

But it does create a problem when somebody looks at that scale and decides the quality target should be 90.

When 70 feels like failure

This is where I think our familiarity with percentages works against us. Show somebody a score of 62 out of 100 and it doesn't feel very good. Seventy sounds passable, 80 sounds pretty good, and 90 sounds like where a strong team ought to be.

But we're importing meaning that isn't in the score.

In our data, a call scoring around 75 is already among the better calls being made. A score in the low 80s is exceptional, while 90 is an outlier. Describing a 75 as "only 75%" makes about as much sense as telling someone in the 75th percentile that they only got three quarters of the answer right. The numbers look similar, but they mean completely different things.

This becomes particularly problematic when you turn 90 into a target. A rep could improve substantially over several months while almost every call they make remains below the line. If your dashboard dutifully paints all of those calls red, you've built a target that's remarkably good at hiding improvement.

You could solve that by making CallCoach more generous. Shift all the scores upwards until strong calls get 90 and exceptional calls regularly get 100. Everyone's dashboard would look much nicer, but we haven't actually improved anybody's performance. We've mostly repainted the ruler, and we'd still need to decide where on the newly painted ruler good performance begins.

I'd rather keep the information in the scale and put the target somewhere sensible.

Where to draw the line

Around one call in four in our data scores 70 or better. That makes 70 an interesting place to start thinking about a target: it's achievable often enough that a rep can demonstrate it consistently, but high enough to distinguish stronger calls from ordinary ones.

I wouldn't blindly make 70 the CallCoach target either. Different organisations measure different behaviours, sell different things and have different distributions. I'd start with your own calls and ask where the middle sits, where the top quarter begins and what the best 10% look like. Then decide what level represents the performance you actually want your team to produce consistently.

That changes how you read improvement as well. A rep moving their typical performance from 60 to 67 hasn't gone from a bad percentage to a slightly less bad percentage. They've moved a meaningful distance through the distribution of actual calls. That's exactly the sort of change a quality score should make visible.

So there's nothing wrong with a scale that runs to 100, any more than there's something wrong with a Garmin sleep score that's almost impossible to max out. Just don't confuse the end of the scale with the place you're supposed to be.

CallCoach is how we score every call, on the ruler described above. Try it on your own calls and see where your middle call, your top quarter and your best 10% actually sit before you pick a target.


All CallCoach data in this post is aggregated, anonymised performance data from hundreds of thousands of calls across a variety of industries. We report the trends we spot; full statistical analysis is beyond the scope of a blog post.