Skip to content
ArticleUpdated 4 min read

Are your lead scores honest? How to check in an afternoon

A lead score is honest if higher scores actually reply more. To check, group everyone you contacted by their predicted score band, compute the reply rate in each band, and see whether it rises with the score. Have at least 50 contacts per band before reading anything into the gaps; with fewer, the noise is larger than the differences.

The question nobody asks their scoring

Every lead tool gives you a score. Few tell you whether it works. A score of 85 means nothing on its own. It only means something if people scored 85 reply more than people scored 60, and people scored 60 reply more than people scored 30. That property is called calibration, and you can check it yourself with a spreadsheet.

Step 1: gather the data

You need, for everyone you contacted in a period:

  • The score they had when you contacted them (not a score recalculated later).
  • Whether they replied.
  • Ideally, whether the reply was positive.

Use one message, or one family of messages, for the period. If half the list got a brilliant message and half got a bad one, you are measuring the copy, not the score.

Step 2: group by score band

Four bands are enough. Fewer makes the check crude; more spreads your data too thin.

Here are two illustrative results for 300 contacts each. The first scoring is honest. The second is not.

Score bandContactedReplies (honest)Rate (honest)Replies (broken)Rate (broken)
0 to 494012.5%12.5%
50 to 696035.0%46.7%
70 to 84120108.3%65.0%
85 to 100801113.8%33.8%

In the honest case, the rate climbs with every band. In the broken case, the top band does worst. Something the scoring rewards is anti-correlated with replying: perhaps it loves big-company titles who never answer cold messages.

Notice that you contacted fewer low scorers. That is normal, since you were using the score to choose. But it means the bottom bands have the least data, which leads to step 3.

Step 3: check the noise before believing the gaps

Small samples lie. The margin of error on a reply rate depends on the number contacted. A rough 95% range is 1.96 × √(rate × (1 − rate) ÷ contacted).

Reply rateContactedMargin of error (about)
10%50± 8 points
10%200± 4 points
5%50± 6 points
5%200± 3 points

So with 50 people in a band at a 10% reply rate, the true rate could be anywhere from about 2% to 18%. Two bands at 8% and 14% with 50 people each are not reliably different. This is why the answer above says at least 50 per band, and why around 200 per band is where smaller differences become believable.

Decision rule: only act on a gap bigger than the margins of both bands. If you cannot see a clear gap, you do not have a scoring problem yet; you have a sample-size problem. Keep collecting.

Step 4: read the pattern

Four shapes come up again and again.

  1. Rising steadily. The scores are honest. Use the threshold with confidence, and consider raising it if the top band has plenty of volume.
  2. Flat. The score is not predicting anything. Either the profile describes the wrong people, or the message is so weak that nobody replies regardless.
  3. Rising, then falling at the top. The scoring over-rewards something, often seniority or company size. The very best-fitting on paper are the hardest to reach.
  4. Noisy, no clear shape. Not enough data. See step 3.

Step 5: fix what the data shows

  • Flat across bands: test the message first. Send the same group a better message; if replies rise in every band equally, the copy was the problem.
  • Top band underperforms: look at the top-band non-repliers. What do they share? Remove or soften that trait in your ideal customer profile.
  • Bottom band surprises you: look at who replied from it. They may be a segment your profile does not describe yet.
  • Everything checks out: move the threshold. If 70 to 84 replies nearly as well as 85 and up, contacting more of the 70s is a cheap way to grow volume.

Re-run the check every month or every 300 contacts, whichever comes first.

Why most tools do not show you this

Calibration is uncomfortable to publish. If a vendor's scores do not predict replies, a reply-rate-by-score chart shows it immediately. Asking a vendor for this number is one of the most useful questions in a demo: "Show me reply rate by predicted score for your customers, or for me after a month." The glossary entry on lead scoring covers the basics.

How Sluice shows it

Sluice scores each person 0 to 100 against your profile, plus whether they are worth approaching and whether their words show the problem now; 70 is the line for worth contacting. Insights shows reply rate by predicted score, alongside bounce rate and cost per lead worth contacting by source, so steps 2 to 4 are a chart rather than a spreadsheet. It does not do the statistics for you on small samples, so the margin table above still applies. Sluice is new, and we do not yet have large-volume results to publish. See pricing, and what a lead actually costs for the cost side of the same question.

Do this today

Export your last month of outreach with the scores at time of contact, build the four-band table, and add the margin of error beside each rate. If the top band is not clearly ahead, stop trusting the score until you know why.

Questions people ask

What does it mean for a lead score to be calibrated?
That the score means what it claims: people scored higher really do reply, meet or buy more often than people scored lower, in the order the scores predict.
How many leads do I need to test my scoring?
At least 50 contacted per score band to see large differences, and around 200 per band to trust smaller ones. At a 10% reply rate, 50 contacts gives a margin of error of roughly plus or minus 8 percentage points.
Should I measure replies or meetings?
Start with replies because there are more of them, so the numbers settle faster. Move to positive replies or meetings once you have the volume.
What if the scores are not predicting anything?
Check whether the message, not the score, is the problem; then look at which traits in your profile the high scorers share with the non-repliers, and remove or reweight them.

Try it on your own market

Sluice quotes the worst-case price before anything runs and charges only for lookups that found something, so finding out costs close to nothing.

Get started