We Lost Deals That Scored in the 90s. So I Backtested the Model.
Every scoring model is assumptions until you test it against reality. The fit score everyone trusted barely beat a coin flip while the strongest predictor was not weighted at all.
Forwarded to you? I am Heath. I build go-to-market systems and put AI to work in sales, the right way, then I write down exactly what I built, what broke, and what it moved. One story per week, receipts only. This one is about the scoring model everyone trusted until I tested it against the deals we actually won and lost.
| 1 | The score was decorating. Several accounts we lost had scored in the 90s on fit. A model that grades your losses as high as your wins is not scoring. |
| 2 | The flatterers carried the weight. The fit score and CRM presence both sat near a 1.1x lift, barely better than a coin flip, and nobody had checked. |
| 3 | The real predictors were behavioral and unweighted. How the buyer actually worked separated wins from losses, not the firmographics we had been scoring on. |
| 4 | The top signal was a 2.67x lift. It showed in 86% of won deals versus 31% of lost ones, and zero lost deals ever reached the top of the funnel. |
The reflex: trust the score because the high accounts look right
Every team that builds a scoring model reaches for the same move. You pick the traits that feel like a good customer, weight them on instinct, and trust the output because the accounts at the top look like the ones you want to win. The score feels like math, so you stop questioning it.
I did exactly that. We had a fit score, and everyone trusted it. Then I pulled the won and lost deals and measured what actually separated them. The fit score barely moved the needle. Several accounts we lost had scored in the 90s on fit. A model that scores your losses as highly as your wins is not scoring. It is decorating.
The block was never the math. It was that nobody had tested the assumptions. Every scoring model is a set of assumptions until you check it against reality, and ours assumed fit predicted wins. It had never been backtested. So it kept flattering the accounts that looked right and stayed silent on the ones that actually bought.
The AE feels this as "the score says call this account." The CSM feels it as "the health score is green so we are fine." The marketer feels it as "this segment fits, so it must convert." Same reflex every time: a number trusted because it feels right, never tested against what happened.
The reframe: backtest the score before you trust it
A scoring model is not truth. It is a hypothesis, and a hypothesis is worthless until you test it against outcomes. The move is not to argue about which traits matter. It is to measure how often each signal shows up in a win versus a loss, and keep only the ones that actually separate the two.
A model that scores your losses as high as your wins is not scoring. It is decorating.
Same move, other seats. An AE stares at a hot lead score and drills it down: has anyone checked that this score ever predicted a close, or does it just reward big logos. A CSM stares at a green health score and drills it down: is green measured against renewals that actually happened, or against logins. A marketer stares at a high-fit segment and drills it down: did that segment convert, or did it just match the firmographic filter. The number is never the thing to trust. The lift underneath it is.
How the best teams frame it
I am not the first to argue that a score is a hypothesis. The operators who have built scoring at scale mostly agree on where the leverage is, and it is not the model. It is testing the model against outcomes.
SOURCE
Kyle Poyar / OpenView, "Lead Scoring Models: Testing the Model"
What it argues. A scoring model is only as good as the validation behind it. Poyar walks through testing a model against real conversion data instead of trusting the point weights you assigned on instinct. The traits teams assume matter often do not survive contact with the outcomes.
My take. Agree, and this is the step everyone skips. Building the score is the easy part. Testing whether it predicted anything is the work, and most teams never do it because the model already feels right.
SOURCE
Kumo, "Lead Scoring Beyond Firmographics: From Company Size to Buying Signals"
What it argues. Firmographics tell you who could buy. Behavior tells you who will. Scoring on company size and industry alone leaves the strongest predictive signal, what the buyer actually does, on the table. Behavioral signals typically carry multiples more lift than firmographic ones.
My take. Extend. This is exactly what my backtest found. The firmographic fit score sat near a coin flip while the behavioral signal carried a 2.67x lift. The signal with the most predictive power was the one we were not weighting.
SOURCE
Federico Presicci, "Win/Loss Analysis: From Deal Data to Strategic Action"
What it argues. Most teams run on self-reported loss reasons that are biased and vague. The real value is in the deal data itself: measure what actually separated the deals you won from the ones you lost, and let the pattern, not the anecdote, drive the change.
My take. Agree hard. The reason my score was decorating is that nobody had run this. The deals held the answer the whole time. You just have to measure the wins against the losses instead of asking the rep why they think it slipped.
Even a curator has to concede when the field agrees: nobody who has actually built a scoring model thinks the model is the hard part. Testing it against real wins and losses is.
The method: Solve, Stack, Split
SOLVE THE CRUX
What is the real problem, framed as work and not a headcount?
The problem is not "we need a better score." It is "we never tested the one we have." So the first work is measurement: for every signal, how often does it show up in a won deal versus a lost one. That ratio is lift. For the AE it is the signal that precedes a close, for the CSM the one that precedes a renewal, for the marketer the one that precedes a real conversion, and in every seat the only way to know is to measure it against what actually happened.
STACK THE CONTEXT
What tech and signals turn raw deals into a tested model?
Not a shopping trip. The won and lost outcomes live in Salesforce. The behavioral signals live in Amplitude, sequence activation the loudest of them. Snowflake holds the deal-and-signal history the backtest runs across. Claude measures the lift per signal, wins versus losses, not vibes. The new weights ship into Deepline, the scoring pipeline the reps actually run on. The model gets tested before anyone trusts it again.
SPLIT · CUT THE DRAG
What low-judgment work goes to the system?
Pulling every deal, joining every signal, and counting how often each one shows up in a win versus a loss. Across the whole history, every signal, every deal. This never scaled by hand, which is why nobody had done it, and it is exactly the mechanical counting AI does perfectly.
SPLIT · KEEP THE JUDGMENT
What stays human?
Deciding what the lift means and what earns weight. The system says this signal carries a 2.67x lift and this one barely beats a coin flip. The operator decides which signals ship into the model, which combinations to trust, and which flatterers to retire. One owner on the model, a review of what the backtest showed, and only the earned weights go live.
The workflow: the board that runs it
Solve, Stack, Split is the shape. Here is the actual board, lane by lane: what the agents run, what stays human, and the tool at each step. Once it is wired, the backtest reruns on the latest deals without anyone rebuilding it from scratch.
What it moved
At a company I was at, a B2B SaaS with a scoring model everyone trusted, this is the exact build I ran. Pull the won and lost deals, measure lift per signal, keep only what earned its weight.
The model did not get more complex. It got tested. The signals that felt important and the signals that were important turned out to be different lists, and once the behavioral one carried its real weight, a Warm-and-above account won at 94% while the fit score everyone trusted sat there barely beating a coin flip.
Both of my receipts are new-business. Drop your own workflow in. A CSM runs the same backtest on renewals: which signals separated the accounts that renewed from the ones that churned, and does the health score actually predict it. A growth marketer runs it on conversion: which behaviors separated the segments that bought from the ones that ghosted, weighted by lift instead of by fit.
WHAT I LEARNED
1. A scoring model is a set of assumptions until you backtest it. Ours scored several losses in the 90s and nobody had checked.
2. Measure lift, not opinion. A signal that shows up as often in losses as in wins is noise, no matter how good it feels.
3. The strongest predictors are often behavioral and unweighted. The best one here carried a 2.67x lift and was not in the model.
4. Signals compound. One behavioral signal plus an ICP threshold stacked to a 3.56x lift, stronger than either one alone.
Run this one this week
Do not rebuild your scoring model. Pull your last fifty closed deals, wins and losses. Pick one signal you assume matters and count how often it shows up in each. If it shows up as much in your losses as your wins, it is a flatterer, and it should stop carrying weight. That one measurement is the backtest in miniature, and it tells you whether the whole model is worth trusting before you bet another quarter of pipeline on it.
Two builds that sit next to this one:
- Scoring the TAM — "You cannot prioritize a market by scoring one deal at a time. Put the entire TAM through one rubric, and make reps work the ranked list top-down." The backtest is how you prove that rubric earns its weight.
- The Expansion Score — "Check the score against your CSMs gut, because the divergences are where the money is." Same instinct: never trust a score you have not tested against what actually happened.
This is one build from the Build Log. Every week I take one sales or revenue problem, run it through the loop, and show the receipts. If someone forwarded this, the subscribe button is right below. Keep building. Heath.
You bought the signal. You never built the motion.
Everyone can capture intent now. The pipeline leaks in the gap between knowing and acting. What I got wrong, and what I am asking Adam Robinson on air.
The Seventh Analyst: The Agent That Reads the Other Six
Anyone can stand up six AI analysts. The one that makes the system compound is the seventh, the meta-analyst that reads the others' audit trails and turns every miss into a rule the next run enforces.
The Governance File: How a Pipeline Stops Repeating Its Worst Week
When you automate, your failures go silent. This build turned every past break into an enforced gate, so a fixed bug stays fixed and nothing broken ships confidently again.