Forwarded to you? I am Heath. I build go-to-market systems and put AI to work in sales, the right way, then I write down exactly what I built, what broke, and what it moved. One story per week, receipts only. This one is about the scoring model everyone trusted until I tested it against the deals we actually won and lost.

TL;DR · THE GIST · 30 SECONDS
1The score was decorating. Several accounts we lost had scored in the 90s on fit. A model that grades your losses as high as your wins is not scoring.
2The flatterers carried the weight. The fit score and CRM presence both sat near a 1.1x lift, barely better than a coin flip, and nobody had checked.
3The real predictors were behavioral and unweighted. How the buyer actually worked separated wins from losses, not the firmographics we had been scoring on.
4The top signal was a 2.67x lift. It showed in 86% of won deals versus 31% of lost ones, and zero lost deals ever reached the top of the funnel.

The reflex: trust the score because the high accounts look right

Every team that builds a scoring model reaches for the same move. You pick the traits that feel like a good customer, weight them on instinct, and trust the output because the accounts at the top look like the ones you want to win. The score feels like math, so you stop questioning it.

I did exactly that. We had a fit score, and everyone trusted it. Then I pulled the won and lost deals and measured what actually separated them. The fit score barely moved the needle. Several accounts we lost had scored in the 90s on fit. A model that scores your losses as highly as your wins is not scoring. It is decorating.

The block was never the math. It was that nobody had tested the assumptions. Every scoring model is a set of assumptions until you check it against reality, and ours assumed fit predicted wins. It had never been backtested. So it kept flattering the accounts that looked right and stayed silent on the ones that actually bought.

The AE feels this as "the score says call this account." The CSM feels it as "the health score is green so we are fine." The marketer feels it as "this segment fits, so it must convert." Same reflex every time: a number trusted because it feels right, never tested against what happened.

The reframe: backtest the score before you trust it

A scoring model is not truth. It is a hypothesis, and a hypothesis is worthless until you test it against outcomes. The move is not to argue about which traits matter. It is to measure how often each signal shows up in a win versus a loss, and keep only the ones that actually separate the two.

DRILL-DOWN · FROM A NUMBER YOU TRUST TO ONE YOU TESTED
LOUD
A 90-plus fit score Looks like a buyer. Several of the accounts we lost scored right here.
NARROWER
The signals that feel like they separate wins from losses Warmer. But feel is not lift. You have to measure it against the deals.
THE CRUX
One behavioral signal that shows in 86% of wins and 31% of losses, a 2.67x lift, sitting in the data with no weight on it That is the signal worth trusting, and it was not in the model.
A model that scores your losses as high as your wins is not scoring. It is decorating.

Same move, other seats. An AE stares at a hot lead score and drills it down: has anyone checked that this score ever predicted a close, or does it just reward big logos. A CSM stares at a green health score and drills it down: is green measured against renewals that actually happened, or against logins. A marketer stares at a high-fit segment and drills it down: did that segment convert, or did it just match the firmographic filter. The number is never the thing to trust. The lift underneath it is.

How the best teams frame it

I am not the first to argue that a score is a hypothesis. The operators who have built scoring at scale mostly agree on where the leverage is, and it is not the model. It is testing the model against outcomes.

SOURCE

Kyle Poyar / OpenView, "Lead Scoring Models: Testing the Model"

What it argues. A scoring model is only as good as the validation behind it. Poyar walks through testing a model against real conversion data instead of trusting the point weights you assigned on instinct. The traits teams assume matter often do not survive contact with the outcomes.

My take. Agree, and this is the step everyone skips. Building the score is the easy part. Testing whether it predicted anything is the work, and most teams never do it because the model already feels right.

SOURCE

Kumo, "Lead Scoring Beyond Firmographics: From Company Size to Buying Signals"

What it argues. Firmographics tell you who could buy. Behavior tells you who will. Scoring on company size and industry alone leaves the strongest predictive signal, what the buyer actually does, on the table. Behavioral signals typically carry multiples more lift than firmographic ones.

My take. Extend. This is exactly what my backtest found. The firmographic fit score sat near a coin flip while the behavioral signal carried a 2.67x lift. The signal with the most predictive power was the one we were not weighting.

SOURCE

Federico Presicci, "Win/Loss Analysis: From Deal Data to Strategic Action"

What it argues. Most teams run on self-reported loss reasons that are biased and vague. The real value is in the deal data itself: measure what actually separated the deals you won from the ones you lost, and let the pattern, not the anecdote, drive the change.

My take. Agree hard. The reason my score was decorating is that nobody had run this. The deals held the answer the whole time. You just have to measure the wins against the losses instead of asking the rep why they think it slipped.

Even a curator has to concede when the field agrees: nobody who has actually built a scoring model thinks the model is the hard part. Testing it against real wins and losses is.

The method: Solve, Stack, Split

SOLVE THE CRUX

What is the real problem, framed as work and not a headcount?

The problem is not "we need a better score." It is "we never tested the one we have." So the first work is measurement: for every signal, how often does it show up in a won deal versus a lost one. That ratio is lift. For the AE it is the signal that precedes a close, for the CSM the one that precedes a renewal, for the marketer the one that precedes a real conversion, and in every seat the only way to know is to measure it against what actually happened.

STACK THE CONTEXT

What tech and signals turn raw deals into a tested model?

Not a shopping trip. The won and lost outcomes live in Salesforce. The behavioral signals live in Amplitude, sequence activation the loudest of them. Snowflake holds the deal-and-signal history the backtest runs across. Claude measures the lift per signal, wins versus losses, not vibes. The new weights ship into Deepline, the scoring pipeline the reps actually run on. The model gets tested before anyone trusts it again.

SPLIT · CUT THE DRAG

What low-judgment work goes to the system?

Pulling every deal, joining every signal, and counting how often each one shows up in a win versus a loss. Across the whole history, every signal, every deal. This never scaled by hand, which is why nobody had done it, and it is exactly the mechanical counting AI does perfectly.

SPLIT · KEEP THE JUDGMENT

What stays human?

Deciding what the lift means and what earns weight. The system says this signal carries a 2.67x lift and this one barely beats a coin flip. The operator decides which signals ship into the model, which combinations to trust, and which flatterers to retire. One owner on the model, a review of what the backtest showed, and only the earned weights go live.

The workflow: the board that runs it

Solve, Stack, Split is the shape. Here is the actual board, lane by lane: what the agents run, what stays human, and the tool at each step. Once it is wired, the backtest reruns on the latest deals without anyone rebuilding it from scratch.

THE WIN/LOSS BACKTEST · WORKFLOW
AI MEASURES THE LIFT · THE OPERATOR DECIDES WHAT EARNS WEIGHT
01 · PULL THE TRUTH
AI Salesforce
Pull won and lost deals
The outcomes the model gets tested against, wins and losses side by side.
AI Amplitude
Pull the behavioral signals
How the buyer actually worked, sequence activation the loudest.
02 · MEASURE LIFT
AI
Measure lift per signal ClaudeHow often each signal shows in wins versus losses. Lift, not vibes.
AI Snowflake
Run it on the deal history
The deal-and-signal history the backtest ran across.
03 · DECIDE THE WEIGHTS
HUMAN The operator
Keep what compounds, retire the flatterers
One behavioral signal plus an ICP threshold stacked to 3.56x; the fit score barely beat a coin flip, so it lost its weight.
BOTH Deepline
Reweight the model
The new weights ship into the scoring pipeline the reps run on.
THE SPLIT
AI measures lift across every signal. The operator decides what earns weight. The pipeline reweights.WHAT IT SHIPSPer-signal lift table · wins vs losses, measured The compounding stack · signals that beat either alone A reweighted model · weight only where it is earned

What it moved

At a company I was at, a B2B SaaS with a scoring model everyone trusted, this is the exact build I ran. Pull the won and lost deals, measure lift per signal, keep only what earned its weight.

THE NUMBERS
86% vs 31%
the top signal in won deals versus lost ones, a 2.67x lift
94% vs 62%
win rate of Warm-and-above accounts versus the base rate, a 1.5x lift
0
lost deals that ever reached Hot or Strike, so the top of the funnel is clean

The model did not get more complex. It got tested. The signals that felt important and the signals that were important turned out to be different lists, and once the behavioral one carried its real weight, a Warm-and-above account won at 94% while the fit score everyone trusted sat there barely beating a coin flip.

Both of my receipts are new-business. Drop your own workflow in. A CSM runs the same backtest on renewals: which signals separated the accounts that renewed from the ones that churned, and does the health score actually predict it. A growth marketer runs it on conversion: which behaviors separated the segments that bought from the ones that ghosted, weighted by lift instead of by fit.

WHAT I LEARNED

1. A scoring model is a set of assumptions until you backtest it. Ours scored several losses in the 90s and nobody had checked.

2. Measure lift, not opinion. A signal that shows up as often in losses as in wins is noise, no matter how good it feels.

3. The strongest predictors are often behavioral and unweighted. The best one here carried a 2.67x lift and was not in the model.

4. Signals compound. One behavioral signal plus an ICP threshold stacked to a 3.56x lift, stronger than either one alone.

Run this one this week

Do not rebuild your scoring model. Pull your last fifty closed deals, wins and losses. Pick one signal you assume matters and count how often it shows up in each. If it shows up as much in your losses as your wins, it is a flatterer, and it should stop carrying weight. That one measurement is the backtest in miniature, and it tells you whether the whole model is worth trusting before you bet another quarter of pipeline on it.

Two builds that sit next to this one:

  • Scoring the TAM"You cannot prioritize a market by scoring one deal at a time. Put the entire TAM through one rubric, and make reps work the ranked list top-down." The backtest is how you prove that rubric earns its weight.
  • The Expansion Score"Check the score against your CSMs gut, because the divergences are where the money is." Same instinct: never trust a score you have not tested against what actually happened.
THE WIN/LOSS BACKTEST · BUILD LOG
Test your scoring model against the deals you actually won and lost.
YOU LEAVE WITH
A per-signal lift table measured across wins and losses, the compounding stack of signals that beat either alone, and a reweighted model that carries weight only where it is earned.
RUNS ON   Salesforce · Amplitude · Claude · Snowflake · Deepline
PROVEN · a 2.67x lift on the top signal, 86% of wins vs 31% of losses
Read the full build, with the workflow board

This is one build from the Build Log. Every week I take one sales or revenue problem, run it through the loop, and show the receipts. If someone forwarded this, the subscribe button is right below. Keep building. Heath.