The claim, and the field under it
A memo going to the board put market coverage at 86 percent, and the recommendations for inbound and outbound were built on top of it.
The number came out of a market sizing model, and the model needed one thing from our CRM: a way to tell which companies belong in the addressable market. That is where it went wrong, and not at the arithmetic.
Why I used a substitute field
The fields the model was designed around did not exist in our CRM, and the revenue field that would have replaced them was empty across the whole portal.
So I used a tier field from an enrichment tool instead. It was the only field in the system that sorted companies by anything resembling size, it sat on about 44 percent of companies, and it had never been checked against anything. Under the circumstances that was a reasonable decision.
What I did not do is write the sentence that says so. The substitution never appeared next to the result, so by the time the figure reached the memo it read as a measurement.
11,503 accounts against 981
The model counted 11,503 accounts in the addressable market. A validated list, built separately from our own criteria, held 981.
| What was counted | Accounts |
|---|---|
| The model, resting on the unvalidated tier field | 11,503 |
| The validated list, built from our own criteria | 981 |
Roughly a factor of 12. The 86 percent had nothing underneath it, and it had already been repeated.
A coverage claim that big cannot be corrected downward. It has to be withdrawn.
From six months inside a B2B GTM organization with 36,000 accounts in its CRM.
The caveat that has to travel with the 981
The validated list is stricter than the market definition it was being compared against, so dividing one by the other produces a number that means nothing.
That matters because the 981 look like a clean replacement figure. 926 of them carried a discovery note, which is what makes them credible as a working list. The list also shows zero active customers, and that is a property of how the list was assembled rather than a conversion result. Anybody reading it as a funnel would draw the wrong conclusion twice.
So what the exercise produced was 981 accounts worth working, and the knowledge that the old number could not be repaired.
Why this is a disclosure problem rather than a data quality problem
Every individual step here was defensible, which is what makes it worth writing down.
The fields were missing, the substitute was the best available, the model was built correctly on top of it, and the arithmetic was right. The only missing piece was a line saying which field carried the segmentation and what its status was. One sentence would have made the 86 percent readable as a rough proxy.
A data cleanup would not have caught this. The tier field was populated on the companies it covered and internally consistent. What was wrong was what it was being used to mean, and that lives in the documentation.
What to do instead
- State four things before the first sum. The field name, its coverage across the object, whether it has been validated against anything, and the reason this field was chosen.
- Agree the substitute before you calculate. A proxy nobody objected to is fine. A proxy nobody was told about becomes a fact in the second document it appears in.
- Build one validated list and put it next to the model. It was a single query against the CRM's own criteria, and it moved the number by an order of magnitude.
- Send the caveat with the figure, every time. If the validated list is stricter than the market definition, say so in the same sentence, because the reader will otherwise divide one by the other.
- Prefer the smaller honest list to the larger claim. 981 accounts, 926 of them with a discovery note, is a list somebody can work. 11,503 is a slide.
The field coverage question underneath this is worth running on its own, and in the same CRM most of the schema turned out to be empty or nearly empty. The same discipline applied to closed deals produced a second surprise, because almost none of the closed deals could answer the question being asked of them.