7 spend data quality problems (and the fixes)

Last updated: 2026-08-14

Published 14 August 2026 · 6 min read

Every spend analysis runs into the same handful of data problems. None are hard to fix once you know to look, and all of them silently distort the answer if you do not.

1. Duplicate supplier names

The big one. Same vendor recorded under several spellings, so your top-supplier report understates concentration and your consolidation analysis misses real overlap. Fix by normalising strings and stripping legal-entity suffixes before anything else — full method here.

2. Part numbers inside descriptions

“BOOT WELLIE DANE GRN K680211-10”. The commodity is “boot”; the rest is a SKU that appears once in the entire dataset. Any grouping that treats the full string as the identity will produce one category per row. Strip tokens containing digits before you cluster or match.

3. Credit notes and reversals

Negative lines are legitimate but they wreck totals if you sum blindly — and they wreck averages badly. Decide explicitly whether to net them against the original transaction or exclude them, and say which in your notes.

4. Intercompany transactions

Charges between entities in the same group are not third-party spend. Leaving them in inflates addressable spend, sometimes dramatically, and the resulting saving target is unachievable because there is no external supplier to negotiate with.

5. Mixed currencies with no rate column

A file with GBP, EUR and USD amounts in one column and no currency indicator is unusable and often not obvious until the totals look strange. Check for a currency column early. If rates were applied at transaction date, keep that rate — retranslating everything at today's rate changes historic comparisons.

6. Blank or useless descriptions

Rows with an amount, a supplier and either nothing or “MISC” / “SUNDRY” / “ADJUSTMENT” in the description. These cannot be classified from the line alone. Two options: infer from the supplier's other transactions, or bucket them honestly as unclassified. Do not let a tool guess — a confident wrong category is worse than an admitted gap.

7. Freight, tax and surcharges as separate lines

Shipping and duty often appear as their own transactions. Classified naively, you end up with a large “freight” category that is really the delivery cost of everything else. Decide whether to allocate these back to the parent line or keep them separate — either is defensible, inconsistency is not.

A ten-minute pre-flight check

Before classifying anything, run these on the raw file:

  • Count distinct supplier names, then count again after normalising. A large gap means duplication.
  • Sum the amount column. Does it match what finance expects? If not, stop.
  • Count negative amounts. Any surprise there?
  • Count blank descriptions as a percentage of rows.
  • Check for more than one currency.
  • Check the date range is what you asked for.

Six checks, ten minutes, and they catch most of what would otherwise surface halfway through the analysis.

Further reading


← All posts