AI spend classification: how it works and where it fails

Last updated: 2026-08-26

Published 26 August 2026 · 6 min read

AI spend classification uses language models to read a transaction description and assign it a procurement category, replacing the supplier-to-category mapping tables that spend analysis ran on for twenty years. It is genuinely better at the job. It is also oversold, and the failure modes are specific enough to be worth knowing before you buy anything.

How does AI spend classification work?

Three steps, in most implementations:

  1. Normalisation. Descriptions are cleaned — case standardised, part numbers and size codes stripped, supplier name variants collapsed. Unglamorous and responsible for a large share of the final accuracy.
  2. Grouping. Similar items are clustered so the model decides once per commodity rather than once per row. This is why a 10,000-row file does not cost 10,000 model calls, and why cost per row falls as datasets grow.
  3. Assignment. The model places each group in a category hierarchy, either one you supply or one generated from your data.

Why it beats rules and lookup tables

A mapping table keys on the supplier. That breaks the moment a supplier sells across categories — Amazon appears under IT, stationery, facilities and catering, and a supplier-level rule puts the whole vendor in one bucket.

A language model reads “NITRILE GLOVE BLUE PWDR FREE BX100” and recognises hand protection without anyone writing a rule for nitrile, or for the eleven abbreviations your ERP uses for gloves. That is the whole advantage: it generalises to descriptions nobody anticipated, which is exactly what the long tail consists of.

Where it still fails

  • Descriptions that are only part numbers.“K680211” carries no information. No model recovers meaning that was never in the data — a tool claiming 100% coverage on a file like that is guessing.
  • Internal jargon and brand names. If a product is only ever referred to by trade name, the model may categorise by the brand rather than the commodity.
  • Company-specific boundaries. Whether packaging is direct or indirect depends on your business, not on language. The model cannot know your convention unless you tell it.
  • Instability between runs. The one people miss. Re-running can produce a slightly different taxonomy, which makes year-on-year comparison impossible. Ask any vendor directly whether re-running the same file reproduces the same categories.

What good accuracy actually looks like

Treat any single headline percentage sceptically, because accuracy depends far more on your data than on the model. A more useful set of questions:

  • What share of rows were left uncategorised, and what were they?
  • How did the long tail do — not the top 100 suppliers?
  • Can you see a sample of your own data classified before paying?
  • Are corrections retained, or do they vanish on the next run?

A vendor comfortable answering all four is more convincing than one quoting 99% accuracy with no denominator.

Human review is not optional

Classification is a judgement task with real money attached. The right shape is machine-first, human-verified: let the model do the bulk, review a statistically meaningful sample, correct what is wrong, and keep the corrections. Anyone selling fully autonomous categorisation with no review step is selling you an unaudited number to put in front of a CFO.

Under the EU AI Act this kind of system is low risk — it categorises commercial transactions, it does not make decisions about people. The transparency obligations still apply, which in practice means you should be told when output is machine-generated. Any vendor cagey about that is a bad sign for unrelated reasons.

Further reading


← All posts