Spend Data Classification: Create Usable Procurement Data

Procurement teams often have access to large volumes of financial and purchasing data. The problem is that the data is rarely ready for procurement analysis.

Supplier names may be inconsistent. Transactions may be recorded in different currencies. Descriptions may be incomplete. General-ledger accounts may not reflect the supply market, and expenditure may be spread across accounts-payable systems, purchase orders, payment cards, expense systems and local spreadsheets.

Before procurement can identify sourcing opportunities, supplier concentration, contract leakage or purchasing patterns, the underlying data must be collected, cleansed, normalised and classified.

This article explains how to turn fragmented expenditure records into usable procurement data—and how to ensure that the finished dataset supports real procurement decisions.

LHTS procurement framework

Primary role: Tactical buyer
Supporting role: Procurement management
Procurement process: Category management—spend analysis and opportunity identification
Learning level: Advanced
Related course: Spend Analysis Foundations

Quick answer: What is usable spend data?

Usable spend data is sufficiently complete, accurate, consistent, timely and traceable to support a defined procurement decision.

Creating usable spend data normally involves:

  1. Defining the procurement question and analysis scope.
  2. Collecting expenditure data from relevant systems.
  3. Reconciling the extracted data with Finance.
  4. Cleansing and standardising transaction fields.
  5. Normalising supplier names and corporate structures.
  6. Designing a procurement taxonomy.
  7. Classifying transactions.
  8. Validating the results.
  9. Turning the findings into procurement actions.
  10. Governing and refreshing the data.

Spend classification does not create savings by itself. It helps procurement identify where investigation and action are justified.

What makes spend data usable?

A dataset is not usable simply because it has been loaded into a dashboard.

Procurement should evaluate spend data against six practical tests.

Complete

Have all relevant data sources, legal entities, business units and transaction types been included?

Missing payment-card or expense data, for example, may understate expenditure in categories such as travel, professional services or low-value purchasing.

Accurate

Do the amounts, currencies, suppliers, dates and descriptions reflect the original transactions?

Errors in tax treatment, exchange rates or credit notes can materially distort category totals.

Consistent

Are supplier names, currencies, date formats, units of measure and category codes presented consistently?

The same supplier should not appear as several unrelated suppliers because of spelling differences or local naming conventions.

Timely

Is the information recent enough to support the intended decision?

Annual spend data may be sufficient for a category review, while a compliance dashboard may require monthly or even weekly updates.

Traceable

Can the analyst move from a dashboard figure back to the underlying transaction?

Traceability is essential when category managers, Finance or business stakeholders question a result.

Decision-ready

Can procurement use the information to decide what to investigate or do next?

A technically perfect dataset that does not support a sourcing, compliance, risk or category decision has limited practical value.

Spend analysis and spend-data preparation are different

Spend analysis is the use of expenditure data to answer procurement questions.

Examples include:

  • Which categories offer the strongest sourcing opportunities?
  • Where is expenditure fragmented across many suppliers?
  • How much spend is outside contracts?
  • Which suppliers create concentration risk?
  • Where are different business units purchasing similar requirements separately?
  • Which categories have low purchase-order compliance?

Spend-data preparation is the work required to make those answers reliable.

It includes:

  • Data extraction
  • Reconciliation
  • Cleansing
  • Supplier normalisation
  • Enrichment
  • Classification
  • Validation
  • Governance

A dashboard is therefore not the beginning of spend analysis. It is one of the outputs of a controlled data-preparation process.

Where usable spend data fits in the procurement process

Spend data supports several parts of tactical procurement and category management.

Opportunity identification

Spend analysis helps procurement decide where to allocate limited sourcing resources.

A category with high expenditure, fragmented suppliers and limited contract coverage may deserve further investigation.

Category planning

Category managers need a reliable view of suppliers, demand, business units, locations and transaction patterns before developing a category strategy.

Sourcing

Historical expenditure can help define the scope of an RFQ, estimate commercial value and identify relevant stakeholders.

Negotiation

Supplier-group spend, volume trends and purchasing patterns can strengthen negotiation preparation.

Contract and compliance management

Spend data can reveal off-contract purchasing, non-PO invoices, unauthorised suppliers and purchases made through the wrong channel.

Supplier management

Normalised supplier data can help procurement identify supplier concentration, parent-company exposure and the organisation’s total relationship with a supplier group.

Step 1: Define the procurement question and scope

A common mistake is to begin by extracting every available transaction.

The first step should be to define what the analysis needs to support.

Possible questions include:

  • Which categories should enter the sourcing pipeline?
  • Where can supplier consolidation be investigated?
  • How much expenditure is outside negotiated contracts?
  • Which suppliers are used by several business units?
  • Where does procurement have single-supplier exposure?
  • Which purchases are made without purchase orders?

The team should then define the scope.

Analysis period

Decide whether the analysis will cover:

  • A calendar year
  • A financial year
  • The last twelve months
  • Several years for trend analysis
  • A shorter period for operational compliance

Organisational scope

Specify which legal entities, countries, sites and business units are included.

Spend scope

Determine whether the analysis includes:

  • Direct materials
  • Indirect goods and services
  • Capital expenditure
  • Logistics
  • Employee expenses
  • Payment-card transactions
  • Non-PO invoices

Inclusion and exclusion rules

Not every financial payment is addressable procurement spend.

The team should define how it will treat:

  • Payroll
  • Taxes
  • Intercompany transactions
  • Grants and donations
  • Refunds
  • Credit notes
  • Supplier rebates
  • Inventory movements
  • Duties and freight
  • One-time payments

These rules should be documented before the analysis begins.

Otherwise, two analysts may produce different spend totals while both appear to be correct.

Step 2: Map and extract the data sources

Procurement spend data may be distributed across several systems.

Common sources include:

  • Accounts-payable data
  • Purchase-order data
  • ERP systems
  • Procurement platforms
  • Payment-card systems
  • Employee-expense systems
  • Contract databases
  • Supplier master records
  • Inventory or production systems
  • Local spreadsheets

The extraction should preserve transaction-level detail wherever possible.

Aggregated account totals may be useful for reconciliation, but they are normally insufficient for detailed procurement classification.

Minimum useful spend-data fields

The required fields depend on the business question, but a practical dataset may include:

Data fieldProcurement purpose
Supplier IDLinks transactions to the supplier master
Supplier nameSupports supplier analysis and normalisation
Invoice or transaction numberEnables traceability and duplicate checks
Transaction dateSupports trends and period analysis
Line descriptionProvides evidence for category classification
AmountShows transaction and category value
CurrencyAllows values to be normalised
Legal entityIdentifies the purchasing organisation
Business unit or cost centreIdentifies the demand owner
General-ledger accountSupports reconciliation and classification
Purchase-order numberSupports PO-compliance analysis
Contract referenceSupports contract-compliance analysis
Quantity and unitSupports price and demand analysis
Category codeEnables category aggregation
Country or siteSupports geographic and risk analysis

Not every organisation will have every field.

The objective is to determine which fields are necessary for the intended procurement decision and which data-quality gaps must be acknowledged.

Step 3: Reconcile the dataset

Before cleansing and classification, the extracted data should be reconciled with an agreed financial control total.

Procurement and Finance should agree:

  • The reporting period
  • Included legal entities
  • Tax treatment
  • Currency conversion rules
  • Treatment of credit notes
  • Treatment of intercompany transactions
  • Excluded payment types
  • The acceptable reconciliation variance

Reconciliation confirms that the dataset is sufficiently complete.

Without this control, procurement may create a detailed and attractive dashboard based on only part of the organisation’s expenditure.

The reconciliation result should be documented and repeated whenever the dataset is refreshed.

Step 4: Clean and standardise the transactions

Raw transaction data normally contains formatting and quality problems.

Typical cleansing activities include:

  • Standardising date formats
  • Converting currencies
  • Removing unsupported symbols
  • Separating tax from net expenditure
  • Correcting negative and positive amounts
  • Matching credit notes to invoices
  • Removing duplicate transactions
  • Standardising units of measure
  • Cleaning incomplete descriptions
  • Identifying cancelled or reversed transactions

Each transformation should be controlled and traceable.

The team should be able to explain how the original value became the value used in the analysis.

Currency normalisation

Where several currencies are involved, values must be converted into an agreed reporting currency.

The organisation should define:

  • The exchange-rate source
  • Whether transaction-date, monthly average or period-end rates are used
  • How currency effects are treated in trend analysis
  • Whether original-currency values are retained

Retaining both the original and converted amounts makes later validation easier.

Step 5: Normalise supplier information

Supplier normalisation means connecting different supplier records to the correct supplier entities and groups.

A single supplier may appear under several names because of:

  • Spelling differences
  • Abbreviations
  • Local subsidiaries
  • Trading names
  • Remit-to entities
  • Previous company names
  • Different tax registrations
  • Duplicate supplier-master records

For example, transactions recorded against “ABC Industrial AB,” “ABC Ind.,” and “ABC Group Sweden” may or may not belong to the same legal entity.

The relationship must be verified rather than assumed.

Important supplier levels

A useful supplier structure may distinguish between:

  • Supplier master record
  • Legal entity
  • Trading name
  • Local subsidiary
  • Production site
  • Ultimate parent
  • Corporate supplier group

Procurement may need more than one view.

Contracting-entity view

This shows which legal entity holds the contract and receives payment.

It is important for contract ownership, legal obligations and operational supplier management.

Supplier-group view

This consolidates related subsidiaries under an ultimate parent.

It can help procurement understand:

  • Total negotiation leverage
  • Supplier concentration
  • Corporate-group exposure
  • Cross-business relationships

Supplier normalisation should retain the underlying legal-entity detail. Consolidating everything into a parent-company view can hide local risks and contractual differences.

Step 6: Enrich the data

Internal transaction records may not contain all the information required for procurement analysis.

Useful enrichment fields may include:

  • Supplier country
  • Ultimate parent company
  • Industry
  • Supplier size
  • Risk rating
  • Diversity status
  • Sustainability information
  • Contract status
  • Preferred-supplier status
  • Category owner
  • Payment terms

Enrichment can come from internal master data, contract systems or approved external sources.

The source and date of each enrichment field should be recorded. External information may become outdated and should not be treated as permanently accurate.

Step 7: Design a procurement taxonomy

A procurement taxonomy is a structured way of grouping expenditure into categories that procurement can manage.

General-ledger accounts are rarely sufficient because they are designed for financial reporting rather than supply-market analysis.

For example, one GL account called “Consulting” may contain:

  • Management consulting
  • IT consulting
  • Engineering services
  • Legal advice
  • Recruitment support

These services may require different category strategies, suppliers and commercial approaches.

Principles for taxonomy design

A practical taxonomy should be:

  • Relevant to procurement decisions
  • Understandable to category owners
  • Detailed enough to support action
  • Stable enough for trend analysis
  • Flexible enough to handle new requirements
  • Mappable to financial and external classifications where needed

A hierarchy might include:

  • Level 1: Broad category
  • Level 2: Category
  • Level 3: Subcategory
  • Level 4: Commodity or service type

The number of levels should reflect business needs. More detail is not automatically better.

A taxonomy that is too broad hides opportunities. A taxonomy that is too detailed becomes difficult to maintain and may produce categories with no clear owner or strategic relevance.

Internal and external taxonomies

Standard classification structures can support benchmarking and external data exchange.

However, an organisation may still need an internal procurement taxonomy that reflects:

  • Its category-management structure
  • Its products and services
  • Its supply markets
  • Its sourcing responsibilities
  • Its risk priorities

External and internal codes can be mapped rather than forcing every requirement into one structure.

Step 8: Classify the transactions

Classification assigns each transaction to the most appropriate category.

Several methods can be used.

Rule-based classification

Rules use known fields and keywords.

Examples include:

  • Supplier-to-category mappings
  • GL-account mappings
  • Material groups
  • Product codes
  • Description keywords

Rules are transparent and useful for recurring transactions, but they require maintenance and may struggle with ambiguous descriptions.

Machine-learning classification

Machine-learning models can identify patterns in descriptions, supplier histories and other transaction fields.

They may improve speed and coverage where transaction volumes are high.

However, their output must still be validated. A confidence score is not proof that a classification is correct.

Human classification

Category specialists may need to review:

  • High-value transactions
  • Low-confidence classifications
  • New suppliers
  • Ambiguous descriptions
  • Large “Other” categories
  • Transactions with conflicting evidence

A hybrid approach is often the most practical:

  1. Apply rules to known transactions.
  2. Use automated models for remaining records.
  3. Prioritise exceptions by value and confidence.
  4. Ask category owners to validate material items.
  5. Feed approved corrections into future classifications.

Step 9: Validate the result

Validation should assess more than the percentage of transaction lines that received a category.

A dataset can appear highly classified while still misclassifying the transactions that matter most.

Validate by transaction count

This shows how many records have been classified.

It can help identify operational data-quality issues and large numbers of unclassified lines.

Validate by spend value

This shows how much expenditure has been classified correctly.

It gives greater weight to high-value transactions that may materially affect a category strategy.

Validate by category owner

Category managers should review whether the results reflect their understanding of suppliers and demand.

Unexpected findings may be genuine opportunities, but they may also indicate data or classification errors.

Review common exception areas

Pay particular attention to:

  • High-value unclassified spend
  • Large “Other” categories
  • Suppliers allocated to several unrelated categories
  • One-time suppliers
  • High-value manual overrides
  • Transactions with poor descriptions
  • Unusual price or volume movements

There is no universal accuracy target suitable for every organisation.

The required quality depends on:

  • The decision being supported
  • The value at risk
  • Category complexity
  • Available evidence
  • The cost of an incorrect conclusion

A strategic sourcing decision involving a critical category requires stronger validation than a high-level opportunity scan.

Step 10: Turn spend insight into procurement action

Spend analysis should end with a prioritised procurement response—not only a dashboard.

Spend insightPotential procurement action
Fragmented expenditure across many suppliersInvestigate consolidation or a sourcing event
High off-contract spendReview contracts, channels and stakeholder compliance
High supplier-group concentrationConduct a risk review and assess alternatives
Similar items bought at different pricesReview specifications, demand and commercial terms
Large unclassified or “Other” categoryImprove descriptions, taxonomy or classification rules
Low purchase-order coverageInvestigate the P2P process and buying channels
Many small transactions across many suppliersConsider a tail-spend solution
Different business units buying the same serviceAssess aggregation and coordinated sourcing
Rapid category-spend increaseValidate demand, price movements and scope changes

The finding is the beginning of the procurement investigation.

Before launching a sourcing initiative, procurement should confirm:

  • Stakeholder requirements
  • Demand forecasts
  • Contract status
  • Supply-market conditions
  • Switching costs
  • Operational risk
  • Resource availability
  • Realistic implementation timing

Spend data helps identify where value may exist. It does not guarantee that an opportunity can or should be implemented.

Practical example: From fragmented data to a sourcing opportunity

A category manager wants to understand facilities-services expenditure across several sites.

The initial extraction contains 14,000 transaction lines and more than 400 supplier names.

The data comes from accounts payable, purchase orders and payment cards. Supplier names are inconsistent, some invoices lack purchase-order references and several descriptions contain only general wording such as “service work.”

The procurement and Finance teams agree the period, scope and exclusions. The extracted values are reconciled with the financial totals.

The data is then cleansed:

  • Currency and date formats are standardised.
  • Credit notes are matched.
  • Duplicate transactions are removed.
  • Supplier records are linked to verified legal entities.
  • Related subsidiaries are connected to parent groups.

The transactions are classified into:

  • Cleaning
  • Security
  • Maintenance
  • Utilities
  • Waste management
  • Other facilities services

The analysis shows that maintenance expenditure is distributed across several local suppliers. Some sites use contracts, while others buy similar services through repeated one-time orders.

Procurement does not immediately conclude that all suppliers should be consolidated.

Instead, the category manager:

  1. Validates the data with local site managers.
  2. Reviews service specifications and response-time requirements.
  3. Checks existing contracts.
  4. Assesses the regional supply market.
  5. Compares the risks of local and consolidated delivery models.
  6. Decides whether a regional sourcing event is appropriate.

The spend analysis creates a structured opportunity for investigation. Procurement judgment determines the next step.

How to measure spend-data quality

A spend-data process should have measurable quality indicators.

Useful measures include:

Spend coverage

The percentage of in-scope expenditure included in the dataset.

Reconciliation variance

The difference between the spend dataset and the agreed financial control total.

Classified spend by value

The percentage of expenditure assigned to an approved category.

Unclassified spend by value

The value that still requires investigation.

“Other” category percentage

A growing “Other” category may indicate weak taxonomy design or incomplete descriptions.

Supplier-normalisation rate

The proportion of expenditure linked to verified legal entities and supplier groups.

Classification accuracy

The percentage of reviewed transactions assigned to the correct category.

This should be measured by both transaction count and spend value.

Data freshness

The time since the latest transaction and supplier information were loaded.

Category-owner acceptance

The extent to which responsible category managers confirm that the data is sufficiently reliable for their work.

Common mistakes when creating usable spend data

Starting without a procurement question

Collecting every available field can create a large project without a clear business outcome.

Begin with the decision the analysis must support.

Assuming accounts-payable data is complete

Accounts payable may not contain payment cards, expenses or some local transactions.

Map all relevant sources before assessing total spend.

Using GL accounts as the procurement taxonomy

Financial accounts are designed for accounting control. They may not represent how procurement manages supply markets.

Merging suppliers only by name

Similar names do not always represent the same legal entity. Different names may also belong to the same corporate group.

Supplier relationships must be verified.

Creating too many categories

An overly detailed taxonomy becomes difficult to maintain and may not support meaningful category ownership.

Measuring classification only by line count

Correctly classifying thousands of low-value transactions does not compensate for misclassifying a major supplier.

Measure quality by both line count and spend value.

Treating AI output as automatically correct

Automated classification can improve efficiency, but material results require human review and source traceability.

Building dashboards without action owners

A dashboard has limited value when nobody is responsible for investigating findings and implementing agreed actions.

Assuming spend analysis automatically creates savings

Spend analysis identifies patterns and opportunities. Value is created only when procurement validates the opportunity and successfully implements an appropriate action.

How this connects to the tactical buyer role

The tactical buyer uses spend data to understand demand, suppliers and sourcing opportunities.

Usable spend data can help the tactical buyer:

  • Build a sourcing pipeline
  • Prepare category analyses
  • Define RFQ scope
  • Identify fragmented expenditure
  • Understand supplier relationships
  • Detect contract leakage
  • Prepare negotiations
  • Prioritise stakeholder discussions

The buyer must be able to question the data rather than simply accept the dashboard.

Important questions include:

  • What is included?
  • What is excluded?
  • How were suppliers normalised?
  • How was the category assigned?
  • Can the result be traced to the original transaction?
  • What procurement action could follow?

How this connects to procurement management

Procurement management is responsible for establishing a sustainable spend-data capability.

Management responsibilities may include:

  • Defining data ownership
  • Agreeing governance with Finance and IT
  • Selecting technology
  • Approving the procurement taxonomy
  • Assigning category ownership
  • Setting data-quality expectations
  • Prioritising analytical resources
  • Measuring whether insights lead to action
  • Ensuring that classification changes are controlled

The objective is not to create one perfect annual report.

The objective is to create a repeatable process that provides procurement with sufficiently reliable information when decisions need to be made.

Technology for spend-data preparation and analysis

Technology can accelerate extraction, cleansing, supplier normalisation, classification and reporting.

Solutions may include:

  • Data-integration platforms
  • Data-preparation tools
  • Supplier-master and enrichment services
  • Classification engines
  • Specialist spend-analysis platforms
  • Business-intelligence tools

Before selecting a solution, procurement should evaluate:

  • Which source systems can be connected?
  • Can line-level transaction detail be retained?
  • How are supplier entities identified?
  • Can the organisation maintain its own taxonomy?
  • How are confidence scores presented?
  • Can users override classifications?
  • Is an audit trail retained?
  • Can totals be reconciled with Finance?
  • How are corrections used in later classifications?
  • Can data and classification rules be exported when the contract ends?
  • Can category managers use the output without specialist technical support?

The best technology is not necessarily the platform with the largest number of features.

It is the solution that supports the organisation’s defined procurement decisions, data environment and governance model.

Learn more with the Spend Analysis Foundations course

Creating usable spend data requires both analytical methods and procurement judgment.

The Learn How to Source Spend Analysis Foundations course explains the basic spend-analysis process and how it supports category management and tactical procurement.

The course provides the foundation for:

  • Understanding procurement spend analysis
  • Collecting and structuring expenditure data
  • Identifying patterns and opportunities
  • Connecting analysis to category management
  • Supporting fact-based procurement decisions

Use the course to establish the basic methodology, and then apply the more advanced cleansing, classification and governance practices described in this article.

Frequently asked questions

What is spend-data classification?

Spend-data classification is the process of assigning procurement transactions to relevant categories so that expenditure can be analysed by supply market, supplier, business unit or purchasing purpose.

What is the difference between spend analysis and spend classification?

Spend classification organises transactions into categories. Spend analysis uses the resulting data to answer procurement questions and identify potential actions.

What data is required for spend analysis?

The required fields depend on the objective. Common fields include supplier, amount, currency, date, description, business unit, legal entity, GL account, purchase-order number and category.

Why are GL accounts not enough for procurement analysis?

GL accounts are designed for financial reporting. They may combine purchases that belong to different supply markets or sourcing strategies.

What is supplier normalisation?

Supplier normalisation connects variations of supplier names and records to verified legal entities, subsidiaries and corporate groups.

What is a procurement taxonomy?

A procurement taxonomy is a structured hierarchy used to group expenditure into categories that procurement can manage.

Can AI classify procurement spend automatically?

AI can classify large volumes of transactions and prioritise exceptions. Human review is still required for high-value, ambiguous or low-confidence transactions.

How accurate should spend classification be?

There is no universal target. The required accuracy depends on the decision, category complexity and materiality. Quality should be measured by both transaction count and spend value.

How often should spend data be refreshed?

The refresh cycle should match the intended use. Category reviews may use quarterly or annual data, while operational compliance monitoring may require more frequent updates.

Does spend analysis create savings?

Spend analysis identifies potential opportunities. Savings or other value are created only when procurement validates and implements an appropriate action.

Conclusion

Usable spend data is not created by loading financial transactions into a dashboard.

It requires a controlled process that begins with a procurement question and continues through extraction, reconciliation, cleansing, supplier normalisation, classification and validation.

The finished dataset must be complete enough, accurate enough and traceable enough to support the intended decision.

Most importantly, procurement must connect each insight to an owner and a possible action.

A practical first step is to select one procurement question, define the required scope and test whether every reported figure can be traced back to the original transactions.

That is the foundation of reliable spend analysis.

Spend data development
Spend data development