Custom models for procurement spend classification
A custom model for spend classification is worth considering when supplier rules and your current procurement system confuse different types of purchases. Before training begins, agree on the category structure and check whether the available descriptions distinguish those categories. Choose a solution based on tests with unseen line items and the work required to review its results.
Syntalith
A procurement lead needs a report they can use to prepare a sourcing event or negotiate with a supplier. If one category captures everything purchased from that company, the report may overstate its role in a particular area. Before commissioning a model, identify the purchasing decision affected by those assignments. That gives the first test a focus and helps establish which errors matter.
One supplier, different purchases
Consider an illustrative example of a manufacturer. The same distributor supplies protective gloves and spare parts for production equipment. A rule assigning all purchases from that distributor to safety supplies produces a technically consistent report, but obscures maintenance spending. When the team plans a combined glove purchase across several plants, it is working from the wrong spending baseline.
A model that analyzes individual line items could separate those purchases if it receives enough information to identify them. The description “protective gloves” provides different evidence from “materials as ordered.” The latter may require a link to the purchase order. A more complex model alone cannot reliably fill in that missing context.
The pilot should therefore assess items with clear descriptions separately from those requiring additional documents. If improvement depends mainly on access to purchase orders, connecting the data will be a substantial part of the project. Purchasing model training alone would leave that problem unresolved.
Who decides where a purchase belongs?
A procurement taxonomy is a structured set of spend categories and subcategories. It should support the decisions the procurement team makes. A general ledger account represents a different accounting dimension; its code should not be assumed to identify the category used in supplier negotiations. A project can use both pieces of information while preserving their distinct meanings.
Before selecting technology, appoint someone to own the category structure and resolve disagreements. In the distributor example, procurement and maintenance need to agree on where spare parts belong. If each plant uses the same category name for a different scope, the combined report will still lack comparability after AI is introduced.
Oracle's guidance on business decisions before spend classification addresses stakeholder decisions about taxonomy and training data quality. For a buyer, this means checking whether historical assignments reflect the current category structure. Strong performance at reproducing old categories may have little value when the company is changing them.
What to compare before commissioning a model
Supplier rules deserve consideration where purchases from a supplier fall into a clearly defined category. They are easy to explain to the team. For a distributor serving several purchasing areas, however, they need additional information about each line item.
If your enterprise resource planning system, or ERP, includes spend classification, include that capability in the comparison. Check which category structures it supports and whether its results can feed existing reports. A custom application needs to justify the work required to maintain it and connect it to your systems.
A general-purpose model can be assessed on its ability to understand descriptions and assign them to agreed categories. If recurring errors remain that are specific to the company's purchases, a custom classifier or fine-tuning becomes a candidate for further testing. Fine-tuning adapts an existing model through additional training. It is neither a required step nor a guarantee of better results.
The comparison should also cover exception handling. A solution that correctly recognizes common purchases but sends most difficult items to analysts may still require substantial work. Understand that workload before deciding to deploy it.
How to evaluate the pilot
Ask for the methods to be compared on the same previously unseen line items. These should cover familiar suppliers as well as purchases from companies the model did not encounter during adaptation. This helps distinguish recognition of the purchase type from reliance on a familiar supplier name.
Break down overall performance by category and by the consequences of errors. Confusing two glove subcategories may matter differently from assigning machine parts to safety supplies. Also assess the value of misclassified spend: a few large line items can distort a category even when many small purchases are classified correctly.
Agree on what the system should do when a description does not support a category choice. Someone needs to review those items and receive the context required to make a decision. Oracle's configuration documentation covers prediction confidence and manual correction and approval. It provides a useful reference when assessing the team's review workload. A custom project still needs to test whether confidence indicators help identify errors in the company's own data.
Also establish when a result can enter a report. Readers should be able to distinguish a suggested category awaiting review from an approved assignment. Depending on the agreed pilot scope, deliverables could include a comparison report, an error review, and a recommendation for next steps with the relevant limitations.
Deployment scope and ongoing ownership
New suppliers and changes in what the company buys can alter the descriptions entering the classifier. Restructuring the categories is a separate event, such as splitting spare parts out of a broader maintenance category. Both situations call for reassessment, but only the second changes the definition of a correct answer. The proposal should identify who reports these changes and how the scope of updates will be agreed.
Syntalith builds AI applications and solutions using custom models, as well as automations and custom software. For spend classification, that allows the model, access to purchasing data, and correction handling to be considered together. The scope still needs to specify which elements belong in the pilot and which require a separate deployment. See Syntalith's pricing page for the current offer.
In your inquiry, describe the report you cannot currently trust and the purchasing decision that depends on it. Identify your system, the information available for each line item, and the person responsible for the categories. Bring examples of errors and explain how they are reviewed today. This will help establish whether the first engagement should compare classifiers, clarify the category structure, or address access to missing data.
Match a model to the task you need it to perform
Describe where your current AI falls short. We will compare model customization options, data requirements and the cost of running the resulting system.
Private LLMs and fine-tuning