Skip to content
← Back to blog
fine-tuningArticle

Fine-Tuning for Freight Document Extraction

If a system repeatedly assigns readable freight document details to the wrong fields, fine-tuning may be worth evaluating. First identify whether the cause is image quality, text recognition, field meaning, or missing information, then test any improvement on documents reserved for evaluation.

Author

Syntalith

Published Updated 5 min read

Start by identifying the recurring error

Fine-tuning adapts an existing model to a task using examples with correct answers approved by the company. The purchase question is whether this can reduce a recurring type of mistake enough to outweigh the cost of checking results and correcting records by hand.

Consider a readable bill of lading from a carrier whose layout differs from the others. It lists the carrier and its address, the consignee, and the delivery location. The system reads each line correctly but puts the carrier's address in the delivery-location field. The text recognition worked; the system assigned a readable value to the wrong field. That creates a wrong location in the record, so an employee must reopen the document and correct it. Another part of the workflow may then act on the misleading address.

The same apparent mistake can start elsewhere. A blurry or poorly lit image can make the source unreadable, so the next step may be to request a better scan. If OCR, or optical character recognition, changes characters or drops part of an address, the text-recognition layer needs attention. When the words are clear but land in the wrong field, the problem is interpretation and field mapping. If the document does not state a delivery location, the system has no evidence to supply one. Each cause calls for a different fix.

What a useful result should preserve

For each extracted value, a reviewer should be able to trace it to its source field on the document. When the document does not contain supporting information, the system should show that gap. An operations employee can check the assignment, while the person responsible for the data can determine whether the error came from text recognition, field mapping, or the source itself. A document extraction result does not establish that delivery occurred; that needs separate evidence.

Before evaluation, the process owner should decide which corrections an operator can approve and which need review by someone familiar with the document. This gives evaluators a shared basis for resolving ambiguous cases.

Assess fields that matter to the process separately, such as delivery location and consignee details. The Google Document AI evaluation guide describes comparing extracted entities with annotations and reporting precision and recall by label. Those measures help show which fields are assigned incorrectly and which are missed.

Precision is the share of values assigned to a field that are correct. Recall is the share of expected values that the system found. Read both alongside the number of cases sent for manual review or correction. If delivery location often returns to an operator, a strong average across easier fields will not settle the purchase decision. A stricter confidence threshold may catch more errors while adding cases to the review queue.

When model adaptation may be worth considering

First, group recurring errors by likely cause. If poor images are responsible, improve the input and make it clear when an operator should request a better scan. If the text is readable and the field rules are clear, updating the schema, validation, or settings of an available extractor may be enough. An off-the-shelf extractor may also handle the document without a custom model. Google's Document AI training overview describes multiple custom extractor types and training approaches. A custom extractor in that product can refer to different mechanisms, so the term alone does not mean that an LLM was fine-tuned.

Fine-tuning is worth evaluating when field-assignment errors recur across documents despite readable inputs and recognized text, and the company has enough examples with approved answers to assess the change. The examples should cover layouts where the current system fails and support a credible comparison. A sparse set of examples or errors driven mainly by poor images does not yet justify buying fine-tuning.

How to evaluate the purchase

Keep development examples separate from evaluation documents the model has not seen. The buying team supplies the expected answers and agrees which carriers, layouts, and date variations the evaluation should cover. Include at least one layout outside the development material to see how the system handles an unfamiliar document. Review field-level quality together with the number and type of cases that still require a person.

If the system still assigns a carrier's address as the delivery location on an unfamiliar layout, that result does not support accepting the adaptation for this field. The next step may be to use a simpler rule, revise the schema, or send the field to an operator for review.

Syntalith can assess the error types, compare the current method with an available extractor, and advise whether model adaptation merits a further test. The scope needs to be agreed for the process: which documents and fields are in the integration, what may be written automatically, and when an operator checks the result. This assessment brings together the operator who corrects the record and the person who maintains its schema and system connection. Syntalith AI applications include custom solutions; the project scope and estimate can be discussed after reviewing the workflow. See the pricing page for offer information.

For an initial discussion, bring examples of incorrect field assignments with the approved values, a list of fields that matter most, and a short description of who checks the record and where it goes after correction. Those details help distinguish an image problem from a field-mapping problem and show how much manual work the current process involves. For image quality on European CMR documents, the separate CMR image-quality example before OCR shows how to assess an image before extraction.

Match a model to the task you need it to perform

Describe where your current AI falls short. We will compare model customization options, data requirements and the cost of running the resulting system.

Private LLMs and fine-tuning
Discuss a custom model