Skip to content
← Back to blog
document classificationArticle

One PDF, several documents: how AI can identify them

A law firm receives a PDF containing an agreement, an amendment, and a printed email. The entire file is indexed as an agreement, so anyone looking for the amendment has to open it and browse the pages. AI can suggest separate entries identifying each document type and its location in the file. That organization helps the matter team reach the material they need directly.

Author

Syntalith

Published Updated 4 min read

An index that shows what the file contains

Suppose a firm receives a bundle in which pages one through four contain a service agreement, pages five and six contain a separate amendment, and page seven is an email printout. Assigning one document type to the whole file hides the other two parts from someone scanning the index.

A useful suggestion would show three entries: the agreement on pages 1–4, the amendment on pages 5–6, and the email on page 7. An employee opens those locations and checks the division. Once it is accepted, the person handling the matter can jump from the index straight to the amendment. The original bundle remains available, preserving the context of which materials arrived together.

The boundary between documents matters. If the agreement’s last page is included in the amendment entry, getting the type name right offers limited help. The reviewer needs to see where each part begins and ends, and be able to correct its page range.

A model performing this task is called a classifier. It assigns a document to an agreed type, such as an agreement or correspondence. That label does not determine whether the document is valid, signed, or has a particular legal effect. It helps the team find material for further work.

Do you need a custom tool?

First, look at how documents arrive. If the sender can supply separate, clearly named files, some organizing work disappears at intake. Existing rules or templates may be sufficient for recurring forms. The document management system, or DMS, may also offer indexing and bundle-splitting features the team has yet to use.

A model is worth trying when bundles routinely combine different materials and an employee has to identify their contents each time. Microsoft describes a custom Document Intelligence classifier that recognizes trained document types and their page ranges within a bundle.

Recognizing the type is one part of the solution. The firm also needs a place for an employee to review the proposed index, adjust boundaries, and save the result with the matter. If suggestions remain in a separate file that has to be copied manually, much of the original work remains.

What adaptation adds

Start by testing an available tool on bundles like those the team organizes every day. If it recognizes the required documents and locates their pages correctly, further customization may be unnecessary. Assess a wrong document label separately from an incorrect split: they require different corrections from the reviewer.

Adapting a classifier becomes worth considering when mistakes recur. For example, a tool may regularly merge an amendment with the preceding agreement in bundles containing visually similar documents. The firm’s staff then provide examples of the correct division and agree on the names the index should use. The types should reflect what employees actually look for.

Some material will not fit the agreed types confidently. A short cover note need not belong to the agreement category simply because it precedes its pages. A useful review screen leaves that part for an employee to identify, with its location in the original file visible. The material stays in the index even while its type needs clarification.

Does the index help the matter team?

Compare results on separate bundles that were not used to adapt the classifier. The person handling documents checks whether every part is visible and whether links open the right pages. They can also assess how much correction the index needs and whether preparing it is easier than their current method.

If the organized documents will later support a chronology, see our article on local AI for legal matter documents. It explains how to work with dates and return to the passages supporting them.

What you can commission from Syntalith

Syntalith builds AI applications and adapts models. For a law firm, we propose a tool that prepares an index of the documents inside an incoming PDF, with links to the relevant pages and a way to correct the division. We also propose saving that result in the firm’s document workflow so the matter team can use the index.

Your staff contribute examples of divisions they currently make by hand and identify the document types they need. We use those examples to assess the current system and an available classifier. Recurring errors then help establish whether adaptation is justified.

Start the first conversation by describing one bundle in which an attachment was difficult to find. Explain how an employee separated it and where the result was saved. If materials are needed for further assessment, we will agree on how to use them. Engagement information is on Syntalith’s pricing page.

Match a model to the task you need it to perform

Describe where your current AI falls short. We will compare model customization options, data requirements and the cost of running the resulting system.

Private LLMs and fine-tuning
Discuss a custom model