Skip to content
← Back to blog
local AIArticle

Small local AI models for internal ticket routing

Requests for printer setup and application access reach the help desk in the morning. A staff member reads them before assigning each to the responsible team. A small local model could suggest the destination, but it needs to keep up with the morning workload on the company's equipment. Assess both the accuracy of its suggestions and how long tickets wait for assignment.

Author

Syntalith

Published Updated 5 min read

Suppose desktop support handles printer setup, while a business application's administrators handle access requests. Employees describe what they need in their own words without selecting a category on a form. The person covering the help desk has to decide where each request belongs.

A model could show a short team suggestion beside each ticket: desktop support for printer setup, or the application's administrators for an access request. The staff member reads the request and suggestion together, checks them, and confirms the assignment. The team name is enough for this task. A long explanation of possible solutions does not help make that decision.

The morning workload shows whether the hardware fits

One suggestion produced on an otherwise idle computer reveals only part of the job. When employees start their day and submit requests close together, tickets may have to wait for the model. Help desk staff need to know whether a suggestion will arrive before they would read and assign the ticket themselves.

A useful trial therefore runs on the computer or server intended for ongoing use, with arrivals resembling the morning workload. Measure the time from a ticket's arrival to a usable suggestion. Timing only the generation of a team name misses the wait before processing starts. It is also worth observing startup, when the model still needs to load into memory.

Ollama's documentation describes how concurrent processing depends on available memory. Processing several requests at once requires additional memory, and an overloaded queue can reject further requests. That gives the buyer a reason to assess the chosen model on the company's actual hardware and workload. Calling a model “small” does not establish when the last suggestion in a morning batch will arrive.

The help desk manager can then compare how many tickets received the right suggestion in time and how many staff assigned themselves. A model that responds correctly after someone has already distributed the tickets does not reduce that part of the work. A fast model that sends requests to the wrong teams creates reassignment work. Both outcomes matter to the purchase decision.

Some tickets can bypass the model

When a form already identifies the requested service clearly, an existing rule may identify the responsible team. For example, GLPI documents ticket business rules with conditions and actions applied when tickets are created or updated. Those rules run in a defined order.

It makes sense to use information the employee has already supplied. The model could handle free-text requests that the rules have not assigned. The company can then assess whether it needs an additional tool for the entire queue or only for part of it.

If the local trial reveals repeated category mistakes despite clear team responsibilities, our article on fine-tuning for ticket classification addresses that separate decision. First, the team needs to establish whether the proposed model and hardware fit the daily pace of work.

Staff still see tickets without suggestions

A request saying “I can't work in the application” may not establish whether the issue concerns access or something else. Help desk staff still need to read it. They can assign it or ask the employee for details without waiting for the model to settle the question.

The same applies when a suggestion is late or an overloaded model cannot produce one. Staff can see the ticket in the help desk system and take it over manually. If a suggestion arrives later, the application needs to respect the assignment someone has already made.

The manager also needs to understand how much work remains. If staff still have to read most morning requests themselves, the trial may support a narrower scope or retaining the existing process. The team should be able to run its regular shift when the local model is unavailable.

What stays local

A local model runs on a designated company device. Ollama describes a local-only mode that disables cloud models and web search. That is a setting in the software running the model. It does not determine where every copy of a ticket is stored in the help desk system or how its integrations work.

For this task, establish which ticket content reaches the model and where suggestions are retained. If data-handling requirements are the reason for choosing local processing, the help desk owner and the infrastructure administrator need to agree on that flow too.

Syntalith can help build an AI application for ticket routing and assess a local model on the company's hardware. The proposed scope could cover a narrow group of routine requests, suggestions for staff, and comparison with existing rules. The help desk team contributes its knowledge of the correct destinations and the busiest arrival periods. The results can inform whether to develop the connection to the current ticketing system. Pricing information is available on our pricing page.

The first conversation can start with a typical morning: what enters the queue, and which requests make staff spend the most time deciding where they belong? A description of the available hardware will help define the trial.

Evaluate private AI for your organization

We help businesses and individuals select hardware, deploy a model and test it on their own tasks. Start with a computer you already own or ask us before buying one.

Private LLMs and fine-tuning
Discuss private AI