Skip to content
← Back to blog
local-modelsArticle

How to evaluate an LLM before deployment

Test a private LLM on business documents: quality, abstention, access, latency and error costs. How Syntalith defines deployment acceptance criteria.

Author

Syntalith

Published Updated 2 min read

A model answers five demonstration questions correctly. The next day it encounters an upside-down scan, a missing table and a contract the user is not allowed to read. Those cases reveal whether the system is ready for company work.

Evaluation is a repeatable test of task performance. Before selecting Qwen, DeepSeek or another model, we define a correct result and the errors that block deployment. OpenAI’s evaluation guidance follows the same principle: objectives, datasets and metrics should reflect the actual application.

Start with the user’s decision

A legal answer needs the correct clause and version. A help desk needs the correct destination team. In accounting, a wrong supplier identifier has different consequences from an extra space in a description. Ambiguous production instructions should be reviewed by the responsible person.

We record those differences in acceptance criteria. A style score cannot replace checks on identifiers, categories and evidence. The process owner should help create the evaluation because they know which mistakes create further work or risk.

Build a test the model has not trained on

Include ordinary cases, less frequent exceptions and missing information. Add Polish characters, industry abbreviations, mixed languages and poor-quality documents where those occur in the work. These are proposed test groups, adjusted to the application.

Keep training examples separate from the final test. Related documents from one case can make a model appear more capable than it is on new customers. Record dataset versions and scoring rules so subsequent releases can be compared against the same baseline.

Measure different kinds of failure

AreaWhat to testWhat the result tells you
Task qualityCorrect fields, categories and citationsWhether the output is useful
UncertaintyMissing sources, contradictory documents, unfamiliar casesWhen the system should abstain or escalate
AccessThe same question from users with different rolesWhether permissions are enforced
LoadShort and long inputs, concurrent requestsQueue length and actual waiting time
OperationsRestarts, unavailable GPUs, weight or template changesBehaviour after outages and updates

For classification, show results per category. A strong average can hide poor handling of rare complaints. For extraction, count missing fields separately from invented fields. For RAG, check whether the cited passage supports the answer.

Confidence thresholds need evaluation too

A score of 0.9 does not establish that nine out of ten similar decisions will be correct. We measure that relationship on task data. It is particularly relevant to Jev and decision models, which return probabilities.

Compare quality at different thresholds alongside the proportion of cases sent to a human. A high threshold may reduce errors while leaving most of the workload with operators. The report needs both measures.

What the assessment delivers

A useful report records model version, configuration, category-level results, error examples, load measurements and a recommendation. We retain the previous working version and define rollback conditions. The same evaluation can run after fine-tuning or a quantization change.

Syntalith provides private LLM selection and evaluation before deployment. Describe the process and the consequences of a wrong answer. We can propose a test that supports a practical decision about whether and how to deploy.

Evaluate private AI for your organization

We help businesses and individuals select hardware, deploy a model and test it on their own tasks. Start with a computer you already own or ask us before buying one.

Private LLMs and fine-tuning
Discuss private AI