Buying local AI with a test on your documents
Before buying a local document assistant, compare its answers on tasks your team actually performs. Check whether it finds the right source version, identifies missing information and reduces the time employees spend reviewing results. A focused test can show whether existing search, an approved private service or a local system fits the work.
Syntalith
A polished demo and a model benchmark do not show whether a tool can answer questions from your company's documents. Set the task, allowed sources and acceptance criteria before comparing options. Then assess both the answer and the time a reviewer needs to check it.
One question, two document versions
Consider a purchasing team comparing supplier offers. Its folder contains an older quote with a delivery window and a newer order document. The newer document refers to an amendment that sets the delivery window, but that file is missing from the sources allowed in the trial. An employee asks which window applies to the specified order version.
The assistant may retrieve the old quote and state its delivery window in fluent prose. The quote supports that date for the earlier offer, while the order document shows that an amendment determines the later window. The useful answer should point out the reference to the amendment and say that the approved source set does not contain the evidence needed to identify the current window. A purchasing owner can review the cited documents and decide whether to add the amendment before anyone relies on an answer.
Turn this into a test task with a question, a fixed set of sources, the order version, the expected answer and a description of disqualifying errors. Assigning the old delivery window to the later order is a critical failure. A confident tone or similar wording does not make the answer acceptable.
Set the criteria before comparing
Build test tasks from questions employees ask and mistakes the team has seen. For each task, define which sources the assistant may use, what it should return, how it should show its evidence, and when it should report that information is missing. Include a question that has an answer in the current version and another where the referenced amendment is absent. These cases help distinguish a missing source from a misread passage.
Anthropic's guidance on evaluating AI agents recommends tasks based on observed failures, clear criteria and review of failed cases. For a document assistant, those ideas point to testing source selection and the handling of missing evidence.
Keep examples used to improve a system separate from tasks reserved for acceptance. Compare the same question, context, source set and document versions across options. If one system receives a newer file or extra hint, its result is no longer directly comparable.
Have a purchasing subject-matter owner assess whether the assistant found the right source and preserved its terms. Record how much checking the answer requires and whether the reviewer can retrace its basis. Failure to retrieve the current document calls for a different check from finding the right passage and misstating its condition. That distinction helps the team decide what to change and which task to repeat.
Compare options on equal terms
The comparison can include the search already in use, an approved hosted or private service, and a local system. First decide which options can use the same materials and perform the same task. A language model may add little when existing search quickly finds the current document an employee needs.
Judge results against the criteria the team agreed on. Does the answer identify the right version and source? Does it report that the amendment is missing? How long does an employee spend checking the result? How long does a useful answer take under conditions the team accepts? Decide who will operate the system and review its quality after changes. Benchmark scores and generation speed belong alongside these task results.
Assess source retrieval separately from answer quality. If the assistant cannot find the newer order document, the problem may be access to the file or how it is indexed. If it finds the passage but changes its meaning, the response itself needs attention. Before choosing a local option, map how a document travels from employee to result. Include external components that may process it, such as file-reading or logging services, and identify who can access the materials.
A failed test can change the purchase decision. Retrieving the wrong version may call for better document organization or search. Guessing the delivery window when the amendment is missing calls for a change in how the system handles incomplete evidence, followed by another test of that case. If the revised result still misses the agreed criterion, the buyer can narrow the intended use, return to simpler search or stop the purchase for that task.
Before the trial, agree that its report will state the conditions, accepted and failed examples, and the tasks covered by the decision. It should show tradeoffs among useful answers, review time and the selected environment, then set conditions for another trial, changes or stopping.
What to bring to a local AI discussion
Syntalith develops AI applications for businesses, including solutions that use local and private models. For this purchasing task, a comparative assessment could cover the test conditions, accepted and failed examples, and a recommendation on whether to improve retrieval, develop a local application or stay with the current tool. Agree on that scope before the work begins.
Bring two or three questions employees regularly ask, the documents they use today, an example of a known error, and the person who can approve the expected answer. Those details help define the tasks, choose who will review them and identify which options belong in the comparison. See Syntalith pricing and services when planning the discussion.
If the test supports a local system, the next decision is whether to run it on a workstation or a shared server. The article on local AI on a workstation or shared server covers that infrastructure choice.
Evaluate private AI for your organization
We help businesses and individuals select hardware, deploy a model and test it on their own tasks. Start with a computer you already own or ask us before buying one.
Private LLMs and fine-tuning