Skip to content
Back to blog
PracticeA test for one real business process

Do AI agents work? A buyer's test for 2026

AI agents can handle a narrow process when the input, permitted actions, owner and quality measure are explicit. Use this test to separate an operable system from a vague automation promise.

An agent earns trust through a defined case, permitted actions, escalation and an inspectable record. The model label cannot answer those operating questions.

Author

Syntalith

Published Updated 7 min read

AI agents can work when they handle a defined process with suitable data, approved actions and a person responsible for exceptions. The same technology can disappoint when a company buys a vague promise such as “handle everything” and leaves the process, measure and owner unspecified.

The buying question is therefore concrete: can this system perform one recurring case to an agreed standard, and can the team operate it after release?

What a working agent must accomplish

Describe the process in one sentence: “When this record arrives, prepare this result, update these fields after approval and send the remaining cases to this owner.” Then make six parts explicit:

  • input and required sources;
  • permitted read and write operations;
  • output schema and source references;
  • stop conditions and human handoff;
  • quality measure and error severity;
  • owner for incidents, policy and updates.

An agent can interpret varied language, choose among approved tools and prepare a next step. Those capabilities become useful only when the process tells the system what counts as completion and what requires review.

Inspect the process before the model

A company may need a workflow, a ready SaaS product or a checklist. A custom agent is a candidate when cases vary in language or document structure, a context-dependent decision recurs and the team can provide representative inputs.

Postpone a build when:

  1. the rules live in one person's memory and have no stable owner;
  2. the required records are missing, inaccessible or too inconsistent to evaluate;
  3. every case requires unique judgment that cannot be described for review;
  4. no one can spend time checking exceptions after release;
  5. the volume or consequence cannot justify integration and operating work.

Writing the process down still creates value. It may reveal that a deterministic flow or a small data-quality project solves the actual problem.

Ask the vendor for operating evidence

A credible supplier should explain the case path, the tools, the stop conditions and the maintenance model. Ask for:

  • a representative evaluation set and its expected outputs;
  • the source location retained for important fields;
  • the permissions attached to each tool;
  • the rule that sends a case to a human;
  • logs an operator can inspect after a failure;
  • the rollback or manual path during an outage;
  • the people who own code, instructions and policy after release.

Claims about general intelligence or a polished interface do not answer these questions. Ask how the system handles a missing source, conflicting records, a malformed request and a provider outage.

Measure a complete case

Score the whole process rather than the final prose. Useful measures include:

MeasureWhat to record
Correct completionCases that satisfy every required field and policy
Correct holdCases stopped for the right reason and routed to the right owner
Unsafe releaseCases that reached an action without required approval
Correction effortOperator time and changes per accepted case
Trace qualitySource, policy, tool and decision records available after the run
Operating costModel, infrastructure, support and human review for the complete case

The reference set should include routine cases, missing data, contradictory values, unusual formats and an explicit outage path. Keep the cases used to tune the instructions separate from the final evaluation set.

Let the first process decide

Run the system in shadow mode when an incorrect send or write would be hard to reverse. Compare its proposed results with the current owner's decision. Inspect errors by impact: a style correction and a wrong customer record need different remedies.

Release only the cases that satisfy the process policy. Review the owner workload after the first slice. An agent that produces good drafts but creates an unmanageable review queue needs a narrower scope or a different operating design.

Build a decision record

Before expanding, record:

  • the process and volume measured;
  • the acceptance set and quality threshold;
  • permitted tools and approval roles;
  • error classes and response plans;
  • cost of a complete run;
  • maintenance owner and review date.

This makes a later model or integration change a reasoned re-evaluation rather than a restart from opinion.

Decide whether to buy

An AI agent is a sensible candidate when a named process has recurring volume, useful data, context-dependent steps, measurable review work and an owner. A workflow, SaaS product or manual procedure is often better when the rules are fixed, cases are rare or the organisation cannot support ongoing review.

Bring one process and its current handling data to a free process scan. Syntalith can help define the acceptance set and return a first scope, including the cases that should remain with a person.

Frequently asked questions

Do AI agents work?
They can work for a defined process with a measurable result, approved tools, human escalation and a trace. They need a representative evaluation and an owner after release.
What should a vendor show before a purchase?
Ask for the process scope, acceptance criteria, failure handling, permission model, operating owner and a way to inspect the system's decisions on representative cases.
When should a company postpone an agent?
Postpone the build when the process has no owner, required data is unavailable, rules change faster than they can be documented or the expected value cannot cover implementation and review work.
How do I start evaluating an agent?
Choose one recurring process, record its current volume and handling effort, prepare representative cases and run a limited evaluation with a named reviewer.

Free process scan

Start with a free process scan.

  • A 30-minute call with the engineer who would lead the work.
  • A review of the processes that cost you the most time and money.
  • A written summary of what to automate first and the likely cost range.

The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.

€0

30 minutes · written takeaway within 2 business days

Book a free process scan (30 min)

Times are shown in your own time zone. We work with clients across time zones.

Describe the process in the form