Skip to content
Back to blog
AI qualityControls for unsupported model output

AI hallucinations: controls that hold up

Choose controls for AI output by consequence, evidence, refusal rules and human review. A practical guide for a safe business pilot.

A fluent answer is not evidence. Choose controls around the consequence of an error, the sources available to the system, and the point where a person must approve the result.

Author

Syntalith

Published Updated 7 min read

An AI model can produce a fluent answer without having a reliable basis for it. That is usually called a hallucination. The useful business question is not whether a model is perfect. It is what the system does when evidence is missing, contradictory, stale, or outside the model's remit.

Choose controls for one workflow, define the harm of a wrong result, and pay for only what that workflow needs.

Start with the consequence

Classify the output before choosing a model or a user interface.

OutputWhat can go wrongSensible control
Internal draftA colleague spends time correcting itA person checks the draft before reuse
Operational recommendationA wrong field or instruction sends work down the wrong pathSource links, structured fields, validation, and approval before the action
External or consequential actionA wrong commitment, record, access decision, or contract interpretation reaches another partyEvidence requirement, explicit refusal state, named approver, audit trail, and a rollback path

The NIST AI Risk Management Framework describes trustworthy AI through characteristics such as validity and reliability, safety, security and resilience, accountability and transparency, explainability, privacy, and fairness. Those characteristics are useful design prompts. They do not replace a workflow-specific acceptance test. A control is adequate only when it catches the failures that matter in that workflow.

Keep the answer tied to evidence

Grounding requires a traceable chain. For each answer, the system should be able to show:

  1. which source set it was allowed to use;
  2. which passage, row, or record it retrieved;
  3. which fields it extracted or calculated;
  4. which rule allowed the final answer; and
  5. which person or downstream system accepted it.

For a document assistant, a citation should lead to a stable file and a useful locator such as a page, section, paragraph, or record identifier. For a workflow assistant, the output should use a schema that makes missing fields visible. A citation that cannot be opened or checked provides no control.

Source access also needs boundaries. Apply permissions before content reaches the model, keep current and superseded versions distinguishable, and treat contradictory sources as a state that needs resolution. Adding more text to the prompt does not repair an unclear source hierarchy.

Refusal is a normal result

The system needs a result for "the evidence is not sufficient." That result can be a refusal, a request for a missing document, or a handoff to a named owner. It should not be a blank field that downstream automation interprets as approval.

Define refusal conditions before implementation. They may include:

  • no source passes the relevance threshold;
  • the source is outside the requester's permission;
  • two current sources disagree;
  • a required field is missing or cannot be validated;
  • the request asks for a decision the workflow has not authorized; or
  • a tool, parser, or source connection failed.

Each condition should have an owner and a next action. The wording can be concise, but the state must be visible in logs and in the user interface.

Test the failure modes that matter

Build a small, reviewed question set from the real workflow. Include known answers, paraphrases, questions with no answer, stale versions, conflicting sources, exact identifiers, permission-restricted records, prompt-injection text inside a document, and tool failures. Keep the expected outcome for each case: answer with a citation, ask for clarification, refuse, or route to a person.

Measure more than answer similarity. Check source correctness, field validity, refusal correctness, permission enforcement, latency, and the rate of actions that require a human correction. Separate the test set from the material used to tune prompts. NIST recommends realistic and representative testing, clear measures, and ongoing monitoring after deployment. A single accuracy number cannot represent those controls.

Put a person where the system cannot self-correct

Human review works as a routing rule. Define the events that require approval: an external message, a change to a financial or contractual record, an access decision, a high-impact recommendation, or an unresolved conflict. Give the reviewer the source passages, the proposed action, and a way to reject or amend it.

For low-consequence drafts, a sampled review may be enough. For a consequential action, the system should not be able to mark its own output as approved. Record who approved it, when, and which source version was visible at the time.

Record what changed

A production control loop needs monitoring for source changes, retrieval failures, refusal rates, corrections, and permission errors. Keep model and prompt versions alongside the result. Review a sample after changes to the source set or workflow rules, then repeat the full acceptance set when the change can affect the decision.

Privacy is part of the design. Decide what personal data enters prompts and logs, who can read those logs, how long they remain available, and which processor agreements apply. The GDPR text is the primary legal reference; a deployment still needs a company-specific assessment.

A pilot that can be accepted

Write a one-page pilot specification before building:

  1. name the workflow and the person accountable for it;
  2. list the source systems, permissions, and version rules;
  3. define the output schema and refusal states;
  4. assemble the reviewed question and failure set;
  5. set approval, logging, and rollback rules; and
  6. decide what evidence is sufficient for release.

If the workflow has no clear owner, no stable source, or no safe action after a refusal, pause the build. Resolve that process decision before selecting a model.

When a simple review is enough

You may not need retrieval, tools, or an agent for a low-consequence writing task. A model that drafts text in a document, with a person checking every result, can be the right scope. More machinery becomes justified when the output must be current, source-linked, permission-aware, or connected to an action. Start with that requirement and keep the first release narrow.

The AI process audit service turns these controls into acceptance cases for one workflow. The quotation-control test shows how source checks stop unsupported text before release, using generated requests rather than client data.

FAQ

Are citations enough to stop hallucinations? No. A citation helps a person check the answer, but it can point to the wrong or outdated passage. Test source correctness and version handling as separate requirements.

Should every AI answer be reviewed by a person? Review intensity should follow consequence. A low-risk draft can use sampling. An external or consequential action needs an approval rule that the system cannot bypass.

Does a larger language model remove the problem? Model capability can change failure rates, but it does not create a source hierarchy, permission boundary, or accountable approval step. Those belong in the workflow.

Sources

Free process scan

Start with a free process scan.

  • A 30-minute call with the engineer who would lead the work.
  • A review of the processes that cost you the most time and money.
  • A written summary of what to automate first and the likely cost range.

The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.

€0

30 minutes · written takeaway within 2 business days

Book a free process scan (30 min)

Times are shown in your own time zone. We work with clients across time zones.

Describe the process in the form