WritingSystem design4 min read2026

Designing AI that knows when it is wrong

Confidence is not correctness. Fallback thresholds, uncertainty signals and refusal paths that keep a system dependable when the stakes are high.

Language models are always confident. They produce fluent, well-structured text whether they are right or wrong, and the fluency is the problem: it removes the cues people normally use to judge whether to trust something. A hesitant colleague signals uncertainty. A model does not, unless you build the system to.

Confidence is not correctness, and most of the work of making AI dependable in high-stakes settings is building the machinery that can tell the two apart and act accordingly.

Where uncertainty actually comes from

It helps to be specific about why a system might be wrong. In my experience the causes fall into a few buckets. The information needed is not available, because it was never in the corpus or retrieval failed to find it. The information is available but ambiguous or contradictory. The question itself is ambiguous. Or the task is outside what the model can do reliably, such as precise arithmetic over many values or reasoning across a very long chain of steps.

Each of these needs a different signal and a different response. A single confidence score bolted onto the output cannot distinguish them, which is why it rarely helps.

Signals you can actually use

The most reliable uncertainty signals usually come from the system around the model rather than the model's opinion of itself. Retrieval gives you several: how relevant the top results are according to the re-ranker, whether the results agree with each other, whether the answer can be grounded in specific spans at all. If the best passage you found is only loosely related to the question, that is a strong signal before the model has written a word.

Consistency is another. Ask the same question in slightly different ways, or sample several answers, and compare them. Agreement does not guarantee correctness, but disagreement is a reliable warning that the system is guessing.

Verification is the strongest. If an answer makes claims, check each claim against its source. If it produces structured output, validate it against the schema and against business rules. If it does arithmetic, recompute it in code. Anything that can be checked deterministically should be, and a failed check is the clearest uncertainty signal there is.

Model self-assessment, asking the model how confident it is, is the weakest of these. It can be a useful extra input, but I would never let it be the only gate on a consequential decision.

Thresholds and fallbacks

Signals only matter if they change behaviour. That means designing explicit thresholds and explicit fallback paths, and deciding them with the people who own the outcome rather than tuning them quietly in code.

A typical design has several tiers. When the evidence is strong and verification passes, answer directly with citations. When the evidence is partial, answer the part that is supported and say clearly what could not be confirmed. When the evidence is weak or contradictory, do not answer; explain what was searched and suggest where to look. When the action is irreversible or high-impact, route to a human regardless of confidence.

Where to set each threshold is a product decision about the relative cost of a wrong answer and a missing one. In a legal research tool, a confident wrong answer can be very costly and a refusal is cheap, so the thresholds should be strict. In a brainstorming tool the trade-off runs the other way. The engineering job is to make the trade-off explicit and adjustable, not to pretend it does not exist.

Refusal is a first-class path

Refusal tends to be treated as a failure state, something to minimise. I think that is backwards. A well-designed refusal is one of the most valuable things a system can do, because it is the moment the system protects the user from a bad decision.

Good refusals are specific. Not 'I cannot help with that', but 'I searched the contracts in this matter and found no termination clause; the closest match is the renewal clause in section 9'. That tells the user what happened, what the system did and what they can do next. It turns a dead end into a useful result.

Refusals also need testing. Your evaluation set should contain questions the system must refuse, and a model or prompt change that makes it answer those questions is a regression, even if it looks like an improvement on a helpfulness metric.

Humans in the loop, deliberately

Human review is the last fallback, and it only works if it is designed for. Escalation should carry context: the question, the evidence found, the signals that triggered review and the draft answer if there is one. The reviewer's decision should flow back into the evaluation set. And the volume should be watched, because a system that escalates everything has simply moved the work rather than done it.

In systems that take actions, such as an agent that changes infrastructure, this matters even more. The pattern I trust is to make the agent's permitted actions small and reversible, require approval for anything outside that set, and verify the outcome after every action with an automatic rollback if the system does not recover. The agent does not need to know when it is wrong if the system around it can detect a bad outcome and undo it.

Dependable beats impressive

It is tempting to judge AI systems by how impressive their best answers are. In production, what matters is how they behave at their worst. A system that knows the limits of its evidence, says so clearly and hands off gracefully will earn more trust than one that is occasionally brilliant and occasionally, invisibly, wrong.

That trust is built in the unglamorous layers: retrieval signals, verification, thresholds, refusal paths and review loops. None of it is visible in a demo. All of it is visible the first time the system is wrong.

Next step

Have a system that needs
to hold up in production?