A convincing answer can still be unsupported.
An executive summary includes a market figure without a source that establishes it. The polished wording makes the gap easy to overlook.
Research: OpenAI · Why language models hallucinate (2025).
When an answer sounds right, what makes it ready to rely on?
Symbiain is developing a structured review method for teams using AI to prepare consequential recommendations. Examine the source material, identify unsupported claims and overlooked constraints, and record the checks a decision owner needs before acting.
The starting point is a scoped research pilot with Iain Davie: one recommendation, its source packet and a human-led comparative review.

Four recurring failure modes documented in research. Their prevalence varies by model, task and setting; this is not a frequency ranking.
An executive summary includes a market figure without a source that establishes it. The polished wording makes the gap easy to overlook.
Research: OpenAI · Why language models hallucinate (2025).
The proposed review connects material claims to evidence and identifies what needs qualification or verification.
Unsupported does not mean false: the source may be missing, incomplete or inapplicable. Comparative performance has not been established.
A recommendation overlooks an operating restriction buried in a long source packet. Supplying more material does not itself establish that every condition was used.
Research: Liu et al. · Lost in the Middle (2023); results concern the models and tasks studied.
The proposed review makes critical constraints explicit and checks whether the recommendation accounts for them.
The business example illustrates the problem; it is not a result measured in the cited research. Comparative performance has not been established.
A team asks an AI to support a preferred strategy. It reinforces the premise without adequately testing contradictory evidence or other options.
Research: Anthropic · Pilot alignment evaluation (2025).
The proposed review brings competing interpretations and disconfirming evidence into the same discussion before a judgement is accepted.
The business example illustrates the problem; it is not a result measured in the cited research. Comparative performance has not been established.
An extended research or analysis task may fail along the way. The team needs evidence of completion and an explicit account of unfinished work.
Research: METR · Task-completion time horizons (updated May 2026); software-task scope.
The proposed review records outcomes, unresolved checks and responsible owners so the next decision reflects the actual state of the work.
Making unfinished work visible does not establish improved task-completion reliability. Comparative performance has not been established.

A reviewable account of what is supported, what is assumed and what remains open.
Conflicting evidence and credible alternatives brought into the decision discussion.
A retained correction, an accountable owner and a clear reason to revisit the decision.
Public description of intended outcomes. Internal selection logic and implementation are outside this overview.
A fictional source packet and an authored review. All organisations, records and figures below are invented for teaching; this is not a client result or a measured model test.
“Our target market is growing by 12%. Existing operations can absorb the launch. A full launch is the preferred route. The board paper is ready for approval.”
These four short records are the complete evidence available for this exercise.
“Premium-segment demand grew 12% last year. This note does not estimate growth in the standard segment.”
The proposed expansion targets the standard segment.
“Spare capacity is 5%. The full launch requires 15% additional capacity. A second shift remains unapproved.”
“Management prefers a full launch. A staged launch is feasible in principle; its cost and timing have not been assessed.”
“Standard-segment forecast: pending. Capacity sign-off: pending. Staged-option comparison: not started. Decision owner: board sponsor.”
S1 supports 12% historical growth in the premium segment. It does not establish the target segment’s growth or a forward forecast.
S2 records 5% spare capacity against a 15% requirement, leaving a 10 percentage-point gap on the stated basis.
S3 identifies a staged option but supplies no comparison of cost, timing, capacity or reversibility.
S4 explicitly records three incomplete checks. A finished document is not evidence of a completed review.
Permitted next step: commission the three checks and compare the options. This packet does not support approval of the full expansion.
Accountable owner: board sponsor. Revisit when: relevant market evidence, a capacity sign-off and an option comparison are available; escalate if they remain unresolved at the decision deadline.
What would change this record: evidence that establishes target-segment demand, feasible capacity and the comparative case for an option. A correction can be superseded when new evidence warrants it.
The record demonstrates traceability and conditional judgement. It does not establish that Symbiain detects errors better than a competent reviewer or that expansion should or should not proceed in a real business.

Record what changed, why it changed, where the correction applies and when it should be reviewed. In this example, R2 remains open until a capacity check resolves it; a later approval may supersede it, while preserving the original record.
Carry a correction into another task only when its scope still applies. Retention alone does not establish better judgement.
How we would test the method →Bring one AI-assisted decision process. Define the evidence and review baseline. Test whether Symbiain improves the resulting judgement.
A limited founding pilot begins with one AI-assisted recommendation, available source material and a named decision owner. Scope, timing, responsibilities, confidentiality and fee are agreed before work begins.
Compare a specified Symbiain review with your normal competent review or a simpler checklist, using equivalent source material and accounting for time and effort. Agree the protocol and scoring before examining results.
Include sound recommendations as well as flawed ones. Assess warranted acceptance, not only the number of objections raised. Report null results and failures alongside any improvements.
Read the evaluation principles →Selected research identifies recurring failure modes. It does not validate Symbiain or establish which shortfall is most frequent in a particular client workflow.