Casebook / 29 · AI decision integrity

More capable AI.
More accountable
decisions.

When an answer sounds right, what makes it ready to rely on?

Symbiain is developing a structured review method for teams using AI to prepare consequential recommendations. Examine the source material, identify unsupported claims and overlooked constraints, and record the checks a decision owner needs before acting.

The starting point is a scoped research pilot with Iain Davie: one recommendation, its source packet and a human-led comparative review.

Conceptual architectural artwork: fragmented translucent records resolve into an ordered field of evidence and open questions.
Field plate / From fluent output to inspectable judgement · Concept artwork
Evidence-led researchHuman decision authorityPublic capability overviewPrototype · Benefits under evaluation
01 / Start with the problem

Where capable AI
still needs scrutiny.

Four recurring failure modes documented in research. Their prevalence varies by model, task and setting; this is not a frequency ranking.

01 / GroundingUnsupported certainty
What the team encounters

A convincing answer can still be unsupported.

An executive summary includes a market figure without a source that establishes it. The polished wording makes the gap easy to overlook.

Research: OpenAI · Why language models hallucinate (2025).

Symbiain’s proposed contribution

Know the basis before relying on the claim.

The proposed review connects material claims to evidence and identifies what needs qualification or verification.

Unsupported does not mean false: the source may be missing, incomplete or inapplicable. Comparative performance has not been established.

02 / ContextMissed context
What the team encounters

A relevant document can be present—and still missed.

A recommendation overlooks an operating restriction buried in a long source packet. Supplying more material does not itself establish that every condition was used.

Research: Liu et al. · Lost in the Middle (2023); results concern the models and tasks studied.

Symbiain’s proposed contribution

Keep consequential conditions in view.

The proposed review makes critical constraints explicit and checks whether the recommendation accounts for them.

The business example illustrates the problem; it is not a result measured in the cited research. Comparative performance has not been established.

03 / ChallengeExcessive agreement
What the team encounters

Agreement can crowd out examination.

A team asks an AI to support a preferred strategy. It reinforces the premise without adequately testing contradictory evidence or other options.

Research: Anthropic · Pilot alignment evaluation (2025).

Symbiain’s proposed contribution

Give the alternative a fair hearing.

The proposed review brings competing interpretations and disconfirming evidence into the same discussion before a judgement is accepted.

The business example illustrates the problem; it is not a result measured in the cited research. Comparative performance has not been established.

04 / ExecutionFragile follow-through
What the team encounters

A promising start is not a completed task.

An extended research or analysis task may fail along the way. The team needs evidence of completion and an explicit account of unfinished work.

Research: METR · Task-completion time horizons (updated May 2026); software-task scope.

Symbiain’s proposed contribution

Make completion and open work distinguishable.

The proposed review records outcomes, unresolved checks and responsible owners so the next decision reflects the actual state of the work.

Making unfinished work visible does not establish improved task-completion reliability. Comparative performance has not been established.

Conceptual navy architectural installation with translucent evidence planes, brass connections and an unresolved gap.
Plate 29.2 / Pause before commitment · Concept artwork
  1. RecordsThe panels stand for claims and the material used to support them.
  2. ConnectionsThe threads represent relationships that need to be checked.
  3. The gapA missing evidential link remains open until a check resolves it.
02 / The broader picture

Make the basis for action
visible to the people responsible.

EvidenceJudgementHuman reviewRevision
01 / CLARITY

What can we rely on?

A reviewable account of what is supported, what is assumed and what remains open.

02 / CHALLENGE

What are we missing?

Conflicting evidence and credible alternatives brought into the decision discussion.

03 / CONTINUITY

What changes next?

A retained correction, an accountable owner and a clear reason to revisit the decision.

Public description of intended outcomes. Internal selection logic and implementation are outside this overview.

03 / Inspect the working

One recommendation.
Every condition visible.

A fictional source packet and an authored review. All organisations, records and figures below are invented for teaching; this is not a client result or a measured model test.

Original briefing / fictional record B1

“Approve a full expansion next quarter.”

“Our target market is growing by 12%. Existing operations can absorb the launch. A full launch is the preferred route. The board paper is ready for approval.”

Read the source packet

These four short records are the complete evidence available for this exercise.

S1 / Market note · 1 June · fictional
“Premium-segment demand grew 12% last year. This note does not estimate growth in the standard segment.”

The proposed expansion targets the standard segment.

S2 / Operations note · 3 June · fictional
“Spare capacity is 5%. The full launch requires 15% additional capacity. A second shift remains unapproved.”
S3 / Strategy note · 4 June · fictional
“Management prefers a full launch. A staged launch is feasible in principle; its cost and timing have not been assessed.”
S4 / Review checklist · 5 June · fictional
“Standard-segment forecast: pending. Capacity sign-off: pending. Staged-option comparison: not started. Decision owner: board sponsor.”

Follow each finding into the decision record

R1 / Claim → source → correction

The growth figure covers a different segment.

S1 supports 12% historical growth in the premium segment. It does not establish the target segment’s growth or a forward forecast.

Revised claim
Target-segment growth is unverified in this packet. The premium figure should not be used as its estimate.
Check and owner
Research lead: obtain a dated standard-segment forecast and reconcile its scope.
Decision consequence
Do not rely on the 12% figure to justify the expansion.
R2 / Constraint → discrepancy → check

Available capacity does not meet the stated requirement.

S2 records 5% spare capacity against a 15% requirement, leaving a 10 percentage-point gap on the stated basis.

Revised claim
The full launch depends on additional capacity that has not been approved.
Check and owner
Operations lead: confirm the capacity basis and the feasibility, cost and approval of a second shift.
Decision consequence
Treat launch capacity as an unresolved condition before commitment.
R3 / Preference → alternative → comparison

A preferred option has not earned comparative support.

S3 identifies a staged option but supplies no comparison of cost, timing, capacity or reversibility.

Revised claim
Management preference is recorded; superiority of a full launch is not established.
Check and owner
Strategy lead: compare full launch, staged launch and deferral against the same criteria.
Decision consequence
Bring the comparison to the board sponsor; do not select a winner from this packet.
R4 / Completion → open work → ownership

The paper is drafted; its checks are unfinished.

S4 explicitly records three incomplete checks. A finished document is not evidence of a completed review.

Revised status
Ready for discussion; material conditions remain open before commitment.
Check and owner
Board sponsor: assign due dates and review the evidence returned by each named lead.
Decision consequence
Reopen the decision when the checks are resolved or the deadline requires escalation.
D1 / Authored decision record · revision 1

Discuss the options. Resolve the conditions.

Permitted next step: commission the three checks and compare the options. This packet does not support approval of the full expansion.

Accountable owner: board sponsor. Revisit when: relevant market evidence, a capacity sign-off and an option comparison are available; escalate if they remain unresolved at the decision deadline.

What would change this record: evidence that establishes target-segment demand, feasible capacity and the comparative case for an option. A correction can be superseded when new evidence warrants it.

The record demonstrates traceability and conditional judgement. It does not establish that Symbiain detects errors better than a competent reviewer or that expansion should or should not proceed in a real business.

Conceptual ivory architectural sequence with retained record panels and a branching brass thread across time.
Plate 29.3 / Retain the change · Concept artwork. Panels represent successive records; the branching thread represents a revision whose earlier basis remains inspectable.
Retain the change

A system should carry
its corrections forward.

Record what changed, why it changed, where the correction applies and when it should be reviewed. In this example, R2 remains open until a capacity check resolves it; a later approval may supersede it, while preserving the original record.

Carry a correction into another task only when its scope still applies. Retention alone does not establish better judgement.

How we would test the method →
04 / From demonstration to evidence

One workflow.
A bounded founding pilot.

Bring one AI-assisted decision process. Define the evidence and review baseline. Test whether Symbiain improves the resulting judgement.

What the proposed pilot would include

A limited founding pilot begins with one AI-assisted recommendation, available source material and a named decision owner. Scope, timing, responsibilities, confidentiality and fee are agreed before work begins.

You bring
The original recommendation, its source packet, relevant constraints, decision deadline and your existing review approach.
We agree
A bounded review process, a competent comparison method, scoring criteria and who independently assesses the results where feasible.
The proposed outputs
A claim-and-evidence record, corrections, unresolved checks with owners, and a findings note covering benefit, effort and failures.
The next decision
Continue, revise or stop according to the agreed findings. Useful learning may include showing that the additional process is not worthwhile.
The test that matters

Does this add value beyond an ordinary review?

Compare a specified Symbiain review with your normal competent review or a simpler checklist, using equivalent source material and accounting for time and effort. Agree the protocol and scoring before examining results.

  • Missed material errors: what remains wrong after each review?
  • Incorrect flags: which sound claims or recommendations are challenged without sufficient reason?
  • Correction quality: does the revision resolve the problem without introducing another?
  • Review effort: what additional time and work does the process require?

Include sound recommendations as well as flawed ones. Assess warranted acceptance, not only the number of objections raised. Report null results and failures alongside any improvements.

Read the evaluation principles →
Research basis and limits

Selected research identifies recurring failure modes. It does not validate Symbiain or establish which shortfall is most frequent in a particular client workflow.