Casebook / 24

Before an AI answer
becomes a decision.

Trace its claims to evidence, inspect what remains uncertain, and see what changes when a new record arrives.

Fluency is not provenance.

An original Symbiain Field Manual plate showing a fluent machine-generated document decomposed into supported, contradicted and unresolved claim paths before human review.
Plate 24.1 / Fluency is not provenance. One answer becomes several inspectable claims; supported, contradicted and unresolved paths remain distinct before reliance.
RESEARCH PROTOTYPE

This page demonstrates a way of structuring inquiry. It does not constitute professional advice, verified intelligence, a recommendation or a conclusion about any person, organisation or transaction.

Try the method / fictional teaching case

An answer you
can open up.

Select a claim. Read its evidence. Then introduce a new record and see what changes—and what stays open. These are authored examples, not live AI outputs or client records.

01 / Select part of the original answer

The task: explain the delay and advise the decision owner from the supplied records. The confident answer above contains errors.

02 / Claim C2 · Version 1

The supplier dispatched it on 17 June.

Contradicted by this packetHandling: Qualify

17 June is the receipt date. Collection is recorded on 13 June. The initial answer has substituted arrival for dispatch.

C2S3Human review
Why this matters to the decision
Do not use an arrival date as proof of a late dispatch.
Responsible next step
Correct the chronology: collection recorded on 13 June, receipt on 17 June.

Read the source passages

Highlighted passages bear on the selected claim. A connection shows relevance, not automatic proof.

S1 / 3 June 2026

Purchase order

§ 1Component A. Required delivery: 12 June.

§ 2Collection by the buyer’s nominated carrier. The extract contains no termination clause.

S2 / 11 June 2026

Supplier email

§ 1Component A is packed and available for collection today.

§ 2Please confirm the carrier booking.

S3 / 13–17 June 2026

Carrier and receipt log

§ 1Collected 13 June. Received 17 June.

§ 2A request for a Monday slot for the next consignment is recorded; no booking confirmation is attached.

03 / Change the evidence

One new record.
Not a licence to assume everything.

Add either or both fictional follow-ups. A later record may clarify a booking or introduce a conflict; it does not automatically settle blame.

S4 / Fictional follow-up

Next-slot acknowledgement

The carrier confirms a delivery booking for the next consignment on Monday 22 June.

S5 / Fictional follow-up

Disputed collection entry

A second carrier note gives collection as 17 June and disputes the 13 June entry.

04 / Retain the correction

The account changes.
The earlier version stays.

  1. V1 · Initial three-record packetBaseline: S1, S2, S3Authored baseline / 5 September 2026

These versions last for this visit only. Export the record before leaving if you want to retain it.

Revised answer / V1

The records show receipt on 17 June against a required date of 12 June. Collection is recorded on 13 June. The supplier reported readiness on 11 June. The packet does not establish negligence or the cause of the whole delay. The next delivery remains unconfirmed. Obtain the underlying booking and collection records. Check the full contract with an appropriately qualified adviser before considering remedies.

Still open: the cause of the delay and any right to terminate. Operations checks the records; the responsible owner decides, with qualified advice where required.

JSON: original answer, active sources, claim register, revised answer, open questions and the full session history. No upload, account or AI service.

Evidence definitions, assumptions and limits

Evidence status and handling are different questions.

Supported by this packet
An included record supports the claim as specified. The record has not been independently authenticated.
Contradicted by this packet
An included record conflicts with the claim.
Not established by this packet
The required basis is missing. This is not proof that the claim is false.
Conflicting evidence
Included records disagree; their sequence alone does not resolve the conflict.

Handling tells the reader what to do: retain, qualify, obtain evidence or refer for domain review. All assessments here are pre-authored teaching rules. Nothing is automatically inferred from an uploaded document.

This demonstrates an inspectable process, not a measured improvement in hallucination detection. Evaluation would need predefined cases, ordinary-review comparisons, independent adjudication, missed-error and false-flag measures, reviewer agreement and time taken. Failures and disagreement must be retained.

The case in plain language

Start with the question,
not the answer.

Why this case

This research case asks whether a Symbiain audit can make the evidential structure of an AI-generated answer more inspectable by separating its claims, sources, inferences, uncertainty, consequences and correction path.

Why the method helps

A response can be grammatically convincing while combining accurate statements, invented detail, weak attribution and reasonable but unstated inference. A single true-or-false label hides those different failure modes and the decision risk attached to them.

The method, step by step

Four moves.
One visible chain.

Symbiain keeps the moves separate: establish what is observed, relate the conditions, test the possible reading, and state what would require revision.

  1. 01 / Observe

    What can we responsibly say?

    NIST describes generative-AI confabulation as confidently presented erroneous or false content, including output that diverges from its input or contradicts earlier output. OpenAI research argues that training and evaluation can reward guessing rather than an appropriate expression of uncertainty.

  2. 02 / Relate

    What may connect?

    A Symbiain audit would decompose an answer into material claims; retain the exact prompt, model and tool context; connect each claim to source evidence; test whether the source actually entails the claim; record alternatives and uncertainty; and route consequential gaps to accountable human review.

  3. 03 / Test

    What would distinguish the readings?

    Against expert review and a defined evidence set, does the audit improve claim-level error detection, attribution, appropriate abstention and correction without creating a second layer of confident but ungrounded judgement?

  4. 04 / Revise

    What would change the account?

    Revise the audit when expert disagreement, source changes, evaluator error, domain shift, missed contradictions, false alarms or downstream outcomes reveal a weakness in the method.

Why use Symbiain here?

From method
to practical value.

The insight is what becomes visible. The feature is what the method does. The benefit is what the user gains. The value is what can improve in the topic at hand.

  1. 01

    Insight

    The dangerous unit is often not the whole answer but the unsupported claim that passes unnoticed inside a persuasive one.

  2. 02

    Feature

    Maps prompt context, atomic claims, source provenance, entailment, competing explanations, uncertainty and decision consequence in one review field.

  3. 03

    Benefit

    Shows exactly what can be retained, what needs qualification, what must be checked and where a human decision owner is required.

  4. 04

    Value

    Provides a source-bounded assurance process for consequential AI use without claiming to certify truth or eliminate model error.

What to examine

Five conditions
to hold together.

  1. 01Exact prompt, model version, system context, tools and retrieval state
  2. 02Claim decomposition: fact, attribution, quotation, calculation, inference and recommendation
  3. 03Primary-source provenance, date, scope and claim-to-source entailment
  4. 04Contradiction, omission, alternative explanation, uncertainty and appropriate abstention
  5. 05Decision consequence, accountable reviewer, correction, regression test and retained failure

Tensions to hold

Fluent usefulness ↔ warranted confidence

Automated scale ↔ expert adjudication

Evaluator assistance ↔ evaluator hallucination

Important boundary

Research-prototype case only. A Symbiain audit would not prove truth, certify a model, eliminate hallucinations or replace domain experts, primary sources, safety testing, legal duties or accountable human judgement. AI-assisted graders can themselves be wrong; high-stakes claims require source inspection and qualified review. Any reported performance would require a pre-specified benchmark, comparison baseline, documented adjudication and independent replication.

Sources & method

Sources support the stated observations only. All analytical readings remain provisional and should be tested against a defined purpose, scope and evidence base.