S1 / 3 June 2026
Purchase order
§ 1Component A. Required delivery: 12 June.
§ 2Collection by the buyer’s nominated carrier. The extract contains no termination clause.
Casebook / 24
Trace its claims to evidence, inspect what remains uncertain, and see what changes when a new record arrives.
Fluency is not provenance.

This page demonstrates a way of structuring inquiry. It does not constitute professional advice, verified intelligence, a recommendation or a conclusion about any person, organisation or transaction.
Try the method / fictional teaching case
Select a claim. Read its evidence. Then introduce a new record and see what changes—and what stays open. These are authored examples, not live AI outputs or client records.
01 / Select part of the original answer
The task: explain the delay and advise the decision owner from the supplied records. The confident answer above contains errors.
02 / Claim C2 · Version 1
17 June is the receipt date. Collection is recorded on 13 June. The initial answer has substituted arrival for dispatch.
Highlighted passages bear on the selected claim. A connection shows relevance, not automatic proof.
S1 / 3 June 2026
§ 1Component A. Required delivery: 12 June.
§ 2Collection by the buyer’s nominated carrier. The extract contains no termination clause.
S2 / 11 June 2026
§ 1Component A is packed and available for collection today.
§ 2Please confirm the carrier booking.
S3 / 13–17 June 2026
§ 1Collected 13 June. Received 17 June.
§ 2A request for a Monday slot for the next consignment is recorded; no booking confirmation is attached.
03 / Change the evidence
Add either or both fictional follow-ups. A later record may clarify a booking or introduce a conflict; it does not automatically settle blame.
The carrier confirms a delivery booking for the next consignment on Monday 22 June.
A second carrier note gives collection as 17 June and disputes the 13 June entry.
04 / Retain the correction
These versions last for this visit only. Export the record before leaving if you want to retain it.
Revised answer / V1
The records show receipt on 17 June against a required date of 12 June. Collection is recorded on 13 June. The supplier reported readiness on 11 June. The packet does not establish negligence or the cause of the whole delay. The next delivery remains unconfirmed. Obtain the underlying booking and collection records. Check the full contract with an appropriately qualified adviser before considering remedies.
Still open: the cause of the delay and any right to terminate. Operations checks the records; the responsible owner decides, with qualified advice where required.
JSON: original answer, active sources, claim register, revised answer, open questions and the full session history. No upload, account or AI service.
Handling tells the reader what to do: retain, qualify, obtain evidence or refer for domain review. All assessments here are pre-authored teaching rules. Nothing is automatically inferred from an uploaded document.
This demonstrates an inspectable process, not a measured improvement in hallucination detection. Evaluation would need predefined cases, ordinary-review comparisons, independent adjudication, missed-error and false-flag measures, reviewer agreement and time taken. Failures and disagreement must be retained.
The case in plain language
This research case asks whether a Symbiain audit can make the evidential structure of an AI-generated answer more inspectable by separating its claims, sources, inferences, uncertainty, consequences and correction path.
A response can be grammatically convincing while combining accurate statements, invented detail, weak attribution and reasonable but unstated inference. A single true-or-false label hides those different failure modes and the decision risk attached to them.
The method, step by step
Symbiain keeps the moves separate: establish what is observed, relate the conditions, test the possible reading, and state what would require revision.
NIST describes generative-AI confabulation as confidently presented erroneous or false content, including output that diverges from its input or contradicts earlier output. OpenAI research argues that training and evaluation can reward guessing rather than an appropriate expression of uncertainty.
A Symbiain audit would decompose an answer into material claims; retain the exact prompt, model and tool context; connect each claim to source evidence; test whether the source actually entails the claim; record alternatives and uncertainty; and route consequential gaps to accountable human review.
Against expert review and a defined evidence set, does the audit improve claim-level error detection, attribution, appropriate abstention and correction without creating a second layer of confident but ungrounded judgement?
Revise the audit when expert disagreement, source changes, evaluator error, domain shift, missed contradictions, false alarms or downstream outcomes reveal a weakness in the method.
Why use Symbiain here?
The insight is what becomes visible. The feature is what the method does. The benefit is what the user gains. The value is what can improve in the topic at hand.
The dangerous unit is often not the whole answer but the unsupported claim that passes unnoticed inside a persuasive one.
Maps prompt context, atomic claims, source provenance, entailment, competing explanations, uncertainty and decision consequence in one review field.
Shows exactly what can be retained, what needs qualification, what must be checked and where a human decision owner is required.
Provides a source-bounded assurance process for consequential AI use without claiming to certify truth or eliminate model error.
What to examine
Tensions to hold
Fluent usefulness ↔ warranted confidence
Automated scale ↔ expert adjudication
Evaluator assistance ↔ evaluator hallucination
Important boundary
Research-prototype case only. A Symbiain audit would not prove truth, certify a model, eliminate hallucinations or replace domain experts, primary sources, safety testing, legal duties or accountable human judgement. AI-assisted graders can themselves be wrong; high-stakes claims require source inspection and qualified review. Any reported performance would require a pre-specified benchmark, comparison baseline, documented adjudication and independent replication.
Sources & method
Sources support the stated observations only. All analytical readings remain provisional and should be tested against a defined purpose, scope and evidence base.