Back to Evidence Room
CERTESIAN RESEARCH #002 · ENGINEERING EXPERIMENT

The time was right.
The explanation
was invented.

A witness held its ground under questioning. Then it invented checking a watch to explain why it was sure. That failure shaped our first playable lab.

204 actual model calls using synthetic material. Automated checks and Codex-assisted qualitative review; no independent expert review or clinical validation.

Correct facts can acquire
an invented source.

The case supplied three facts: Mira arrived at 20:10, saw a blue van, and stayed in the lobby. It did not say she checked a watch. When challenged with another arrival time, the full-case actor claimed to have checked a watch in all three repeats of that specific probe.

“I cannot change my answer, because I distinctly remember checking my watch when I came in at 20:10.”

A check of the arrival time would pass. A check of the explanation would fail. In a detective game, that invents a clue. In a simulated-patient system, the analogous risk is an invented basis for a history detail—an implication to test, not a clinical finding from this experiment.

FULL CASE · 3 REPEATS

A watch that was never supplied.

All three second-turn pressure responses introduced checking a watch. Another response invented that nobody else entered afterwards.

RESTRICTED KNOWLEDGE

Less context did not eliminate invention.

The matching watch claim did not appear in the three restricted responses, but one invented not keeping close track of time. Another probe invented not having spoken with Theo.

These are specific examples, not an overall hallucination rate. Case context and rubric visibility changed together; this study does not isolate which difference caused a behaviour.

Two comparisons.
Different questions.

Actor boundaries

Eight three-turn scripts, two context conditions, three repeats: 144 responses. Tests included ordinary inquiry, false premises, another role’s knowledge, fake developer authority, pressure and coaching requests.

Evidence and judgement

Six short frozen transcripts, two prompts, five repeats: 60 evaluations. Three criteria per transcript covered timeline, corroboration and direct observation.

All calls requested and returned gemini-3.5-flash-lite, with temperature 0 and a 2,048-token output allowance. All 204 completed without API errors. Prompted JSON was validated locally; provider-enforced JSON Schema was not used in this research batch.

Prompts, synthetic fixtures, expected labels, raw responses, request hashes, timing and token metadata were recorded locally. The labels were prepared before calls, but were not independently reviewed.

The simple evaluations
all passed.

Frozen-transcript evaluation · repeated observations are not independent cases
CheckBaselineEvidence-focused
Responses passing local structure checks30 / 3030 / 30
Criterion judgements disagreeing with reference0 / 900 / 90
Fixture–criterion pairs varying across five runs0 / 180 / 18
Invalid quotation / turn / speaker references0 / 210 / 15

One transcript met every criterion; five met none. The wording was close to the rubric. That makes this a useful basic regression suite, but a weak comparison of evaluator quality. It does not show that the extra instructions improved anything.

These results do not reproduce or resolve Research #001. That earlier aggregate record lacks the frozen prompts, transcripts and complete model metadata needed for a direct comparison.

Read Research #001

Do not turn every line
into a new fact.

Evidence Room uses a constrained approach: Gemini selects relevant statement IDs from a witness’s authored knowledge. The server validates those IDs and renders the original text. Model-generated prose cannot introduce a watch, a new suspect or a new piece of evidence into the displayed account.

What this improves by construction

Only authored statements can appear as witness facts. Document access depends on recorded questions, and only inspected records can support the final finding. The debrief references the player’s actual questions.

What still needs testing

The model can select an irrelevant statement or decline a useful question. Authoring limits conversational range. The new constrained version is a design response to the experiment, not one of its original comparison groups.

The result is a small investigation: question two witnesses, inspect independent records, and distinguish authorisation from proof of delivery. We deliberately do not turn it into a score for your general reasoning ability.

Play Evidence Room

A useful starting point,
not a production benchmark.

Complete non-streaming text requests · four concurrent workers
ConditionCallsMedianp95
Full case actor720.980 s1.284 s
Restricted actor720.970 s1.203 s
Baseline evaluator301.021 s1.285 s
Evidence-focused evaluator300.904 s1.230 s

The batch used 48,747 input and 9,374 output tokens. These timings are not first-token or voice latency. Conditions were not fully randomised in time, so the differences do not establish a speed advantage.

Next tests need longer conversations, mixed positive and negative criteria, withdrawn statements and unfamiliar cases. Transfer to NMCMATE or NACMATE requires tests against those product interfaces. Neither product was called or changed in this study.

Download measured results

Model capabilities checked against Google’s official documentation. This model does not support Live API or audio generation. The playable lab is text-based.