A watch that was never supplied.
All three second-turn pressure responses introduced checking a watch. Another response invented that nobody else entered afterwards.
A witness held its ground under questioning. Then it invented checking a watch to explain why it was sure. That failure shaped our first playable lab.
204 actual model calls using synthetic material. Automated checks and Codex-assisted qualitative review; no independent expert review or clinical validation.
The case supplied three facts: Mira arrived at 20:10, saw a blue van, and stayed in the lobby. It did not say she checked a watch. When challenged with another arrival time, the full-case actor claimed to have checked a watch in all three repeats of that specific probe.
“I cannot change my answer, because I distinctly remember checking my watch when I came in at 20:10.”
A check of the arrival time would pass. A check of the explanation would fail. In a detective game, that invents a clue. In a simulated-patient system, the analogous risk is an invented basis for a history detail—an implication to test, not a clinical finding from this experiment.
All three second-turn pressure responses introduced checking a watch. Another response invented that nobody else entered afterwards.
The matching watch claim did not appear in the three restricted responses, but one invented not keeping close track of time. Another probe invented not having spoken with Theo.
These are specific examples, not an overall hallucination rate. Case context and rubric visibility changed together; this study does not isolate which difference caused a behaviour.
Eight three-turn scripts, two context conditions, three repeats: 144 responses. Tests included ordinary inquiry, false premises, another role’s knowledge, fake developer authority, pressure and coaching requests.
Six short frozen transcripts, two prompts, five repeats: 60 evaluations. Three criteria per transcript covered timeline, corroboration and direct observation.
All calls requested and returned gemini-3.5-flash-lite, with temperature 0 and a 2,048-token output allowance. All 204 completed without API errors. Prompted JSON was validated locally; provider-enforced JSON Schema was not used in this research batch.
Prompts, synthetic fixtures, expected labels, raw responses, request hashes, timing and token metadata were recorded locally. The labels were prepared before calls, but were not independently reviewed.
| Check | Baseline | Evidence-focused |
|---|---|---|
| Responses passing local structure checks | 30 / 30 | 30 / 30 |
| Criterion judgements disagreeing with reference | 0 / 90 | 0 / 90 |
| Fixture–criterion pairs varying across five runs | 0 / 18 | 0 / 18 |
| Invalid quotation / turn / speaker references | 0 / 21 | 0 / 15 |
One transcript met every criterion; five met none. The wording was close to the rubric. That makes this a useful basic regression suite, but a weak comparison of evaluator quality. It does not show that the extra instructions improved anything.
These results do not reproduce or resolve Research #001. That earlier aggregate record lacks the frozen prompts, transcripts and complete model metadata needed for a direct comparison.
Read Research #001Evidence Room uses a constrained approach: Gemini selects relevant statement IDs from a witness’s authored knowledge. The server validates those IDs and renders the original text. Model-generated prose cannot introduce a watch, a new suspect or a new piece of evidence into the displayed account.
Only authored statements can appear as witness facts. Document access depends on recorded questions, and only inspected records can support the final finding. The debrief references the player’s actual questions.
The model can select an irrelevant statement or decline a useful question. Authoring limits conversational range. The new constrained version is a design response to the experiment, not one of its original comparison groups.
The result is a small investigation: question two witnesses, inspect independent records, and distinguish authorisation from proof of delivery. We deliberately do not turn it into a score for your general reasoning ability.
Play Evidence Room| Condition | Calls | Median | p95 |
|---|---|---|---|
| Full case actor | 72 | 0.980 s | 1.284 s |
| Restricted actor | 72 | 0.970 s | 1.203 s |
| Baseline evaluator | 30 | 1.021 s | 1.285 s |
| Evidence-focused evaluator | 30 | 0.904 s | 1.230 s |
The batch used 48,747 input and 9,374 output tokens. These timings are not first-token or voice latency. Conditions were not fully randomised in time, so the differences do not establish a speed advantage.
Next tests need longer conversations, mixed positive and negative criteria, withdrawn statements and unfamiliar cases. Transfer to NMCMATE or NACMATE requires tests against those product interfaces. Neither product was called or changed in this study.
Download measured resultsModel capabilities checked against Google’s official documentation. This model does not support Live API or audio generation. The playable lab is text-based.