The patient must not see the mark scheme. The marker must justify each judgement. And the final report must reveal what the system actually knows.
Kent TanFirst published 3 September 2026Revised 9 September 2026
Revised for Certesian with current product distinctions and the NACOSCE follow-up. This is an author-reported engineering note, not a peer-reviewed clinical validation study.
01 / THE ARCHITECTURE
Different stages need different information.
NMCMATE’s actor produces labelled patient and examiner speech in one response. The actor receives scenario information and a code-selected examiner question, but no marking grid. Later, a separate marking call receives the transcript and assessment material. Code derives the report from its judgements.
The patient cannot see the marking grid.
AVAILABLE TO THIS STAGE
Scenario & patient brief
Clinical data
Code-selected examiner question
OUTSIDE THIS STAGE
Marking criteria
Essential flags
Expected answer
NMCMATE: one model response carries labelled patient and examiner speech. A parser preserves speaker attribution.
Why separate the responsibilities?
A simulated patient with access to the rubric can inadvertently coach the learner. Keeping those fields outside the actor interface reduces that failure path. It does not guarantee the patient will never coach or disclose too much: behavioural probes are still needed.
The examiner’s scheduled questions are chosen by code. Speaker labels are parsed into attributed transcript rows, because a parsing error can become a marking error. Criteria are saved with the report so later edits do not silently change the standard a past attempt was assessed against.
We changed the marking prompt to stop it citing reference data as evidence. Select a version to see the trade-off we recorded.
RUN 01PASS
RUN 02PASS
RUN 03PASS
RUN 04PASS
RUN 05PASS
0criteria varied across runs
5reported evidence problems
Shipped prompt. The verdict stayed the same, while citation problems remained.
Author-reported aggregate results from Research #001. One motivating transcript; five runs per prompt. Not a clinical-validity or accuracy benchmark. These experiments were not rerun for this website.
Four prompt configurations were each used to mark the motivating transcript five times. The record reports final verdict sequences, the count of criteria that varied across runs, and evidence problems. Evidence-problem figures are reported counts, not percentages or a per-run accuracy measure.
The baseline produced five passes but retained five evidence problems. Stronger citation instructions removed the recorded evidence problems while introducing verdict changes. All three edits were reverted. This supports a narrow conclusion: those edits did not improve the combined behaviour on that transcript.
The public aggregate record does not include the exact marking-model identifier, the full frozen transcript, raw responses or all execution metadata. Independent reproduction requires those materials. We do not convert these results into a reliability or accuracy percentage.
03 / WHAT DID NOT WORK
REPEATED MARKING
Temperature zero still varied.
In five runs of a 99-turn transcript, judgements varied on two essential criteria and one red flag. All five final verdicts remained FAIL because other grounds for failure remained.
Stable final verdicts can conceal unstable underlying judgements.
MAJORITY VOTING
More calls did not resolve it.
A three-pass voting configuration was tested across five runs. The recorded comparison still found unstable items, at roughly three times the marking-token use. The production configuration was returned to one pass.
This does not prove voting never helps, or that the model performs at chance. Expert-labelled reference answers would be needed to assess correctness.
04 / A SECOND ENGINE, A DIFFERENT FINDING
The first explanation did not simply transfer.
NACOSCE’s 4 September engineering note tested three frozen transcripts and 53 subconditions with gemini-3.8-flash. The negative-condition instability seen in the NMC experiments did not recur in the three non-occurrence conditions tested here.
NACOSCE recorded stability probes · separate from the NMC experiment above
Configuration
Runs per transcript
Unstable subconditions
Before polarity change
5
0 / 53
After polarity change
5
1 / 53
After change, additional run
3
2 / 53
The variation appeared in compound requirements and quality judgements instead. The sample is too small to attribute the change to the polarity fix or establish general reliability. The fix also addressed a separate design problem: demanding a quotation to prove something never occurred.
A second practical finding: reasoning-token headroom
The note records report generation failing at a 4,000-token ceiling when reasoning consumed much of the budget. Subsequent implementation raised the ceiling to 32,000 and recorded reasoning usage separately. That ceiling is an allowance, not the actual cost of each call.
This finding concerns the tested provider configuration. It is not a claim that every reasoning model needs that ceiling.
05 / TWO REPORT POLICIES
Assessment must respect what a session can show.
NMCMATE
A simulated verdict.
The implementation computes a simulated outcome using its own essential-criterion policy and raised red flags. That policy is not presented as the NMC’s official scoring formula or a prediction of an examiner’s result.
NACMATE · IN TESTING
A descriptive report.
Subcondition judgements are combined in code into criterion and competency feedback. The report separates demonstrated, partly demonstrated, not demonstrated and competencies this station did not assess. It gives no score or pass/fail verdict.
Deterministic report rules mean identical inputs produce identical reports. They do not make the model’s judgements deterministic, accurate or clinically validated.
The content pipeline records provenance and verification for generated material. Verification records are tied to a content hash so edited text cannot inherit an approval for different wording. Automated review and safety-veto rules are engineering controls; they do not substitute for qualified educator review.
What this work supports
Implemented information boundaries and report logic
Recorded failure modes and prompt trade-offs
Separate findings from two product implementations
Record provenance: founder-owned Research #001, NMCMATE’s report and marking implementation, and NACOSCE’s engineering note dated 4 September 2026. No new model evaluations or live learner-record access were performed for this revision.
BUILD WITH US
Can this help your educators?
Start with a bounded pilot and agreed review criteria.