Revalidate the same raw outputs
No new model calls were needed for this comparison. The revised validator flagged all seven scope conflicts among the same 108 saved outputs. It did not magically correct the model’s judgement.
A poor answer does not remove a criterion from the session. Our evaluator sometimes acted as if it did.
Synthetic fixtures with author-defined references, automated checks and Codex-assisted analysis. No independent expert review, clinical validation or readiness prediction.
We expanded the earlier suite to include mixed outcomes, negated and withdrawn offers, a split plan across turns, paraphrasing, a vague deadline and a 20-turn distraction case. Each of 12 fixtures ran three times against Gemini 3.5 Flash-Lite.
One fixture put the useful words in the actor’s turn. The learner only said “Okay.” Instead of consistently marking the in-scope behaviours as not demonstrated, the model returned not_assessed.
| Successful requests | 36 / 36 |
|---|---|
| Criterion outputs disagreeing with references | 7 / 108 |
| Fixture–criterion pairs varying across repeats | 1 / 36 |
| Invalid citation issues detected | 0 |
All seven disagreements were scope decisions on this one fixture. The non-occurrence criterion changed from not assessed to met across repeats. Correct citations alone did not expose that error.
The revised contract defaults each criterion to in scope. Only an explicit assessed: false permits a not-assessed result. A conflicting model output is routed to needs_review, rather than silently accepted or converted to a failure.
No new model calls were needed for this comparison. The revised validator flagged all seven scope conflicts among the same 108 saved outputs. It did not magically correct the model’s judgement.
Twelve calls covered the known wrong-speaker case, split plan, withdrawn offer and an added explicitly out-of-scope criterion. With the revised prompt, all 30 criterion outputs matched their references; none required review.
The follow-up is targeted and includes known failures. It is not a held-out test or proof of general improvement. More unfamiliar, independently labelled material is needed.
A read-only audit of the actual product source showed why a universal rule would be wrong. NMC stores positional criteria and a separate simulated verdict. NAC stores subconditions with all/any logic and non-occurrence rules. PASSHHA’s interview debrief has prose findings, not a structured criterion rubric.
A NAC condition such as “does not interrupt” can be satisfied without a quote. If it is breached, the contradictory behaviour is what needs evidence. The shared contract preserves this polarity.
Existing raw quotes lack turn references. The adapters assign a turn only when there is one exact learner match. Repeated quotes, wrong-speaker text and missing fragments stay unresolved.
Offline checks executed NMC’s actual pure verdict function and NAC’s actual pure reporting functions against synthetic inputs, then passed NAC output through the new adapter. PASSHHA was checked with source-shaped synthetic fixtures. No product databases or learner records were accessed, and no production product was modified.
Common Ground and the research workbench use the same versioned assessment module. It validates quotations, attribution, missing or duplicate decisions, absence evidence and criterion scope. It keeps model status alongside the validated result.
This is an implemented evidence layer with source-compatible adapters. It is not a deployed common engine across all products, a commercial multi-tenant API or a validated competency model.
The initial batch used 12,252 input tokens and 5,743 output tokens. Full requests, outputs and hashes are kept locally. The public downloads contain the synthetic fixtures and measured aggregates.
Next gates are independent reference review, unfamiliar cases, product-level integration checks, then an external workflow pilot. Longitudinal readiness predictions require outcome-linked data that this project does not have.