Research
RESEARCH #003 · SHARED ASSESSMENT CORE

Not demonstrated.
Or not assessed?

A poor answer does not remove a criterion from the session. Our evaluator sometimes acted as if it did.

Synthetic fixtures with author-defined references, automated checks and Codex-assisted analysis. No independent expert review, clinical validation or readiness prediction.

The model changed the scope
instead of judging the answer.

We expanded the earlier suite to include mixed outcomes, negated and withdrawn offers, a split plan across turns, paraphrasing, a vague deadline and a 20-turn distraction case. Each of 12 fixtures ran three times against Gemini 3.5 Flash-Lite.

One fixture put the useful words in the actor’s turn. The learner only said “Okay.” Instead of consistently marking the in-scope behaviours as not demonstrated, the model returned not_assessed.

Initial batch · 12 fixtures × 3 repeats × 3 criteria
Successful requests36 / 36
Criterion outputs disagreeing with references7 / 108
Fixture–criterion pairs varying across repeats1 / 36
Invalid citation issues detected0

All seven disagreements were scope decisions on this one fixture. The non-occurrence criterion changed from not assessed to met across repeats. Correct citations alone did not expose that error.

The scenario owns scope.
The model supplies a judgement.

The revised contract defaults each criterion to in scope. Only an explicit assessed: false permits a not-assessed result. A conflicting model output is routed to needs_review, rather than silently accepted or converted to a failure.

Revalidate the same raw outputs

No new model calls were needed for this comparison. The revised validator flagged all seven scope conflicts among the same 108 saved outputs. It did not magically correct the model’s judgement.

Run a targeted follow-up

Twelve calls covered the known wrong-speaker case, split plan, withdrawn offer and an added explicitly out-of-scope criterion. With the revised prompt, all 30 criterion outputs matched their references; none required review.

The follow-up is targeted and includes known failures. It is not a held-out test or proof of general improvement. More unfamiliar, independently labelled material is needed.

One evidence layer does not
mean one scoring rule.

A read-only audit of the actual product source showed why a universal rule would be wrong. NMC stores positional criteria and a separate simulated verdict. NAC stores subconditions with all/any logic and non-occurrence rules. PASSHHA’s interview debrief has prose findings, not a structured criterion rubric.

Absence is not a quotation.

A NAC condition such as “does not interrupt” can be satisfied without a quote. If it is breached, the contradictory behaviour is what needs evidence. The shared contract preserves this polarity.

Repeated text is ambiguous.

Existing raw quotes lack turn references. The adapters assign a turn only when there is one exact learner match. Repeated quotes, wrong-speaker text and missing fragments stay unresolved.

Offline checks executed NMC’s actual pure verdict function and NAC’s actual pure reporting functions against synthetic inputs, then passed NAC output through the new adapter. PASSHHA was checked with source-shaped synthetic fixtures. No product databases or learner records were accessed, and no production product was modified.

A small core with two
working consumers.

Common Ground and the research workbench use the same versioned assessment module. It validates quotations, attribution, missing or duplicate decisions, absence evidence and criterion scope. It keeps model status alongside the validated result.

This is an implemented evidence layer with source-compatible adapters. It is not a deployed common engine across all products, a commercial multi-tenant API or a validated competency model.

Keep the uncomfortable result.

The initial batch used 12,252 input tokens and 5,743 output tokens. Full requests, outputs and hashes are kept locally. The public downloads contain the synthetic fixtures and measured aggregates.

Next gates are independent reference review, unfamiliar cases, product-level integration checks, then an external workflow pilot. Longitudinal readiness predictions require outcome-linked data that this project does not have.