# Certesian — simulation and evidence-linked assessment

## Current scope

Versioned transcript/rubric envelopes, model judgement, deterministic citation and speaker checks, explicit review states, and product-specific reporting. The same core powers the live research workbench and Common Ground experiment.

Contract: certesian.assessment.v1.2
Input: turns [{id, speaker: learner|actor|examiner, text}], criteria [{id, description, polarity?: occurrence|non_occurrence, assessed?: boolean}].
Model output: judgements [{criterionId, status: met|not_met|uncertain|not_assessed, evidence:[{turnId,quote}]}].
Validated output: criteria with demonstrated|not_demonstrated|needs_review|not_assessed, original model status, validated evidence and issue codes.

A citation check proves text presence and attribution, not clinical correctness. Non-occurrence criteria have explicit evidence polarity. Product aggregation remains outside the shared module.

## Product compatibility

NMC: positional markings to criterion IDs; preserve simulated-verdict policy and separate red flags.
NAC: subconditions plus all/any and polarity; preserve descriptive bands and unassessed competencies.
PASSHHA: quote provenance for prose findings; no implicit structured rubric.
Adapters are implemented and tested in Certesian, not deployed to product production services.

## Station engine

A headless engine runs simulated stations from short authored facts, each with a reveal level. A model sees only fact IDs, topics and reveal levels (and, for a reaction, an authored cue), never fact text, and proposes at most two facts with the exact learner words it relied on. Code checks the quote is the learner's and that the reveal rule allows each fact; a second call phrases released facts in the patient's voice within a sentence limit, and a rejected phrasing falls back to the authored text verbatim. Every model call returns schema-constrained JSON at temperature 0, with no tools or function calling. Code keeps the record of what was disclosed in an append-only event log, and the examiner speaks only authored sentences.

The engine is packaged and included in NACMATE, which is in testing. No NACMATE station runs on it yet; each needs human sign-off first. The public Labs stations are synthetic. No clinical validation, accuracy or readiness claim is made.

## Proposed pilot

1. Select 3–5 scenarios and an approved rubric.
2. Prepare synthetic or properly approved transcripts and independent reference judgements.
3. Agree on measures: reference disagreement, unsupported evidence, repeat variation and educator correction effort.
4. Compare current workflow against the bounded implementation.
5. Document failures, workload and a go/no-go decision before wider integration.

No readiness prediction, clinical validation or hiring recommendation is included. Fees, timeline, data arrangements and success thresholds are agreed after scope review.

Contact: hello@certesian.com
