THE CERTESIAN TECHNOLOGY PROGRAMME

Understand performance.
Start with evidence.

We are building reusable simulation and assessment components through professional products, public experiments and tests we can inspect.

A judgement needs a traceable path.

Versioned transcriptCriterion judgementEvidence checksProduct report policy
SIMULATION

Control the world.

Keep role knowledge, progression and explicit memory in the application. Labs tests constrained characters and authored case facts.

ASSESSMENT

Keep the source.

Validate turn IDs, exact quotations and speaker attribution. Preserve unassessed, absent and unresolved states.

REPORTING

Respect the context.

A common evidence format does not require a common score. Existing products retain their own reporting rules.

Try the evaluation core

Three ways to develop
the same questions.

EnvironmentWhat it contributesCurrent relationship
NMCMATEProfessional simulation and essential-criterion reporting.Source contract inspected; synthetic checks execute its pure verdict function.
NACMATESubcondition judgements, absence rules and descriptive bands.In testing. Source contract inspected; pure reporting functions tested with the adapter. The station engine is packaged and vendored into NACMATE, and NACMATE checks report quotes against candidate turns in its own code. No NACMATE station runs on the engine yet.
PASSHHAInterview dialogue and prose debrief findings.Source-shaped adapter tests; no fabricated criterion coverage.
Certesian LabsAccessible investigation, communication and character experiences, and three clinical stations.Live experiments with synthetic content. Common Ground uses the shared evaluator; the three stations run on the station engine.

These relationships describe inspected code and current experiments. They do not mean every product already runs one shared engine, or that product user data has been pooled.

Implemented and testable

  • Versioned transcript and criterion contract
  • Evidence provenance checks and review states
  • Three source-compatible product adapters
  • A station engine: authored facts, code-side disclosure rules and a replayable event log
  • Six playable public experiments, three of them stations on the engine
  • Synthetic fixtures and recorded Gemini runs

Further validation required

  • Educator agreement on real assessment material
  • Longitudinal competency and readiness modelling
  • Production adoption of the adapters and the station engine
  • External pilot outcomes and commercial demand
  • Multi-tenant API operations and support

The long-term research direction is human performance intelligence. Today’s deliverable is a narrower evidence layer, not a validated prediction of someone’s future performance.

The patient can only say
what was written.

Every station is written as short authored facts, each with a reveal level: said to an open question, said when its topic is asked, or said only to a precise question.

  • A model sees each fact’s ID, topics and reveal level, and for a reaction the author’s cue for when it applies. It never sees the fact text.
  • It proposes at most two facts, with the exact words of the learner’s that it relied on.
  • Code checks that those words are the learner’s and that the reveal rule allows each fact. Anything else is not released.
  • A second call phrases the released facts in the patient’s voice within a sentence limit. Code rejects a phrasing that is empty, runs over the limit, changes a number or a negation, or uses any word of three or more letters that is not in the authored text, and the patient then says the authored text verbatim. Very short words and word order are not checked.
  • No model call has tools or function calling. Every call returns schema-constrained JSON at temperature 0.
  • The record of what was disclosed is kept by code, in an append-only event log.

Fixes from reading
test transcripts.

  • “Tell me more” gets the next part of the story, not “I already told you”.
  • A red-flag screen gets an answer to every part the candidate asks about.
  • The patient says who she is when asked.
  • She never denies something she was not asked about precisely.
  • When the candidate explains a plan, she can give an authored reaction to it, one a turn.
  • The examiner speaks only fixed, authored sentences.

These are engineering fixes found by reading test transcripts. They are not measured improvements.

Where the station engine
stands today.

  • The engine is packaged and included in NACMATE, which is in testing. No NACMATE station is live on it yet, and each one needs human sign-off first.
  • No clinical validation, accuracy, readiness prediction or learner outcomes are claimed.
  • The Labs stations and the test fixtures are synthetic.
  • Research #004 is paused until enough real attempts exist. Its rehearsal ran on model-written transcripts and is not published.

Bring a rubric.
Choose one question to test.

A first pilot can compare evidence attribution, judgement consistency and educator correction effort on an agreed set of synthetic or approved scenarios. Agree on measures before building, and keep human review in the loop.

A self-serve commercial API and readiness prediction are not currently offered. The workbench is a bounded experiment.