Control the world.
Keep role knowledge, progression and explicit memory in the application. Labs tests constrained characters and authored case facts.
We are building reusable simulation and assessment components through professional products, public experiments and tests we can inspect.
Keep role knowledge, progression and explicit memory in the application. Labs tests constrained characters and authored case facts.
Validate turn IDs, exact quotations and speaker attribution. Preserve unassessed, absent and unresolved states.
A common evidence format does not require a common score. Existing products retain their own reporting rules.
| Environment | What it contributes | Current relationship |
|---|---|---|
| NMCMATE | Professional simulation and essential-criterion reporting. | Source contract inspected; synthetic checks execute its pure verdict function. |
| NACMATE | Subcondition judgements, absence rules and descriptive bands. | In testing. Source contract inspected; pure reporting functions tested with the adapter. The station engine is packaged and vendored into NACMATE, and NACMATE checks report quotes against candidate turns in its own code. No NACMATE station runs on the engine yet. |
| PASSHHA | Interview dialogue and prose debrief findings. | Source-shaped adapter tests; no fabricated criterion coverage. |
| Certesian Labs | Accessible investigation, communication and character experiences, and three clinical stations. | Live experiments with synthetic content. Common Ground uses the shared evaluator; the three stations run on the station engine. |
These relationships describe inspected code and current experiments. They do not mean every product already runs one shared engine, or that product user data has been pooled.
The long-term research direction is human performance intelligence. Today’s deliverable is a narrower evidence layer, not a validated prediction of someone’s future performance.
Every station is written as short authored facts, each with a reveal level: said to an open question, said when its topic is asked, or said only to a precise question.
These are engineering fixes found by reading test transcripts. They are not measured improvements.
A first pilot can compare evidence attribution, judgement consistency and educator correction effort on an agreed set of synthetic or approved scenarios. Agree on measures before building, and keep human review in the loop.
A self-serve commercial API and readiness prediction are not currently offered. The workbench is a bounded experiment.