Grounding
Clinical assertions must link to allowed transcript or chart evidence, or remain blocked from a ready state.
AI evaluation
Krasyn evaluates the work product, its evidence, and its failure behavior. Model names and aggregate accuracy numbers cannot replace traceability, clinician review, versioned release gates, and a tested rollback.
Current publication boundary
This page publishes the evaluation method and release requirements. It does not publish a clinical-benefit, diagnostic-accuracy, or comparative superiority result. Any future result will identify the evaluated product version, dataset, metric definition, exclusions, and expiry.
A system cannot look safer by producing an empty note, over-warning on every line, or hiding a failure behind a single average.
Clinical assertions must link to allowed transcript or chart evidence, or remain blocked from a ready state.
A release corpus permits no invented medication, diagnosis, numeric value, laterality, risk, or plan presented as fact.
Safety is not measured by returning less. Material supported facts omitted from a draft are measured separately.
We measure clinician changes and time to finalization so a technically grounded but unusable draft cannot pass.
Negation, contradiction, numbers, multiple speakers, poor audio, prompt injection, and missing context have frozen cases.
Latency, cost, upstream failure, retry, rollback, and retention behavior are part of the same release decision.
Each evaluation identifies the input and transcript version, allowed context, template, model deployment, prompt policy, safety rules, timestamps, output hash, and grounding result. The record preserves provenance without storing secrets or hidden model reasoning.
A candidate runs against frozen cases and adversarial cases, is compared with production, receives human review of critical deltas, and enters a monitored cohort only with a kill switch and tested rollback. A failed gate is not averaged away.