Skip to main content

AI evaluation

A Clinical Draft Should Show Why It Deserves Trust

Krasyn evaluates the work product, its evidence, and its failure behavior. Model names and aggregate accuracy numbers cannot replace traceability, clinician review, versioned release gates, and a tested rollback.

Current publication boundary

This page publishes the evaluation method and release requirements. It does not publish a clinical-benefit, diagnostic-accuracy, or comparative superiority result. Any future result will identify the evaluated product version, dataset, metric definition, exclusions, and expiry.

The release scorecard is deliberately multi-dimensional

A system cannot look safer by producing an empty note, over-warning on every line, or hiding a failure behind a single average.

Grounding

Clinical assertions must link to allowed transcript or chart evidence, or remain blocked from a ready state.

Critical invention

A release corpus permits no invented medication, diagnosis, numeric value, laterality, risk, or plan presented as fact.

Omission

Safety is not measured by returning less. Material supported facts omitted from a draft are measured separately.

Edit burden

We measure clinician changes and time to finalization so a technically grounded but unusable draft cannot pass.

Difficult inputs

Negation, contradiction, numbers, multiple speakers, poor audio, prompt injection, and missing context have frozen cases.

Operations

Latency, cost, upstream failure, retry, rollback, and retention behavior are part of the same release decision.

What makes a run reproducible

Each evaluation identifies the input and transcript version, allowed context, template, model deployment, prompt policy, safety rules, timestamps, output hash, and grounding result. The record preserves provenance without storing secrets or hidden model reasoning.

What makes a rollout governable

A candidate runs against frozen cases and adversarial cases, is compared with production, receives human review of critical deltas, and enters a monitored cohort only with a kill switch and tested rollback. A failed gate is not averaged away.

The clinician remains the authorization boundary

  • The draft remains visibly distinct from a signed clinical note.
  • Transcript or chart evidence is available at the point of review when supporting evidence exists.
  • Clinician edits remain distinguishable from generated text.
  • No model signs, prescribes, orders, sends, or writes to a chart on its own.
  • Degraded states explain what failed, what remains safe, and the next safe action.