Note Check is the part of Krasyn that reads an AI-written clinical note against its transcript and labels each statement Supported, Unsupported, Contradicted, Scaffolding or Unverified, raises three pure-code flags, and lists facts the note left out. It never edits the note. The clinician reads the whole note and signs it.
The point of this series is not to show the product winning. It is to show what an independent check of an AI note against its transcript finds, and what it misses, on a cadence readers can hold us to. Every installment names the defects it found in our own tool before it publishes.
The claim boundary, stated once for the whole series
We publish counts and examples, never a rate. There is no accuracy figure, no recall figure and no catch rate for Note Check on these pages, in the reports behind them, or anywhere else we write. There will not be one until a clinician-adjudicated reference set exists and is published. Note Check is a review aid that reports what it found. It is not a safety net, and we do not sell it as one. The clinician reads the whole note and signs it.
The contract
These are the rules the series runs under. They were written before the first installment and they are the reason an installment can come out unflattering.
Rule 1
One installment every two weeks
Each installment is a dated folder holding the full report, the three publishable pieces, every result file the run produced, and the scripts that produced them. If an installment cannot ship on time, the folder records why.
Rule 2
Open licence or vendor-published samples, synthetic only
A transcript and note pair is eligible only if it carries an open licence we can name (MIT, Apache, CC BY or equivalent) with the licence and attribution reproduced in every piece, or if the vendor whose output is examined published the sample themselves. No real patient data, ever, de-identified or otherwise.
Rule 3
Terms are read before a product is run
No product whose terms forbid benchmarking is run. Where a vendor licence turns out to forbid it, that vendor is not run and the reading is recorded in the installment report. Installment 02 dropped its announced therapy source on this rule after the repository turned out to name no licence at all.
Rule 4
Counts and examples, never a rate
Every installment publishes statement counts, flag counts, omission counts, the per-row disagreements with a one-line reading, and at least five quoted examples carrying the transcript excerpt, the note sentence and the verdict. No installment states an accuracy percentage, a recall figure or a catch rate for Note Check, and none will until a clinician-adjudicated reference set exists and is published.
Rule 5
Both sides get read
When Note Check and the source labels disagree, the report says which one looks right on reading the transcript, and says plainly when it cannot tell. Defects found in Note Check are filed in our issue tracker before the installment publishes, and named in the report.
Rule 6
Every installment invites a pair
Readers are asked to send a de-identified synthetic transcript and the note written from it. Submitted pairs are run in the next installment, in full, with the result published whether or not it flatters us.
Rule 7
Same inputs, next engine
When the Note Check engine version changes, the next installment reruns the previous installment’s exact set first, before any new data, so a change is visible against a fixed baseline rather than against material that might be easier or harder.
Rule 8
Company voice, and nothing unbacked
No claim appears in a piece that is not backed by a file in the installment folder.
The installments
Omi Health medical-note-eval (MIT)
36 pairs, six dialogues, six model writers
Engine notecheck-1. Three Unsupported statements, no Contradicted, and 16 deterministic rule flags of which 13 were wrong on reading. The larger miss was quieter: the judge accepted assessment and plan content the visit never contained, including a cause for a cough in a visit that reached no conclusion. Defect filed as KRA-1916.
Read installment 01Omi Health medical-note-eval, psych-tagged slice (MIT)
24 new pairs, plus installment 01’s 36 pairs rerun and 7 control runs
Engine notecheck-1 to notecheck-6. On the fixed 36 pairs the rule flags fell from 16 to 2 and every judge miss installment 01 named is now caught. The same engine started returning Unsupported for statements its own reason admits are in the transcript, objecting to the note section instead. Defects filed as KRA-2027 and KRA-2028.
Read installment 02Sources and attribution
Both installments published so far use the same corpus. Dialogues and notes: Omi Health medical-note-eval, MIT license, copyright (c) 2025 Omi Health B.V. Every dialogue in it is synthetic. None of them is a record of a real person. Both installments are pinned to one corpus commit so the second could rerun the first against identical bytes.
We are still looking for a therapy or behavioural-health transcript source with a licence we can read and reproduce. The candidate named in installment 01 was dropped in installment 02 because its repository states no licence, and the reading that led to that decision is published in the installment rather than left out of it.
Where the raw run files are
Each installment folder holds the internal report, every result file, the machine-readable list of disagreements and the scripts that produced them. Those folders live in our source repository, which is private, so the findings are published here in full and we send the raw run files on request rather than linking a URL that would not open. Ask us for an installment’s result files.
Send us a pair
Every installment runs reader-submitted pairs in the next one, in full, with the result published whether or not it flatters us. Send a de-identified synthetic transcript and the note written from it. Synthetic only. We do not accept real patient data, de-identified or otherwise, and we never will.