How we evaluate clinical AI before it touches a chart
No model ships to a practice until it survives our gauntlet: accuracy rubrics, hallucination hunts, and practicing dentists grading every draft.
Why "sounds right" isn't good enough
Fluency is cheap; correctness is everything. A note that reads beautifully but invents a filling on #30 is worse than no note at all. So we don't grade our AI on prose — we grade it on clinical facts, tooth by tooth, code by code.
The three gates
- Rubric-scored accuracy. Every draft is checked against a rubric: correct teeth and surfaces? Findings supported by the conversation? Codes matched to documentation? A draft must clear strict thresholds before it can ship.
- Hallucination hunting. We deliberately probe edge cases — noisy operatories, overlapping voices, rare procedures — and measure invented content at near-zero tolerance. Anything the model can't support, it must omit.
- Clinician-in-the-loop review. Practicing dentists grade drafts blind, and nothing reaches customers without their sign-off. Their corrections feed directly back into the next evaluation round.
Evaluation never ends
Shipping is the starting line. Every model update re-runs the full gauntlet, and production drafts are continuously sampled and scored. If quality ever regresses, we know within days — not from a complaint, but from the numbers.
Trust in clinical AI isn't claimed. It's measured, published internally, and re-earned with every release. That's the standard we hold ourselves to before OpenClinic touches a single chart.