New: OpenClinic now supports Orthodontics & Oral Surgery — book a live demo
← All articles

How we evaluate clinical AI before it touches a chart

No model ships to a practice until it survives our gauntlet: accuracy rubrics, hallucination hunts, and practicing dentists grading every draft.


Why "sounds right" isn't good enough

Fluency is cheap; correctness is everything. A note that reads beautifully but invents a filling on #30 is worse than no note at all. So we don't grade our AI on prose — we grade it on clinical facts, tooth by tooth, code by code.

The three gates

  • Rubric-scored accuracy. Every draft is checked against a rubric: correct teeth and surfaces? Findings supported by the conversation? Codes matched to documentation? A draft must clear strict thresholds before it can ship.
  • Hallucination hunting. We deliberately probe edge cases — noisy operatories, overlapping voices, rare procedures — and measure invented content at near-zero tolerance. Anything the model can't support, it must omit.
  • Clinician-in-the-loop review. Practicing dentists grade drafts blind, and nothing reaches customers without their sign-off. Their corrections feed directly back into the next evaluation round.

Evaluation never ends

Shipping is the starting line. Every model update re-runs the full gauntlet, and production drafts are continuously sampled and scored. If quality ever regresses, we know within days — not from a complaint, but from the numbers.

Trust in clinical AI isn't claimed. It's measured, published internally, and re-earned with every release. That's the standard we hold ourselves to before OpenClinic touches a single chart.


Bring your practice home on time.

See OpenClinic on your own cases in a 20-minute live demo.