OCR systems ranked on the
open 472-page Tibetan OCR benchmark created by
BDRC (Buddhist Digital Resource Center). This is a
fair generalist board: every model is scored on
all pages, and any page a model
skips (including the pages script-specialist models deliberately avoid) counts as a full error.
Lower Character Error Rate (CER) is better. CER is
normalized per page and capped at 100% —
a page can never count as more than fully wrong, so runaway/repetition output can't distort the average.