Tencent's Hunyuan and SSV Digital Culture Lab, in collaboration with the Institute of Information Engineering at the Chinese Academy of Sciences, have released Chronicles-OCR, the first benchmark designed to evaluate multimodal AI models on ancient Chinese script recognition. The benchmark includes 2,800 expert-annotated images covering seven major calligraphic styles — from oracle bone script to cursive — and quantifies recognition difficulty across these historical forms.
The research team tested 28 leading multimodal large language models, and the results were grim for ancient scripts. On the cross-era character detection task, GPT-5 and Gemini 2.5 Pro scored near zero. The best-performing model managed only 16.5. Even when researchers removed the localization step by providing tight bounding boxes, top accuracy reached just 27.1%, with Gemini 3.1 Pro hitting only 14.0% on oracle bone script.
These numbers confirm that modern models rely heavily on clean, modern typography. Faced with unconstrained, noisy ancient physical media, their text segmentation mechanisms break down entirely. Font classification analysis further revealed that models often identify the texture of the medium — like tortoise shell or bronze patina — rather than true character strokes.
The study also uncovered a counterintuitive phenomenon: enabling reasoning modes actually degraded accuracy on ancient scripts. In side-by-side comparisons, nearly every model that supports "thinking mode" performed worse with it turned on. When the underlying visual perception is missing, chain-of-thought reasoning doesn't fix the gap — it amplifies hallucinations, producing confident but wrong answers.