known·good

Notes / ocrjapanesevlmbenchmarktransformerscudapython

Vertical Japanese (tategaki) OCR benchmark: dots.ocr-1.5 vs LightOnOCR-3 vs chandra-ocr-2

· by shaun, written up with claude

On 33 vertical Japanese pages (21k chars, exact ground truth), dots.ocr-1.5 had 1.27% CER, LightOnOCR-3-4B 2.59%, chandra-ocr-2 10.96% (repetition loops; turns small kana っ into つ). LightOnOCR-3-1B ships mis-keyed weights and 0.8B ships use_cache=false.

Tested onArch Linux (kernel 7.2); NVIDIA RTX 5080 16 GB, driver 610.57; PyTorch 2.10.0+cu128 (2.7.0+cu128 for dots.ocr); transformers 5.19.0 (LightOnOCR-3), 5.3.0 (chandra-ocr 0.2.0), 4.56.1 (dots.ocr); LightOnOCR-3 snapshots of 2026-10-08; dots.ocr-1.5; chandra-ocr-2; Chromium 152 for rendering

Symptoms #

The question was which local OCR model to use for vertical Japanese (縦書き, tategaki): novels, newspapers, old books. Two things made it hard to answer:

  • Easy benchmarks are saturated. On clean English, tables and horizontal Japanese, every recent OCR VLM scores 0.0–0.6% character error rate (CER). A single vertical paragraph separated them by only 1–3 characters, which is noise.
  • The failures that matter are not visible in averages of short pages. On longer vertical pages some models loop (repeat a phrase until the token limit), transcribe bleed-through from the back of the page, or silently change the spelling system (っ → つ, 來 → 来). These only show up with enough text and enough page types.

So a dedicated vertical-Japanese test set was built, with exact ground truth, and five models were run on it.

Results (CER = edits ÷ ground-truth characters; lower is better):

Model CER Median page Pages > 10% CER Time, 33 pages
dots.ocr-1.5 (3B) 1.27% 0.96% 0 271 s
LightOnOCR-3-4B 2.59% 1.52% 0 222 s
LightOnOCR-3-0.8B 4.15% 3.48% 1 154 s
LightOnOCR-3-1B 9.00% 4.41% 7 194 s
chandra-ocr-2 10.96% (90.6% uncapped) 1.74% 3 629 s

Each page's edits are capped at its own length, so one runaway loop costs at most one page of error. Without the cap, chandra's single 18,500-character loop dominates everything (90.6%).

By page type (same metric):

Category Pages LO3-4B LO3-0.8B LO3-1B chandra dots
clean block, 6 fonts 6 1.51 3.02 8.63 1.51 0.88
paperback page 5 1.16 2.39 7.36 1.76 0.92
furigana (ruby) 4 3.40 3.78 30.46 9.24 1.29
old orthography 旧字旧仮名 4 3.68 6.27 6.75 0.70 1.71
numbers / Latin 3 2.85 0.57 33.67 3.71 0.14
2–3 tier newspaper 段組 4 2.99 4.72 2.93 22.08 1.49
brush / pen fonts 2 1.47 2.94 3.67 0.92 1.28
simulated scans 5 2.28 4.82 8.80 8.58 1.18

Six of the 33 test pages: paperback, furigana, old orthography, 3-tier newspaper, numbers and Latin, simulated phone photo

Root cause #

The differences come from model behaviour, not from reading order.

Reading order is fine for all of them. The score was also computed order-free (a multiset comparison of characters). The gap between ordered and order-free error is only 0.25–1.2 points for every model. All five read right-to-left columns and multi-tier newspaper pages in the right sequence.

chandra-ocr-2 normalizes modern text to pre-war spelling. It enlarges small kana in modern text: っ → つ, ゃ → や, ょ → よ (123 times across the set), and sometimes writes old kanji forms (気 → 氣). Example from a paperback page:

textground truth: 糊代りに使った飯粒の / 貼らさってあった / な筈がなかった
chandra:      糊代りに使つた飯粒の / 貼らさつてあつた / な筈がなかつた

The same bias makes it the best model on genuinely old-orthography pages (0.70%). It also produced three blowups: a repetition loop on a 2-tier page (…だっただっただった… for 18,546 characters, 233 s), mirrored bleed-through transcribed as text on a thin-paper scan, and furigana emitted inline as <sup>reading</sup>, sometimes attached to the wrong kanji.

LightOnOCR-3 does the opposite: it modernizes old orthography. On 旧字旧仮名 pages the 4B writes 來 → 来, 樂 → 楽, 雜 → 雑, あつた → あった, and even rewrites classical grammar (かへらば → かへれば). It also sometimes reformats prose into Markdown (a recipe's 、-separated ingredient list became - bullets, plus an invented 「作り方」 heading). The 0.8B drops dakuten (ござ → ごさ) and leaks simplified Chinese (済 → 济). The 1B is not usable for vertical text, even after the checkpoint repair described below.

dots.ocr-1.5 has the least spelling bias. Its errors are isolated misreadings (檸檬 → 檜櫟 in one font, ゐ → る on old text). It had no page above 2.7% CER.

Furigana: if transcribed readings are forgiven, the ruby category is dots 0.91%, LightOnOCR-3-4B 1.87%, chandra 2.54%. dots and the 4B mostly drop the ruby, as wanted; the 4B occasionally writes it inline as 扇(あお).

Fix #

Use dots.ocr-1.5 for vertical Japanese. It was the most accurate in 6 of 8 categories and never failed badly. LightOnOCR-3-4B is the next choice: zero blowups and a bit faster, but expect modernized spelling on pre-war text. It also needs about 9 GB of VRAM in bf16, versus about 6 GB for dots. Avoid chandra-ocr-2 for tategaki, or at least guard it with a max-length check for loops.

If you run LightOnOCR-3 with Hugging Face transformers, three problems in the 2026-10-08 release need workarounds.

1. LightOnOCR-3-1B outputs garbage (pk pk pk pk …, 是很是很是很…) and the load log shows a LOAD REPORT with every language_model.model.layers.* key UNEXPECTED and every model.language_model.layers.* key MISSING. The checkpoint stores the language model weights under the wrong prefix, so they are dropped and the model runs with random weights. Older transformers versions behave the same. Rename the keys once:

pythonfrom safetensors.torch import load_file, save_file
sd = load_file("LightOnOCR-3-1B/model.safetensors")
sd = {("model.language_model." + k[len("language_model.model."):]
       if k.startswith("language_model.model.") else k): v for k, v in sd.items()}
save_file(sd, "LightOnOCR-3-1B-fixed/model.safetensors", metadata={"format": "pt"})
# copy the other files (config, tokenizer, processor) next to it; 310 keys are renamed

Load it with LightOnOcrForConditionalGeneration / LightOnOcrProcessor. Its config says model_type: mistral3, so AutoModelForImageTextToText picks the wrong class.

2. LightOnOCR-3-0.8B is about 10× slower than the 4B. Its config.json and generation_config.json ship with "use_cache": false, so every generated token recomputes the whole sequence. One page went from 208 s to 14.7 s with identical output:

pythonmodel.generate(**inputs, max_new_tokens=6144, do_sample=False, use_cache=True)

3. The 4B and 0.8B (Qwen3.5 architecture) fall back to slow PyTorch kernels:

text`causal_conv1d_fn` is falling back to its reference PyTorch implementation because `causal_conv1d` is not installed.
`chunk_gated_delta_rule` is falling back to its reference PyTorch implementation because `flash-linear-attention` is not installed.

Install both. flash-linear-attention is pure Triton. For causal-conv1d, use the prebuilt wheel matching your torch, so that it doesn't try to compile against a different system CUDA. Both worked on an RTX 5080 (sm_120):

bashpip install flash-linear-attention
pip install "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl"

For plain transcription, call LightOnOCR-3 with the image only (empty prompt) and enable_thinking=False in apply_chat_template (4B/0.8B). Resize the longest edge to 2048 px.

How it was found #

Test set (33 pages, 20,937 ground-truth characters). Passages come from public-domain Aozora Bunko texts (Akutagawa, Dazai, Sōseki, Kenji, Edogawa Ranpo, Higuchi Ichiyō, Mori Ōgai, Okamoto Kidō, Terada Torahiko and others). Three pages were written by hand for numbers, acronyms and URLs. Pages were rendered writing-mode: vertical-rl by headless Chromium in 11 fonts: Noto Serif/Sans CJK, IPA Mincho, and Google Fonts' Shippori Mincho, Zen Old Mincho, BIZ UD Mincho and Gothic, Kaisei Tokumin, Zen Maru Gothic, Klee One and Yuji Syuku.

  • For each page the passage is binary-searched to the longest prefix that fits the page box with no overflow. Overflow is measured in the browser with --dump-dom and a script that compares scrollWidth/scrollHeight. The cut is then backed off to a clause boundary, so the ground truth is exactly the visible text.
  • Furigana come from Aozora's own ruby markup (|漢字《かんじ》). Paragraphs with gaiji (※[#…]), layout annotations or /\ repeat marks are skipped.
  • Multi-tier pages use CSS multi-column inside vertical-rl, which stacks the columns top-to-bottom like a newspaper.
  • Scans are rendered, then degraded with OpenCV: skew + noise + JPEG, 150 dpi binarized, phone-photo perspective with uneven light and a gutter shadow, mirrored bleed-through, and a worn photocopy.

Scoring. Model markup (HTML, Markdown, LaTeX, grounding tokens) is stripped. Then NFKC (which folds vertical presentation forms such as ︒ ﹁ ﹂ and full-width digits), dash variants are unified, and all whitespace is removed. Variant kanji (旧/舊) and historical kana (ゐ/ゑ) are not folded: writing 旧字 as 新字 counts as an error. The metrics are strict CER, CER without punctuation, order-free error, and a ruby-tolerant CER that forgives readings written in <rt>, <sup>, brackets or inline. Every model ran in plain full-page OCR mode with greedy decoding.

Caveats. These are rendered pages, not physical scans; the scan category only simulates degradation. Aozora texts are public and probably in every model's training data. The test set is 33 pages, so differences under about 0.5 points within one category are not meaningful.

References #

  • dots.ocr-1.5: https://huggingface.co/rednote-hilab/dots.ocr-1.5
  • LightOnOCR-3: https://huggingface.co/lightonai/LightOnOCR-3-4B (also -1B, -0.8B)
  • chandra-ocr-2: https://huggingface.co/datalab-to/chandra-ocr-2
  • Aozora Bunko: https://www.aozora.gr.jp/
  • causal-conv1d prebuilt wheels: https://github.com/Dao-AILab/causal-conv1d/releases