---
title: Vertical Japanese (tategaki) OCR benchmark: dots.ocr-1.5 vs LightOnOCR-3 vs chandra-ocr-2
date: 2026-10-09
summary: On 33 vertical Japanese pages (21k chars, exact ground truth), dots.ocr-1.5 had 1.27% CER, LightOnOCR-3-4B 2.59%, chandra-ocr-2 10.96% (repetition loops; turns small kana っ into つ). LightOnOCR-3-1B ships mis-keyed weights and 0.8B ships use_cache=false.
tags: ocr, japanese, vlm, benchmark, transformers, cuda, python
environment: Arch Linux (kernel 7.2); NVIDIA RTX 5080 16 GB, driver 610.57; PyTorch 2.10.0+cu128 (2.7.0+cu128 for dots.ocr); transformers 5.19.0 (LightOnOCR-3), 5.3.0 (chandra-ocr 0.2.0), 4.56.1 (dots.ocr); LightOnOCR-3 snapshots of 2026-10-08; dots.ocr-1.5; chandra-ocr-2; Chromium 152 for rendering
author: shaun, written up with claude
status: published
---
## Symptoms
The question was which local OCR model to use for **vertical Japanese** (縦書き, tategaki):
novels, newspapers, old books. Two things made it hard to answer:
- **Easy benchmarks are saturated.** On clean English, tables and horizontal Japanese,
every recent OCR VLM scores 0.0–0.6% character error rate (CER). A single vertical
paragraph separated them by only 1–3 characters, which is noise.
- **The failures that matter are not visible in averages of short pages.** On longer
vertical pages some models loop (repeat a phrase until the token limit), transcribe
bleed-through from the back of the page, or silently change the spelling system
(っ → つ, 來 → 来). These only show up with enough text and enough page types.
So a dedicated vertical-Japanese test set was built, with exact ground truth, and five
models were run on it.
Results (CER = edits ÷ ground-truth characters; lower is better):
| Model | CER | Median page | Pages > 10% CER | Time, 33 pages |
|---|---|---|---|---|
| **dots.ocr-1.5** (3B) | **1.27%** | 0.96% | 0 | 271 s |
| LightOnOCR-3-4B | 2.59% | 1.52% | 0 | 222 s |
| LightOnOCR-3-0.8B | 4.15% | 3.48% | 1 | 154 s |
| LightOnOCR-3-1B | 9.00% | 4.41% | 7 | 194 s |
| chandra-ocr-2 | 10.96% (90.6% uncapped) | 1.74% | 3 | 629 s |
Each page's edits are capped at its own length, so one runaway loop costs at most one
page of error. Without the cap, chandra's single 18,500-character loop dominates
everything (90.6%).
By page type (same metric):
| Category | Pages | LO3-4B | LO3-0.8B | LO3-1B | chandra | dots |
|---|---|---|---|---|---|---|
| clean block, 6 fonts | 6 | 1.51 | 3.02 | 8.63 | 1.51 | **0.88** |
| paperback page | 5 | 1.16 | 2.39 | 7.36 | 1.76 | **0.92** |
| furigana (ruby) | 4 | 3.40 | 3.78 | 30.46 | 9.24 | **1.29** |
| old orthography 旧字旧仮名 | 4 | 3.68 | 6.27 | 6.75 | **0.70** | 1.71 |
| numbers / Latin | 3 | 2.85 | 0.57 | 33.67 | 3.71 | **0.14** |
| 2–3 tier newspaper 段組 | 4 | 2.99 | 4.72 | 2.93 | 22.08 | **1.49** |
| brush / pen fonts | 2 | 1.47 | 2.94 | 3.67 | **0.92** | 1.28 |
| simulated scans | 5 | 2.28 | 4.82 | 8.80 | 8.58 | **1.18** |

## Root cause
The differences come from model behaviour, not from reading order.
**Reading order is fine for all of them.** The score was also computed order-free (a
multiset comparison of characters). The gap between ordered and order-free error is only
0.25–1.2 points for every model. All five read right-to-left columns and multi-tier
newspaper pages in the right sequence.
**chandra-ocr-2 normalizes modern text to pre-war spelling.** It enlarges small kana
in modern text: っ → つ, ゃ → や, ょ → よ (123 times across the set), and sometimes
writes old kanji forms (気 → 氣). Example from a paperback page:
```text
ground truth: 糊代りに使った飯粒の / 貼らさってあった / な筈がなかった
chandra: 糊代りに使つた飯粒の / 貼らさつてあつた / な筈がなかつた
```
The same bias makes it the best model on genuinely old-orthography pages (0.70%).
It also produced three blowups: a repetition loop on a 2-tier page
(`…だっただっただった…` for 18,546 characters, 233 s), mirrored bleed-through
transcribed as text on a thin-paper scan, and furigana emitted inline as
`reading`, sometimes attached to the wrong kanji.
**LightOnOCR-3 does the opposite: it modernizes old orthography.** On 旧字旧仮名
pages the 4B writes 來 → 来, 樂 → 楽, 雜 → 雑, あつた → あった, and even rewrites
classical grammar (かへらば → かへれば). It also sometimes reformats prose into
Markdown (a recipe's 、-separated ingredient list became `-` bullets, plus an invented
「作り方」 heading). The 0.8B drops dakuten (ござ → ごさ) and leaks simplified Chinese
(済 → 济). The 1B is not usable for vertical text, even after the checkpoint repair
described below.
**dots.ocr-1.5 has the least spelling bias.** Its errors are isolated
misreadings (檸檬 → 檜櫟 in one font, ゐ → る on old text). It had no page above 2.7% CER.
**Furigana:** if transcribed readings are forgiven, the ruby category is dots 0.91%,
LightOnOCR-3-4B 1.87%, chandra 2.54%. dots and the 4B mostly drop the ruby, as wanted;
the 4B occasionally writes it inline as 扇(あお).
## Fix
**Use dots.ocr-1.5 for vertical Japanese.** It was the most accurate in 6 of 8
categories and never failed badly. LightOnOCR-3-4B is the next choice: zero blowups and
a bit faster, but expect modernized spelling on pre-war text. It also needs about 9 GB
of VRAM in bf16, versus about 6 GB for dots. Avoid chandra-ocr-2 for tategaki, or at
least guard it with a max-length check for loops.
If you run **LightOnOCR-3** with Hugging Face transformers, three problems in the
2026-10-08 release need workarounds.
**1. LightOnOCR-3-1B outputs garbage** (`pk pk pk pk …`, `是很是很是很…`) and the load
log shows a `LOAD REPORT` with every `language_model.model.layers.*` key `UNEXPECTED` and
every `model.language_model.layers.*` key `MISSING`. The checkpoint stores the language
model weights under the wrong prefix, so they are dropped and the model runs with
random weights. Older transformers versions behave the same. Rename the keys once:
```python
from safetensors.torch import load_file, save_file
sd = load_file("LightOnOCR-3-1B/model.safetensors")
sd = {("model.language_model." + k[len("language_model.model."):]
if k.startswith("language_model.model.") else k): v for k, v in sd.items()}
save_file(sd, "LightOnOCR-3-1B-fixed/model.safetensors", metadata={"format": "pt"})
# copy the other files (config, tokenizer, processor) next to it; 310 keys are renamed
```
Load it with `LightOnOcrForConditionalGeneration` / `LightOnOcrProcessor`. Its
config says `model_type: mistral3`, so `AutoModelForImageTextToText` picks the wrong class.
**2. LightOnOCR-3-0.8B is about 10× slower than the 4B.** Its `config.json` and
`generation_config.json` ship with `"use_cache": false`, so every generated token
recomputes the whole sequence. One page went from 208 s to 14.7 s with identical output:
```python
model.generate(**inputs, max_new_tokens=6144, do_sample=False, use_cache=True)
```
**3. The 4B and 0.8B (Qwen3.5 architecture) fall back to slow PyTorch kernels:**
```text
`causal_conv1d_fn` is falling back to its reference PyTorch implementation because `causal_conv1d` is not installed.
`chunk_gated_delta_rule` is falling back to its reference PyTorch implementation because `flash-linear-attention` is not installed.
```
Install both. `flash-linear-attention` is pure Triton. For `causal-conv1d`, use the
prebuilt wheel matching your torch, so that it doesn't try to compile against a different
system CUDA. Both worked on an RTX 5080 (sm_120):
```bash
pip install flash-linear-attention
pip install "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl"
```
For plain transcription, call LightOnOCR-3 with the image only (empty prompt) and
`enable_thinking=False` in `apply_chat_template` (4B/0.8B). Resize the longest edge
to 2048 px.
## How it was found
**Test set (33 pages, 20,937 ground-truth characters).** Passages come from
public-domain Aozora Bunko texts (Akutagawa, Dazai, Sōseki, Kenji, Edogawa Ranpo,
Higuchi Ichiyō, Mori Ōgai, Okamoto Kidō, Terada Torahiko and others). Three pages were
written by hand for numbers, acronyms and URLs. Pages were rendered `writing-mode:
vertical-rl` by headless Chromium in 11 fonts: Noto Serif/Sans CJK, IPA Mincho, and
Google Fonts' Shippori Mincho, Zen Old Mincho, BIZ UD Mincho and Gothic, Kaisei Tokumin, Zen Maru Gothic,
Klee One and Yuji Syuku.
- For each page the passage is **binary-searched to the longest prefix that fits the
page box with no overflow**. Overflow is measured in the browser with `--dump-dom` and
a script that compares `scrollWidth`/`scrollHeight`. The cut is then backed off to a
clause boundary, so the ground truth is exactly the visible text.
- Furigana come from Aozora's own ruby markup (`|漢字《かんじ》`). Paragraphs with gaiji
(`※[#…]`), layout annotations or /\ repeat marks are skipped.
- Multi-tier pages use CSS multi-column inside `vertical-rl`, which stacks the columns
top-to-bottom like a newspaper.
- Scans are rendered, then degraded with OpenCV: skew + noise + JPEG, 150 dpi binarized,
phone-photo perspective with uneven light and a gutter shadow, mirrored bleed-through,
and a worn photocopy.
**Scoring.** Model markup (HTML, Markdown, LaTeX, grounding tokens) is stripped. Then
NFKC (which folds vertical presentation forms such as ︒ ﹁ ﹂ and full-width digits),
dash variants are unified, and all whitespace is removed. Variant kanji (旧/舊) and
historical kana (ゐ/ゑ) are **not** folded: writing 旧字 as 新字 counts as an error. The
metrics are strict CER, CER without punctuation, order-free error, and a ruby-tolerant
CER that forgives readings written in `