Wednesday, September 9, 2026
venfeedSubscribe

A 3B model hit 97.6% on a new benchmark for digitising historical books

FineBooks tested 14 open models on 2,165 pages of 18th and 19th century printing. The best result did not come from the largest model.

Venfeed Editor2 min read
ShareXBlueskyLinkedInHNRedditEmail

A new benchmark called FineBooks has tested 14 open models on optical character recognition across 2,165 pages of 18th and 19th century books, with a three-billion-parameter model reaching 97.6 percent accuracy.

The headline finding is not the accuracy. It is that a 3B model produced it, on a task where the reflex assumption is that a larger model does better.

Why historical printing is hard

Modern OCR is effectively solved for clean contemporary documents. Historical printing is a different problem, and the reasons are specific.

Type from the period is irregular — hand-set, with worn and broken sorts, uneven inking and pages that show through from the reverse. Orthography is unstandardised, so spellings vary within a single book. The long s persists in 18th century printing and is easily confused with f. Ligatures are common. Marginalia, catchwords, ornaments and running heads interrupt the text block. And the artefact itself has aged: foxing, water damage, tight bindings that curve the text into the gutter.

A model doing well here is doing something more like reading than pattern matching, using context to resolve characters that are genuinely ambiguous in the image.

Why a small model wins

The result is less surprising on reflection than at first glance. OCR is a bounded task with a well-defined output, and the capability it needs — visual recognition plus enough language modelling to disambiguate — does not obviously scale with the general reasoning that additional parameters buy.

What matters far more is training data appropriate to the domain, and a smaller model trained on relevant material will beat a larger one that has seen little historical printing.

That has a practical consequence for anyone running a digitisation programme. Archives and libraries have millions of pages and small budgets; a 3B model runs on hardware an institution already owns, at a cost per page that makes mass digitisation feasible, without sending collections to a commercial API.

It is the same argument arriving from several directions this fortnight. OpenBMB released MiniCPM5-2B under Apache 2.0 for on-device use. Anker shipped a home hub doing camera analysis on a 26-TOPS chip. The frontier is not always the right tool.

The benchmark is the contribution

Fourteen open models evaluated on a common corpus of 2,165 pages is the useful artefact here, more than any single score.

Historical OCR has been assessed institution by institution, on private corpora, with results that could not be compared. A shared benchmark lets a library choose a model on evidence rather than on a vendor's claim, and lets model developers see whether they have improved.

It also sits awkwardly beside the week's other book-related story. 404 Media documented a warehouse where books are cut apart and scanned for AI training. The same digitisation capability serves an archive preserving a collection and an operation destroying books to feed a model, and the technical work is identical.

The benchmark's authors have not published the licence status of the 2,165 pages, though 18th and 19th century printing is comfortably out of copyright — which is precisely why it is available to benchmark on.

Venfeed Editor
Editor in chief

Runs the newsroom. Rename this profile in the studio to your own byline.

The Feed · weekdays, 6:30am ET

Every weekday, the AI stories that moved money or shipped code.

No cross-posting, unsubscribe anytime. See all newsletters