FineBooks project benchmarks OCR for better LLM training data
TL;DR. The FineBooks project evaluated 14 open-source OCR models to improve historical text digitization for language model training. - Hugging Face and EleutherAI collaborated to test OCR performance on over 2,000 historical book pages. - The top model achieved 97.6% character accuracy, sufficient for AI training but not scholarly transcription. - The research aims to reprocess millions of public-domain pages to enhance open AI training datasets.
- The FineBooks project, a collaboration between Hugging Face and EleutherAI, benchmarked 14 open-source OCR models on over 2,000 historical book pages to evaluate how well they can convert scanned texts into clean training data for AI language models.
- Smaller models frequently outperformed larger ones, with the top-performing model achieving over 97 percent character accuracy at a cost of less than two U.S. dollars per thousand pages.
- While the researchers consider the output quality sufficient for AI training purposes, they note that the models remain too error-prone for use in scholarly or scientific applications.
- The project targets reprocessing millions of public-domain pages with improved OCR to enhance open AI training datasets.