MARTIANGet a free audit
/ Drew Tada

Unlocking Sanskrit: How I Built an OCR System That Beat Every Frontier Model

I built an open Sanskrit OCR system that beat every frontier model we tested, marking a new SOTA for transcribing handwritten Devanagari.

Editorial photography

My Indian friend shared a tweet about how badly the latest AI models were transcribing Sanskrit manuscripts. GPT-6 was getting only roughly 30% of the text right, while earlier models barely managed ten words. The documents came from the MIDF/eGangotri collection: digitized Sanskrit manuscripts, some accompanied by transcriptions prepared by scholars. The handwriting looked legible, and these were the best models available. Existing Sanskrit and Devanagari OCR systems struggled too. The published AnciDev benchmark reports character error rates of 30.06% for its best Tesseract model, 46.33% for Attention-LSTM, and 48.59% for CNN-RNN. Basically no existing model was usable for this corpus, even things built specifically for Sanskrit. My friend, who reads Devanagari, looked through the documents and told me he couldn't understand why the frontier models were so bad at reading handwritten Sanskrit. I decided I could do a better job, and I did.

Looking through the collection, I realized we already had the essential ingredient for training: manuscript images paired with scholarly transcriptions. The data needed substantial cleanup: text sometimes belonged to an adjacent page, two written lines were combined into one, and scholars’ annotations appeared alongside the manuscript’s actual wording. I used AI to build browser tools that let me inspect the images and transcriptions together, move misplaced text, and separate annotations from body text. The resulting corpus grew to 2,879 annotated pages and 31,215 lines, and I released the dataset on Hugging Face for others to download and use. Since its release on Hugging Face, the dataset has been downloaded over 1,100 times.

With usable data in hand, I benchmarked fifteen vision-language models (VLMs), which take an image and generate text from it. GPT and Claude performed badly on these manuscripts. Gemini was the clear exception, with roughly a 10% character error rate, meaning about one character in ten needed correction. This is decent performance but still not great. I fine-tuned a small Qwen VLM on our data to see whether a specialized model could beat it. The transcriptions improved, but it repeatedly got stuck generating the same characters or passages. Our best version reached 20.71% error, compared with Gemini’s 10.49% on the same nine pages. Increasing image resolution and adjusting generation settings helped, but I decided we needed a different architecture.

I switched to a traditional OCR pipeline that divides the task into smaller pieces. First, a fine-tuned Kraken BLLA model finds the written lines on a page, a process called segmentation. Kraken’s polygon tracing then outlines each line and crops it into an individual image. A fine-tuned PP-OCRv6 text recognizer reads those images, and the results are assembled into the page’s transcription. We trained the line-finding and text-reading models using our reviewed examples, testing different approaches and adding data from manuscripts where they struggled. This made failures much easier to fix: I could see whether the system missed a line, cut off part of the writing, or misread a character. As the line finder improved, it supplied suggestions I could quickly approve or correct, accelerating the creation of more training data.

I tested the system on pages kept out of training in two groups. On nine new pages from two manuscripts whose handwriting was entirely absent from the training data, the system reached an 8.88% character error rate. On 220 pages from a manuscript whose handwriting was present in the training data, it reached 3.55%, compared with Gemini’s 9.86% on those same pages: 64% fewer character errors. Across all 229 pages, the combined error rate was 3.71%. The published evaluation documents the two groups and their training history.

This is the most accurate Sanskrit OCR system I tested on this corpus. I’ve released the models, code, and evaluation results so others can run it locally, inspect its mistakes, and improve it. If you have questions or comments, reach out to me at drew (dot) tada (at) martian (dot) engineering, or on twitter @MartianEng

← Back to blog
A broad red desert canyon
Free automation audit

Still figuring out what AI means for your company?

Get a clear plan with a free automation audit.

Get a free audit ↗