Digitize Urdu books with AI OCR: scans to searchable, editable text
Urdu is one of the hardest scripts for machines to read — flowing Nastaliq, joined letters, and dots that change everything. Most OCR tools give up on it. RafayGen's Urdu OCR at rafaygen.space/ocr was built specifically to read Urdu, so a scanned book becomes text you can search, copy and edit.
What it does
- Reads scanned Urdu pages, photos of text, and image-only PDFs.
- Outputs searchable text you can copy, and a rebuilt PDF with a real text layer.
- Handles whole books page by page, keeping the reading order.
- Works for literature, notes, documents and old manuscripts.
Why it exists
It began as a project to digitise Urdu literature — the kind of poetry and prose that only exists on paper and risks being lost. Making Urdu machine-readable means it can be searched, quoted, translated and preserved. That capability is now open to everyone on RafayGen.
How to use it
Go to rafaygen.space/ocr, upload your scanned Urdu file, and let it process. You get back searchable text and a downloadable PDF. Free to try in the browser.