Digitize Urdu books with AI OCR: scans to searchable, editable text

Urdu is one of the hardest scripts for machines to read — flowing Nastaliq, joined letters, and dots that change everything. Most OCR tools give up on it. RafayGen's Urdu OCR at rafaygen.space/ocr was built specifically to read Urdu, so a scanned book becomes text you can search, copy and edit.

What it does

  • Reads scanned Urdu pages, photos of text, and image-only PDFs.
  • Outputs searchable text you can copy, and a rebuilt PDF with a real text layer.
  • Handles whole books page by page, keeping the reading order.
  • Works for literature, notes, documents and old manuscripts.

Why it exists

It began as a project to digitise Urdu literature — the kind of poetry and prose that only exists on paper and risks being lost. Making Urdu machine-readable means it can be searched, quoted, translated and preserved. That capability is now open to everyone on RafayGen.

How to use it

Go to rafaygen.space/ocr, upload your scanned Urdu file, and let it process. You get back searchable text and a downloadable PDF. Free to try in the browser.

Try RafayGen free →