Odia OCR

Odia OCR reads scanned PDFs and photos of printed Odia and turns them into editable Unicode Odia text. Pages with English headings, email addresses or scheme names come out in both scripts, common misreads are corrected automatically, and the result downloads as a Word file. Everything is processed in your browser.

Why a scanned Odia page cannot simply be copied

A scanned letter or a photographed notice is a picture of text, not text. Selecting it in a PDF viewer either selects nothing or copies something unusable. Many phone scanning apps add a hidden layer of English character recognition to every page, so copying an Odia line from such a scan produces strings like ~Wl~Q QIQScll~I instead of the words on the page.

Older documents have a second problem. Files typed in legacy fonts such as Akruti or Sreelipi store Odia letters as English keyboard codes, so their text copies out as Latin gibberish even when the page looks perfect on screen. OCR avoids both problems by reading the letters as they appear and writing them back out in standard Unicode, which works in Word, WhatsApp, email and any website without a special font.

How to convert a scanned Odia document to text

  1. 1
    Add the scan

    Select a scanned PDF or photos of the pages, or paste a screenshot. Up to 40 pages can be read in one batch.

  2. 2
    Straighten the pages

    Rotate any page that is sideways and remove blank ones. A page read upside down produces nothing useful.

  3. 3
    Choose the language setting

    Use Odia + English for letters that contain English lines or words. Use Odia only for pages with no English; it reads roughly twice as fast.

  4. 4
    Check and download

    Each page appears beside its text. Correct anything misread directly in the text box, then copy the text or download it as Word or plain text. Downloads include your corrections.

What it is used for

  • Government letters and office orders. Reusing the wording of a memo, reply or order without retyping it.
  • Recruitment notifications. Pulling dates, post names and conditions out of a scanned advertisement or corrigendum.
  • Application forms. Getting the questions of a printed Odia form into a document you can fill in.
  • Old Akruti and Sreelipi PDFs. Recovering Unicode text from files whose fonts no longer copy correctly.
  • Books, circulars and printed pages. Quoting a passage or making text searchable.
  • Screenshots. Text from an image shared on WhatsApp or a page that will not let you select it.

Corrections made automatically

OCR engines confuse a small number of Odia letter shapes in the same way every time, mostly where the ି sign curls over ଥ, ଧ and ଖ. Because the mistakes are predictable, they can be fixed after reading. The tool shows how many corrections it made on each page, and you can switch them off in the options.

  • ଥୁବା, କରାଯାଇଥୁଲା become ଥିବା, କରାଯାଇଥିଲା, where the ି on ଥ was read as ୁ.
  • ଅଧୃକାରୀ, ଅଧ୍ଵକ become ଅଧିକାରୀ, ଅଧିକ, the same shape on ଧ.
  • ଉଲ୍ଲେଖୁତ becomes ଉଲ୍ଲେଖିତ, the same shape on ଖ.
  • ସଂଖ୍ଯା, ଦ୍ବାରା become ସଂଖ୍ୟା, ଦ୍ୱାରା, with ଯ-ଫଳା and ବ-ଫଳା written in standard Unicode.
  • ବର୍ତମାନ, ସମ୍ପୂର୍ଣ become ବର୍ତ୍ତମାନ, ସମ୍ପୂର୍ଣ୍ଣ, restoring the second ତ or ଣ after ର୍.
  • ୩୦.୦୯. ୨ ୦ ୨ ୬ becomes ୩୦.୦୯.୨୦୨୬, with broken dates joined back together.

Getting a cleaner result

Recognition quality depends far more on the scan than on anything else. A flat, straight, evenly lit page at about 300 DPI reads well; a curled page photographed at an angle under a tube light does not. Black-and-white scans work as well as colour. Crop away the edges of the desk or the next page if they appear in a photo, since stray marks are sometimes read as characters.

Very small print, faded photocopies and pages that have been photocopied several times give noticeably more errors. When a page comes out badly, rescanning it is usually quicker than correcting it.

Check before you rely on it

OCR on clear printed Odia gets most words right, but not every word. Names, reference numbers, amounts and dates are exactly the details where a single wrong character matters, so compare them against the scan before you send or publish the text.

Handwriting, signatures and rubber stamps are not read; the tool skips them by default so they do not fill the text with junk characters. Tables come out as lines of text rather than columns, so a table usually needs rebuilding in Word after conversion.

Read on your device

Pages are processed inside your browser. The files you add are not uploaded to filecrafthub.com or to any other server, are not stored, and are not linked to you. The first time you use the tool, your browser downloads the reading engine once and keeps it, so later visits start faster. There is no account and nothing to install.

This matters for the documents Odia OCR is most often used on. Office correspondence, applications and certificates often carry names, addresses and reference numbers that should not pass through an unknown server.

Frequently asked questions

Is Odia OCR free?

Yes. There is no sign-up, no page charge and no watermark on the text.

Are my files uploaded anywhere?

No. Pages are read inside your browser. Only the reading engine is downloaded, once, and your browser keeps it for later visits.

What kind of Odia text do I get?

Standard Unicode Odia. It works in Word, Google Docs, WhatsApp, Facebook, email and websites without installing any font.

Will it work on old Akruti or Sreelipi PDFs?

Yes, as long as the PDF shows the Odia letters correctly on screen. The tool reads the letters as they look, so the old font encoding does not matter and the result comes out in Unicode.

Can it read a page with both Odia and English?

Yes. Choose Odia + English and the English headings, email addresses and scheme names come out in English instead of garbled characters.

Does it read handwriting?

No. It is built for printed text. Handwritten entries and signatures are skipped.

How many pages can I convert at once?

Up to 40 pages per batch. A page usually takes 5 to 15 seconds, depending on your device.

Why does it say this window blocks the reading engine?

Some in-app browsers, such as the ones inside WhatsApp or Facebook, block the background process the tool needs. Open the page in Chrome, Edge or Firefox instead.

Related tools

Working with a long scanned file? Organise PDF lets you pull out just the pages you need before reading them, Crop Image trims a phone photo down to the page, and Text Editor is handy for tidying the converted text.

Published by the FileCraftHub team. FileCraftHub builds free, privacy-first browser tools for everyday file and document tasks. OCR output should always be checked against the original document before use. Questions can be sent through our contact page.