OCR Online — Image to Text

Extract editable text from a photo, screenshot or scanned page. Pick the language (English, Hindi, Tamil, Bengali and 20+ more), optionally clean up the image, and get the text with a confidence score. It runs the Tesseract OCR engine inside your browser, so the image is never uploaded.

Drop an image here, click to browse, or paste a screenshot with Ctrl+V

JPG, PNG, WebP, BMP, GIF. Printed text works best. For a PDF, convert pages with PDF to JPG first.

Image sent to OCR

Extracted text

How to extract text from an image

  1. Drop the image on the box, browse for it, or take a screenshot and press Ctrl+V (⌘+V on a Mac).
  2. Choose the language printed in the image. Add a second language for mixed documents, such as Hindi with English.
  3. Leave Greyscale + auto-contrast on. Turn on the other clean-up options only if the first result is poor (see the table below).
  4. Click Extract text. Check the confidence badge and tick Highlight doubtful words to see what to proofread.
  5. Edit the text if needed, then copy it or download it as a .txt file.

Tesseract OCR online, in your browser

This tool is built on Tesseract, the open-source OCR engine originally developed at HP and later maintained by Google, which is still the standard free OCR engine used by desktop apps, scanners and command-line workflows. Here it runs as Tesseract.js — the same C++ engine compiled to WebAssembly — with its LSTM neural-network recogniser and the official trained language files. You get the accuracy of tesseract image.png out -l eng --psm 3 without installing anything or sending the image to a cloud API.

The Layout menu maps to Tesseract's page segmentation modes (--psm): Automatic is PSM 3, Single block is PSM 6, Single column is PSM 4, Single line is PSM 7 and Sparse text is PSM 11. Keep spacing sets preserve_interword_spaces=1, which helps with tables and code.

OCR accuracy tips

Problem in the imageWhat to change
Text is small (a screenshot of a long page, a photo taken from far away)Turn on Upscale 2×. Tesseract reads best when capital letters are 25–40 px tall.
Grey text, faded print, uneven lighting on a phone photoKeep auto-contrast on; if shadows remain, add Black & white.
White text on a dark background (dark-mode apps, slides)Turn on Invert.
Scanned sideways or upside downUse Rotate image. Tesseract does not auto-rotate in this mode.
A single line such as a serial number, IFSC code or captcha-like labelSet Layout to Single line.
Scattered labels (a screenshot of an app, a signboard, a product box)Set Layout to Sparse text.
Text comes out jumbled in two columnsCrop to one column with the image cropper and run each separately.
Random letters, wrong alphabetThe language is wrong. Hindi text read as English, for example, gives nonsense.

Two things clean-up cannot fix: motion blur and JPEG damage. If the text looks smeared when you zoom in, retake the photo in good light, holding the phone parallel to the page.

Supported OCR languages

LanguageCodeScriptNotes
EnglishengLatinHighest accuracy; also handles most numbers and symbols.
Hindi, Marathi, Nepali, Sanskrithin, mar, nep, sanDevanagariPrinted text is good; matras on low-resolution images need Upscale.
BengalibenBengaliAlso covers Assamese letters in most documents.
Tamil, Telugu, Kannada, Malayalamtam, tel, kan, malDravidian scriptsUse 300 DPI scans or sharp photos; conjuncts are the usual errors.
Gujarati, Punjabi, Odiaguj, pan, oriGujarati, Gurmukhi, OdiaWorks on printed forms and newspapers.
Urdu, Arabicurd, araArabic (right to left)Printed Naskh is fine; Nastaliq calligraphy is harder.
Spanish, French, German, Italian, Portuguesespa, fra, deu, ita, porLatinChoose the right one to get accents such as é, ü, ñ.
RussianrusCyrillic—
Chinese (Simplified), Japanese, Koreanchi_sim, jpn, korCJKLarger language files; horizontal text works best.

Examples

Copying an error message from a screenshot

Press PrtSc or Win+Shift+S, then Ctrl+V on this page. Screenshots are sharp, so the default settings usually return 99%+ confidence. Tick Keep spacing for stack traces or code so indentation survives.

A photo of a printed Hindi notice with English words

Set Language to Hindi and Second language to English. Leave auto-contrast on. If the photo is from a distance, turn on Upscale 2×. Check highlighted words: numbers and English acronyms are where mixed-script errors usually appear.

A bill or receipt

Thermal receipts fade. Turn on Black & white, and use Layout Single column. Then use Keep spacing if you want item prices to stay lined up with their names.

What OCR can and cannot do

FAQ

Is this OCR tool free and private?

Yes. It is free with no sign-up or page limit, and the Tesseract engine runs inside your browser. Your image is processed on your device and is never uploaded to a server.

Is this the real Tesseract OCR?

Yes. It uses Tesseract.js, which is the open-source Tesseract engine (LSTM model, version 5) compiled to WebAssembly, with the official trained language data. Results match desktop Tesseract with the same language and page segmentation settings.

How accurate is the OCR?

On clean printed text at a readable size, expect 97 to 99 percent of characters to be correct. Accuracy drops with blur, low contrast, text smaller than about 20 pixels high, decorative fonts and handwriting. The confidence score and highlighted words show where to proofread.

Can it read Hindi and other Indian languages?

Yes. Choose Hindi, Marathi, Bengali, Tamil, Telugu, Gujarati, Kannada, Malayalam, Punjabi, Odia, Urdu, Nepali or Sanskrit. For documents that mix scripts, such as a Hindi form with English words, set the second language to English.

Can it read handwriting?

Only neat, separated block capitals work reasonably. Tesseract is trained on printed text, so cursive and joined handwriting usually give poor results.

How do I OCR a PDF?

Convert the PDF pages to images first with the PDF to JPG tool at 300 DPI, then open each image here. If the PDF was created digitally rather than scanned, you can usually just select and copy the text in a PDF reader instead.

Why does the first run take longer?

The browser downloads the OCR engine (about 3 MB) and the trained data for your language (roughly 2 to 15 MB depending on the language). Both are cached, so the next image starts almost immediately.