OCR Online Guide: Turn Images into Editable Text
Optical character recognition, or OCR, detects letter shapes in an image and returns editable text. It is useful for screenshots, scanned documents, receipts, notes, and photographed pages.
Extract text from an image with the free OCR tool →
How OCR works
A recognition pipeline runs in stages. First it preprocesses the image: convert to greyscale, increase contrast, remove speckle noise, and straighten (deskew) tilted text. Then it segments the page into blocks, lines, words, and finally individual glyphs. Each glyph is classified against known letter shapes, and a language model nudges ambiguous results toward real words, so c1own becomes clown in prose. That last step is why isolated codes and numbers are harder than sentences.
Worked example
Feed OCR a screenshot of the line "Invoice #10582 — total $1,240.00 due 5 Oct" and a typical result is clean. Feed it a crumpled thermal receipt of the same text and you might get "lnvoice #1O582 - total $I,24O.OO due 5 0ct": the digit 1 read as letter l, the digit 0 read as letter O, and the reverse. The fix is better input, then a careful proofread of every number.
Getting better results
- Use a sharp, evenly lit image with high contrast between text and background.
- Scan printed pages at about 300 DPI; avoid going under 150 DPI.
- Crop away borders, shadows, logos, and unrelated graphics.
- Keep text horizontal and the camera parallel to the page to avoid perspective distortion.
- Set the correct recognition language, especially for accented or non-Latin scripts.
- Review names, numbers, punctuation, and currency symbols after recognition.
Where OCR fits
| Use case | Reliability |
|---|---|
| Screenshot of on-screen text | High — crisp, aligned pixels |
| Flatbed scan of a printed page | High at 300 DPI |
| Phone photo of a document | Good with even light and no skew |
| Receipts, forms, tables | Mixed — layout and thermal fading hurt |
| Handwriting | Low — expect heavy correction |
OCR is not perfect
Unusual fonts, handwriting, low resolution, dense tables, watermarks, and busy backgrounds all produce errors. Treat OCR output as a first draft and proofread it before using the text in a legal, financial, medical, or operational document.
Privacy
A browser-based OCR tool can process images locally, so nothing is uploaded. Check the specific tool's behaviour before running IDs, contracts, receipts, or private records through it.
Frequently asked questions
What resolution do I need for a good scan?
About 300 DPI for printed pages. Below roughly 150 DPI the strokes blur and error rates climb. For a phone photo, fill the frame and keep the camera parallel to the page.
Why does OCR confuse certain characters?
Similar shapes: 0 and O, 1 and l and I, 5 and S, rn and m. The language model helps in prose but not in codes or tables, so proofread those.
Can OCR read handwriting?
Print recognition is mature; handwriting is far less reliable and depends on neat, separated letters. Treat handwritten results as a rough transcription.
Is my image uploaded?
A browser-based tool can run recognition locally. Confirm the behaviour before processing sensitive documents.
Related NeatJSON tools: Image to ASCII, Image Compressor, PDF to JPG, Image Metadata Viewer.