How does OCR work on scanned documents?
OCR turns a scanned page from a plain image into a document you can search for any word in.
Răspuns scurt
- OCR stands for optical character recognition: the text in an image becomes searchable text.
- A scan produces only an image; without OCR, its content can't be found through search.
- OCR runs automatically when the document is uploaded and adds a layer of text over the image.
- After processing you can search for anything that appears in the document: a name, a registration number, a date or an amount.
What the law says (in brief)
- Law 135/2007 requires that a document kept electronically remain legible and unaltered: OCR adds search capability, it doesn't replace the original document.
- Law 16/1996 places the responsibility for quick retrieval on the holder, so the inventory remains necessary even with OCR.
- Regulation (EU) 2016/679, Art. 5(1)(e): text extracted through OCR contains the same personal data and falls under the same retention periods.
- Scanned financial-accounting documents remain subject to the periods in Law 82/1991 and OMFP Order 2634/2015.
Practical examples
- You search for "acceptance report hall 2" and get four results out of 3,500 scanned pages, in under a second.
- A tax ID written in the body of an incoming letter becomes searchable, even though it doesn't appear anywhere in the file name.
- At 200 scanned pages a day, manual indexing would mean about two hours of typing; OCR does it at upload.
- A personnel file from 2014, scanned retroactively, becomes searchable by the employee's name the same day.
Common mistakes
- Scanning at 150 dpi or lower, which visibly reduces recognition accuracy; 300 dpi is the practical threshold.
- Pages are crooked or photographed with a phone, without automatic straightening.
- Searching without diacritics, even though the recognized text contains them.
- Assuming OCR is exactly 100% accurate: on handwritten documents the result stays partial.
- Giving up the inventory, on the grounds that "there is full-text search anyway".
How 4docs helps
- Automatic OCR on upload, including for bulk imports from an old archive.
- Full-text search inside document content, not just the file name.
- Combined filtering by year, document type, department and retention period.
- The matched snippet is highlighted directly in the page preview.