OCR and full-text search in documents

OCR and full-text search in documents

Text recognition on scanned and photographed documents, automatic indexing and search by content — find a document by what's written in it, not by its file name.

A scanned archive with no text recognition is a box of images: you know the document is there, but you can only find it if you remember where you put it. OCR fixes exactly this limitation. Every uploaded page is analyzed automatically, the text is extracted and indexed, and from that moment the document becomes searchable by any word that appears in it — a name, a number, an amount, an equipment name. Combined with metadata filters, the practical result is that questions like "where is the notice from spring 2022 for the Cluj work point?" get an answer in seconds, not an afternoon.

  • Automatic text recognition for scanned PDFs, images and photographed documents.
  • Support for Romanian, including diacritics — search finds the word even if it was typed without them.
  • Content indexing on upload, with no manual step and no waiting for anyone.
  • Combined search: words from the text plus metadata filters (type, date, department, issuer, category).
  • The result highlighted directly on the document's page, so you can see why it was found.
  • Automatic extraction of common fields (date, number, issuer, amount) to help fill in metadata.
  • On-demand reprocessing of older documents, once scan quality has improved.
  • Search limited to what you have the right to see — results respect access rights, they don't bypass them.

How text recognition works, in short

To a computer, a scan is an image: colored dots, without meaning. OCR — optical character recognition — scans the image, identifies the areas that contain writing, separates lines and characters and turns them into text. On top of this text, an index is built — a structure that allows fast searching without re-reading every document. The quality of the result depends on the source: a printed document, scanned straight, at 300 dpi, is recognized almost perfectly; one photographed at an angle, with shadows, or handwriting, is recognized only partially. That's why the platform always keeps the original image alongside the extracted text — the text serves the search, the image remains the document.

Content search plus metadata filters

Text-only search returns too many results, and metadata-only search assumes someone filled it in correctly. The combination of the two is what works in practice. You start from a word or phrase that definitely appears in the document, then narrow down with the filters you already have: the time range, the department, the records-schedule category, the issuer. Results display the text snippet where the match was found, so you can quickly tell the document you're looking for apart from similar ones. The same search can be saved and reused — useful for periodic checks, where the question repeats month after month.

Metadata filled in automatically from the recognized text

The most time-consuming part of digitizing an archive isn't scanning, it's filling in metadata. OCR cuts this work down significantly: common fields — the document's date, number, issuer, amount — can be extracted automatically from the recognized text. They're proposed, not imposed: the operator confirms or corrects them, and corrections improve the results that follow. For large batches of documents of the same type, extraction templates can be defined once and applied to the whole batch, turning a task of weeks into one of days.

What OCR can't do, and how to avoid surprises

It's only honest to state the limits too. Handwriting is recognized inconsistently, and manually filled-in forms generally remain a problem. Documents with stamps overlapping the text, low-resolution scans or folded pages produce recognition errors. In practice, this means a searched word may be missing from the index even though it's on the paper. The recommendation is simple: for critical categories, fill in a few key metadata fields by hand, so retrieval doesn't depend solely on the recognized text, and scan at a decent quality from the start — rescanning costs more than scanning well the first time.

Legal references

  • Law no. 135/2007 (republished) — archiving documents in electronic form: the document must remain intact and accessible throughout the retention period
  • Law no. 16/1996 — the National Archives Law: keeping records of and retrieving documents from the company's archival fund
  • Regulation (EU) 2016/679 (GDPR) — recognized text may contain personal data; access to search results follows the same rights as access to the document

Frequently asked questions

Which files can be processed with OCR?

Scanned PDFs, images (JPG, PNG, TIFF) and photographed documents. Files that already contain digital text — electronically generated PDFs, text-based documents — are indexed directly, without recognition, because their text already exists.

Does it work on Romanian-language documents, with diacritics?

Yes. Recognition is configured for Romanian, and search is tolerant of diacritics: you'll find the document even if you typed the word without them, or the other way around.

Is handwriting recognized?

Only partially and inconsistently. Printed text is recognized very well; handwriting and manually filled-in forms remain unreliable. For such documents, the recommendation is to fill in a few key metadata fields by hand, so retrieval doesn't depend on OCR.

How long does processing a document take?

Usually a few seconds per page, depending on resolution and page complexity. Processing happens in the background at upload, so it never blocks anyone: the document is available immediately, and content search becomes active as soon as indexing finishes.

Can I search an archive uploaded before OCR was turned on?

Yes. Existing documents can be sent for reprocessing, individually or in batches, and after indexing they become searchable by content just like new ones.

Does the search show me documents I don't have access to?

No. Results respect access rights: a document you're not allowed to see doesn't appear in the list and isn't revealed through the text snippet. Search isn't a loophole around permissions.

Related articles

Make your scanned archive searchable

We turn on text recognition for new documents and reprocess your existing archive, so every document can be found by what's written in it.

4b2b.net
Business Ecosystem
4conta.ro
Accounting
4invoices.net
Invoicing App
4expenses.net
Expense Management
4notify.net
Notifications
4hosting.net
Hosting
4database.net
Databases
4buildsite.net
Website Builder
4myapp.net
App Builder
4avatars.net
AI Avatars
4chaty.net
AI Chatbot
4webagency.net
Web Agency Software
4softedu.net
Education Websites
4softcrm.net
CRM Platform
4softerp.net
ERP System
4softhr.net
Human Resources
4mystaff.net
Staff Portal
4myprojects.net
Project Manager
4docs.net
Document Management
4mycontracts.net
Contracts
4appointments.net
Appointments
4marketingonline.net
Marketing
4insurance.net
Insurance
4property.net
Real Estate
4lawyers.net
Legal Software
4mygarage.net
Auto Service
4driving.net
Driving Schools
4fleet.net
Fleet Management
4myevents.net
Events
4therapy.net
Therapy
4clinics.net
Clinics
4dental.net
Dental Practices
4restaurants.net
Restaurants
4beautify.net
Beauty Salons
4gym.net
Fitness Gyms
4guards.net
Security Companies
4construct.net
Construction Companies
4marketplace.net
Marketplace
4shopy.net
Online Store
4pricing.net
Price Comparison
4salefood.net
Food Delivery
4rentify.net
Rentals
4transports.net
Transport
4agencytravel.net
Travel Agency
4hotel.net
Hotels & Guesthouses
4ong.net
NGO Management
OCR and full-text search in documents | 4docs