Glossary

OCR

Extracting text from images and scans, including Indian scripts and handwriting.

Pre-processing, deskew, denoise, contrast, determines OCR accuracy more than the recognition engine does.

Modern vision-language models substantially outperform classical OCR on poor scans, handwriting and complex layouts, at higher cost per page.

Output structure matters as much as character accuracy. A perfectly transcribed financial table flattened into a paragraph has lost the information that made it useful, so layout preservation belongs in the requirements alongside the accuracy target.

Set the confidence threshold against the cost of a mistake. On a financial document, routing an uncertain field to a human is cheap; on a bulk archive it may not be worth it, and that is a business decision rather than a technical default.

Commonly misunderstood: Indian-script accuracy varies considerably by script and scan quality, and a vendor quoting one figure for 'Indian languages' has not measured properly.

Related terms, in context

The concepts you almost always meet alongside ocr.

Intelligent document processing
Reading structured meaning out of documents, invoices, claims, forms, with confidence scores per field.
Computer vision
Systems that interpret images and video, inspection, counting, safety monitoring, reading.

Where this shows up in our work

OCR is not an abstraction for us. It is a decision we make on live projects. It shows up most directly in ocr & handwriting recognition, document processing & idp, where getting it wrong has a cost someone can measure.

If you are evaluating a vendor on this, the useful question is not whether they can define the term. It is what they measure, what they would refuse to do, and what happens in their system when the assumption behind ocr stops holding.

Questions

What is OCR?

Extracting text from images and scans, including Indian scripts and handwriting.

What do people get wrong about ocr?

Indian-script accuracy varies considerably by script and scan quality, and a vendor quoting one figure for 'Indian languages' has not measured properly.

Does Orqent Labs build this?

Yes, OCR & Handwriting Recognition and Document Processing & IDP. We work across India, covering all 19,238 PIN codes remotely.

Building something that involves ocr?

We will tell you honestly whether it is the right approach for your problem.

Or email bd@dtrasglobal.com · call +91 74118 77878