Glossary
Document processing, OCR & IDP glossary
Plain-language definitions of the Intelligent Document Processing, OCR, and document-automation terms — and the specific document standards — WiseTREND works with every day.
In short
A plain-language glossary of the Intelligent Document Processing, OCR, and document-automation terms WiseTREND works with every day — including the AI concepts behind modern document capture and the specific document standards (CMS-1500, ACORD, MICR, MRZ) we automate for clients.
Core IDP & AI concepts
- Intelligent Document Processing (IDP)
- Intelligent Document Processing is the use of AI — OCR, machine learning, computer vision, and natural language processing — to automatically classify documents, extract structured data, validate it, and deliver it into business systems. IDP turns unstructured documents such as invoices, IDs, and claims into clean, usable data with little or no manual entry. WiseTREND IDP products
- Optical Character Recognition (OCR)
- OCR converts an image of text — a scan, photo, or PDF — into machine-readable characters. It is the recognition layer inside an IDP pipeline, but on its own it only produces text, not the validated, structured data a business system can use. OCR vs IDP
- Intelligent Character Recognition (ICR)
- ICR is OCR for handwriting. It uses machine learning to read hand-printed characters in form fields — common on ACORD forms, intake paperwork, checks, and delivery slips — where traditional OCR, tuned for machine print, struggles.
- Optical Mark Recognition (OMR)
- OMR detects marks rather than characters — checkboxes, filled bubbles, and rating scales on surveys, ballots, and assessment forms. It answers 'which option was selected?' instead of 'what text is written here?'. WiseSURVEY
- Document classification
- Document classification is the automatic identification of what each document is — invoice, remittance, ID, or contract — so the right extraction model runs. It lets an IDP pipeline sort mixed batches and split multi-document files without manual sorting.
- Data extraction
- Data extraction is the step that pulls specific fields, table rows, and line items out of a recognized document — vendor, totals, dates, diagnoses — and outputs them as structured key-value data rather than a wall of text.
- Key-value pair extraction
- Key-value pair extraction captures a label and its value together — for example 'Invoice #: INV-10427' or 'Total: $4,820.00' — so the meaning of each captured value is preserved when it lands in a database or form.
- Data validation
- Validation checks extracted data against business rules and reference data before it is trusted — confirming a vendor exists, that line items sum to the total, or that a date is well-formed. Low-confidence results route to a human; the rest flow through automatically.
- Straight-through processing (STP)
- Straight-through processing, also called touchless processing, means a document is captured, validated, and posted to the system of record with no human intervention. STP rate — the share of documents processed touchlessly — is the headline metric for an IDP deployment's ROI.
- Human-in-the-loop (HITL)
- Human-in-the-loop is the review step where a person verifies or corrects only the fields an IDP system flags as low-confidence. A well-tuned pipeline sends just the genuine exceptions to a reviewer and lets the high-confidence majority post automatically.
- Natural Language Processing (NLP)
- NLP is the branch of AI that interprets human language. In document processing it locates and structures information inside free-form text — contracts, loss runs, medical notes — where there is no fixed layout to rely on.
- Machine learning (ML)
- Machine learning is the technique of training models on examples rather than fixed rules. In IDP, ML lets a single model read the same document type across thousands of layouts — for example invoices from every vendor — without a template for each.
- Generative AI / LLMs in document processing
- Generative AI and large language models (LLMs) add flexible understanding and summarization to an IDP pipeline — reading intent in unstructured documents and answering questions about them. In production they are paired with validation and human review to control accuracy and avoid hallucinated values.
Document capture techniques
- Structured, semi-structured & unstructured documents
- Structured documents have a fixed layout (a tax form). Semi-structured documents carry the same fields in different layouts (invoices from many vendors). Unstructured documents hold meaning in prose (contracts, letters). Each calls for a different capture strategy — templates, machine learning, or NLP.
- Structured vs unstructured data
- Structured data is organized into defined fields and rows a system can query directly. Unstructured data — the text inside a scanned document, email, or image — has no such organization. IDP exists to turn the second into the first.
- Document skill / extraction model
- A document skill (or extraction model) is a pre-trained capability for a specific document type that ships ready to recognize and extract its fields — so a project starts from a working model instead of being built from scratch. WiseTREND's Wise* plugins are tuned document models in this sense. WiseTREND products
- MICR (Magnetic Ink Character Recognition)
- MICR is the special font line printed along the bottom of a check that encodes the routing number, account number, and check number. Reading the MICR line accurately is essential to automating check capture and remote deposit. WiseCHECK
- CAR / LAR (check amount recognition)
- CAR (Courtesy Amount Recognition) reads the numeric amount on a check; LAR (Legal Amount Recognition) reads the written-out amount. Cross-validating the two raises accuracy on handwritten checks. WiseCHECK
- MRZ (Machine-Readable Zone)
- The MRZ is the two or three lines of monospaced characters at the bottom of passports and many IDs, encoded to the ICAO standard. Parsing the MRZ and cross-checking it against the printed text is a core step in identity-document verification. WiseID
- 3-way match
- A 3-way match compares an invoice against its purchase order and the goods-receipt record before payment, confirming that what was ordered, received, and billed all agree. Automating it is central to accounts-payable automation. WiseINVOICE
Document types & standards
- Accounts payable (AP) automation
- AP automation is the use of IDP to capture invoices, match them to purchase orders, route exceptions, and post them to an ERP without manual keying — typically cutting per-invoice processing cost by 60–80%. Finance solutions
- ACORD forms
- ACORD forms are the standardized insurance documents used for submissions and certificates — including ACORD 25 (Certificate of Liability), 125, 126, 127, 130, and 140. Carriers vary the layouts, so reliable capture needs models maintained as forms change. WiseACORD
- CMS-1500 & CMS-1450 (UB-04)
- CMS-1500 is the standard professional healthcare claim form (physicians, suppliers); CMS-1450, also called UB-04, is the institutional claim form (hospitals, facilities). Both carry hundreds of coded fields that healthcare IDP must capture accurately. WiseEOB
- EOB, ERA & the 835
- An Explanation of Benefits (EOB) is the human-readable statement of how a claim was paid; an Electronic Remittance Advice (ERA) is its electronic equivalent, often exchanged as an 835 transaction. Converting paper and image EOBs into 835-compatible data is a key healthcare automation task. WiseEOB
- Bill of lading (BOL) & proof of delivery (POD)
- A bill of lading is the contract and receipt that travels with a freight shipment; a proof of delivery confirms it arrived, often with a signature and any exceptions noted. Both are high-volume documents in transportation automation. WiseTRANS
- KYC (Know Your Customer)
- KYC is the regulated process of verifying a customer's identity during onboarding. Document-based KYC uses IDP to read and validate IDs, passports, and proof-of-address documents as the data-capture step of the wider identity-verification stack. WiseID
ABBYY platform terms
- ABBYY FlexiCapture
- ABBYY FlexiCapture is ABBYY's deeply configurable data-capture platform, deployable on-premises or in the cloud. It is the foundation WiseTREND's Wise* plugins extend, and it suits complex documents and validation at scale. FlexiCapture vs Vantage
- ABBYY Vantage
- ABBYY Vantage is ABBYY's cloud-native, skills-based IDP platform, offering pre-trained document 'skills' and a low-code design experience for fast deployment. It complements FlexiCapture rather than replacing it. FlexiCapture vs Vantage
- ABBYY FineReader
- ABBYY FineReader is ABBYY's OCR and PDF product line — including FineReader PDF for desktops and FineReader Server for high-volume conversion — used to digitize, convert, and compare documents.
- ADRT (Adaptive Document Recognition Technology)
- ADRT is ABBYY's technology for understanding a document's logical structure — headings, tables, and multi-page relationships — so extraction adapts to layout variation instead of relying on a rigid template.
- Process intelligence (process & task mining)
- Process intelligence — including ABBYY Timeline — analyzes event logs and desktop activity to build a data-driven picture of how a business process actually runs, revealing where document automation will pay off most.
Have a document type we didn't list?
If your team wrestles with a document we didn't define here, tell us about it — we have almost certainly automated something like it.