Documentation Index

Fetch the complete documentation index at: https://knowledgecenter.docuware.com/llms.txt

Use this file to discover all available pages before exploring further.

Introduction to Intelligent Document Processing

Prev Next

DocuWare Intelligent Document Processing (IDP) uses artificial intelligence to to process your documents. This includes splitting, classification and data extraction for indexing of documents.

The processing of documents and especially the indexing is the bases for all processes with DocuWare. Legal rules require perfectly indexed documents, document search and business processes are based on index data of documents and can be streamlined once the data is extraced. Index values like an invoice number, a vendor name, or a contract date are what turns a stored file into a searchable, actionable record. IDP takes over this work: it reads the document content, identifies the relevant information, and writes it into your DocuWare index fields.

The problem IDP solves

Most business documents are unstructured. They arrive as scanned PDFs, email attachments, or photographed pages, and each supplier formats the documents differently. The information you need is somewhere on the page, but not in a predictable data structure. Extracting useful data from these files by hand is slow, repetitive, and prone to mistakes. Traditional automation approaches try to solve this with fixed templates, but these break whenever a supplier changes their layout or a new document type appears.

IDP takes a different approach. Instead of relying on rigid rules, it uses artificial intelligence that learns from your actual documents. The AI recognizes document types, finds relevant fields, and handles variations in layout, language, and scan quality without requiring a separate template for each supplier or format. The result is a processing pipeline that adapts to your documents rather than forcing your documents into a fixed structure.

Information: Structured vs. unstructured

Structured data lives in clearly defined schema: a database, spreadsheet, or an XML. Each piece of information has a fixed place and a known format. Unstructured data is everything else. A scanned invoice is unstructured because the invoice number might sit in the top right corner on one supplier's document and in the middle of the page on another's. The same is true for contracts, delivery notes, and most business correspondence. The information is there, but its position varies from document to document.

How IDP processes a document

Every document that enters IDP goes through the similar sequence of steps. The process runs automatically in the background, and in most setups no user interaction is required at all. If your process demands a human check, for example during the early stages of a new IDP setup or for high-risk documents, you can add an optional verification step where a user reviews and corrects the extracted values before archiving.

Pre-processing, splitting, and classification

The journey begins with pre-processing: IDP prepares the incoming file by running optical character recognition (OCR), correcting image quality, and aligning pages. The goal is to produce clean, readable input regardless of how the document was captured.

Next comes splitting. In many real-world scenarios, documents do not arrive one at a time. A scanner operator might feed a stack of 15 invoices into a device and produce a single PDF. IDP detects where one document ends and the next begins, without barcodes or separator pages, and breaks the file into individual documents that can each be processed on their own.

Once the documents are separated, classification takes over. IDP looks at each document and determines what type it is: an invoice, a delivery note, a contract, a purchase order, or whatever document classes you have defined for your organization. This step is what allows IDP to extract content and archive each document.

Extraction and archiving

With the document type identified, IDP moves on to extract the actual content. Every document contains pieces of information that are relevant for your business process. On an invoice for example these are the invoice number, the date, the vendor name, and the total amount. On a contract, they might be the contract ID, the effective date, and the parties involved. In IDP terminology, each of these pieces of information is called a field. You decide which fields IDP should look for.

IDP locates these fields in the document, reads their values, and writes them into the corresponding DocuWare index fields. Once this step is complete, the document is searchable and ready for further processes.

Finally, the fully indexed document is stored in the target file cabinet. If a workflow is configured, say an invoice approval process, it starts automatically as soon as the document arrives. From the user's perspective, documents simply appear in the right place with the right metadata, as if someone had filed them by hand.

AI model types

IDP relies on three types of AI models, each responsible for one of the core processing steps described above. In DocuWare, these models are also referred to as agents.

  • Splitting models detect document boundaries inside a multi-page file. They work purely from content, recognizing when the text, layout, or structure shifts from one document to another. They do not require barcodes, blank pages, or fixed page counts. This makes them particularly useful in mailroom scenarios where operators scan mixed stacks of paper without pre-sorting.

  • Classification models assign each document to a document class that you define. These classes can be as broad or as specific as your organization needs. A simple setup might distinguish between "Invoice", “Delivery Note” and "Other." A more advanced setup might differentiate between "Domestic Invoice," "International Invoice," "Credit Note," and "Pro Forma Invoice." The model learns from examples, so as your document mix changes over time, the classification can evolve with it.

  • Extraction models do the detailed data extraction. You define which fields the model should extract: invoice number, date, line items, totals, or any other piece of information relevant to your process. The model finds these values even when layouts vary between suppliers or when scan quality is poor. Consider a finance team that processes invoices from 200 different suppliers. Instead of maintaining a template for each one, a single extraction model handles all of them.

How to get your models

You do not have to build every model from scratch. Depending on your scenario, you can choose from several paths and combine them within the same IDP setup.

Pre-built models

For common use cases, pre-built models are available out of the box. These cover scenarios like standard invoice extraction or basic document splitting and require no training or configuration. If your documents follow widely used formats, a pre-built model may be all you need to get started.

Models trained from DocuWare file cabinets

If you already have a large volume of documents archived in DocuWare, you can use them as training data. IDP models can be trained directly from DocuWare Configurations > DocuWare IDP. You select the file cabinets that contain your training documents, start the training, and the resulting models are tailored to the specific formats, layouts, and content of your actual files. Training may take up to 24 hours, but you do not need to wait. You can continue configuring your workflow in the meantime.

Custom models on the IDP Platform

For specialized or complex requirements, you can create fully custom models on the standalone IDP Platform. This platform supports any document type, whether or not the documents are archived in DocuWare, and is typically used with the help of your DocuWare partner or contact.

  1. Custom models can be built in two ways. The classic approach: you upload sample documents, mark the relevant fields, train the model, and review the results. This method takes more time and effort, but it produces very high accuracy.

  2. The Gen AI approach works differently and combines simple setup with high accuracy. Before annotating documents, you describe what the model should extract using natural-language instructions, for example "Extract the invoice number," "Return only the domain part of the email address," or "Find the delivery date in the header area." The model starts working immediately, with no training phase at all. This makes it a good fit for many simple use cases but also for proofs of concept or situations where you need results quickly. It is also a good fit If you want to train the model while you are working with the documents and validate them during storing.
    The Model can still be trained with training data, refined and tested to get the highest possible accuracy.

Where IDP fits into your document flow

IDP processes documents at the point where they enter DocuWare. The two most common entry points are email import and the DocuWare Desktop Apps.

  • Email import:
    For email import, you configure DocuWare to monitor a mailbox and import incoming messages automatically. When you add an IDP configuration to this setup, every attached PDF is classified and indexed before it reaches the file cabinet. A typical example is an accounts-payable mailbox that receives dozens of invoice PDFs per day. IDP classifies each attachment, extracts the key fields, and archives the document without manual intervention. Read more about configuring IDP for email import.

  • DocuWare Desktop Apps:
    For DocuWare Desktop Apps, documents added through the Scan or Import plug-ins pass through the same IDP pipeline. Paper invoices captured with DocuWare Scan and existing PDF files brought in through the Import plug-in are automatically split, classified, indexed, and archived. Read more about configuring IDP for DocuWare Desktop Apps.

Frequently asked questions

What accuracy can I expect?

Accuracy depends on several factors: the quality of your scans, the variety of document layouts, and whether you use a pre-built, prompt-based, or annotation-trained model. Pre-built models work well for standard documents. Custom models trained with annotated data typically reach the highest accuracy. In general, the more representative your training data is, the better the results. There is no universal number, but you can monitor confidence scores over time and adjust your setup accordingly.

Which document formats and languages are supported?

IDP processes PDF files, which is the most common format for scanned and emailed business documents. Documents in other formats such as TIFF or JPEG are converted during pre-processing. IDP supports multiple languages. The exact set of supported languages depends on the model and configuration. Contact your DocuWare partner for details on your specific scenario.

How long does it take to get a custom model into production?

With the GenAI approach, you can start extracting data within minutes. There is no training phase. For annotation-based models, the timeline depends on how many documents you annotate and how varied your document layouts are. Training itself can take up to 24 hours once started. In practice, getting from the first annotation to a production-ready model is typically a matter of days, not weeks.

Can I retrain a model when document layouts change?

Yes. If a supplier changes their invoice format or a new document type appears, you can update existing models with additional training data. You do not need to build a new model from scratch. For models trained from DocuWare file cabinets, you can start a new training run that includes the updated documents.

Supported versions: DocuWare Cloud