Access and Feeds

Intelligent Document Capture

By Dick Weisinger

Document capture is the digitization and input of information into a Content or Knowledge repository.  Capture is an important part of any Document and Content Management System and I’ve written about it here in previous blogs.

The capture component of document and content management systems have been around for years, and when discussing benefits and usage of new systems from the end user perspective, it is something that has often been overlooked or taken for granted. 

During the 90’s capture didn’t really change much. But scanners gradually evolved to support faster scans, higher DPI resolutions, more colors and smaller footprints.  This along with improved OCR software backed by computers with faster processors and more memory has allowed capture to evolve in a way that makes it a much more interesting component of the total system solution.

The market is growing fast.  One estimate shows that capture software sales grew by 18 percednt between 2004 and 2005 to roughly $1.1 billion.

A recent article in KMWorld magazine by Robert Smallwood describes a new phase of Document Capture called Intelligent Document Capture (IDC).  [Note that IDC is also known as Intelligent Document Recognition or IDR, as described by Harvey Spencer in the same issue of KMWorld.]  Capture solutions often pair the scan process with a post-processing phase where OCR is used to recognize text fields or strings embedded within the image.

More advanced types of OCR typically are assisted by using pre-defined zone templates and document markings that help identify the location of field data on the images. Today’s OCR software is becoming more and more sophisticated in not having to rely on searching fixed page locations for identifying field data. 

With IDC it is now possible for images to identify field data from images having completely unstructured page layouts or to have semi-structured data, some being in fixed locations and others in differing locations that are not pre-defined. 

An interesting application of unstructured OCR is license plate recognition using images taken by cameras mounted at toll plazas and traffic intersections.  Within the area of information and document management, contracts, resumes and letters are examples of unstructured documents that are good candidates for IDC. 

Semi-structured document examples include insurance claim forms, invoices and application forms.  Quick processing of these kind of forms provides big returns in time and cost savings.  Proper handling of invoice processing, for example, can help maximize dollars.  Invoices paid off quickly often can result in bonuses or discounted pricing by the billing company; invoices paid at the end of the invoice period allow the company to hold and invest their money until the due date.  Prompt processing that includes IDC and automated analysis of the payment terms provides options for the company to choose the best payment strategy.

Major players in the IDC space include Kofax, Captiva and ReadSoft.  The Formtek | Orion RepositoryLink product is an integration of Formtek’s Document Management capabilities with SAP ERP.  Formtek’s SAP-certified solution at one customer’s site manages more than 9 million invoice documents.  On the capture side of the Formtek SAP solution, we have an integration with the Readsoft IDC Invoices capture, allowing invoice data to be intelligently interpretted and sucked into the structured SAP data store.

Being able to recognize in-coming newly captured data automatically allows for automatic integration into workflow and Business Process Management BPM systems.  Metadata and tags can also be autogenerated based from data interpretted from the scanned images.

Digg This
Reddit This
Stumble Now!
Buzz This
Vote on DZone
Share on Facebook
Bookmark this on Delicious
Kick It on DotNetKicks.com
Shout it
Share on LinkedIn
Bookmark this on Technorati
Post on Twitter
Google Buzz (aka. Google Reader)

Leave a Reply

Your email address will not be published. Required fields are marked *

*