Access and Feeds

Open Source OCR

By Dick Weisinger

Google has announced the availability of an Open Source OCR package.  It’s called Tesseract.  In geometry, tesseract is the four-dimensional analog of the three-dimensional cube.  There’s nothing ground-breaking technically with this project, but it represents implementations of algorithms that were considered best-in-class ten years ago, and it’s available now as Open Source.

Tesseract was developed by HP from 1985-1995, and at some point after that someone at HP must have decided that OCR was something outside of HP’s core competency, and no additional development work followed.  Since then it’s been rediscovered and been taken under the wings by Google as a new Open Source project.  Google has fixed some known bugs and then felt that the quality was good enough to make it available publicly.

Google’s explanation for interest in OCR is that having a firm handle on OCR can assist them in their company mission of organizing “the immense amount of information available on the web”.  GooglePrint, Google’s library scan project comes to mind as a potential driver behind their interest.

OCR technology is interesting to Enterprise Content Mangement (ECM) because many of the vendors in this space, like Formtek, have their roots in Document Imaging.  OCR has become a commodity item, a common give-away that comes with new scanners.  But at a higher end, companies like Kofax and ReadSoft have done much to perfect both OCR accuracy and also create simple-to-use interfaces for OCR.  Tesseract in its current state can’t compete with those solutions, but it’s interesting to think of what possibilities might come out of this Open Source project.

Digg This
Reddit This
Stumble Now!
Buzz This
Vote on DZone
Share on Facebook
Bookmark this on Delicious
Kick It on DotNetKicks.com
Shout it
Share on LinkedIn
Bookmark this on Technorati
Post on Twitter
Google Buzz (aka. Google Reader)

Leave a Reply

Your email address will not be published. Required fields are marked *

*