Access and Feeds

ECM: Search Depends on Clean MetaData

By Dick Weisinger

Accurate search depends on clean metadata, and Google’s on-line book project is an example of how not to handle metadata.  After studying the project, researcher Geoffrey Nunberg from the UC Berkeley School of Information came to the conclusion that the metadata Google is using is extremely out of whack.

Nunberg gives one example of book publication dates that are stored in the database, something you’d think would not be that hard to capture when books are digitized. Google’s database lists more than 250,000 books with incorrect dates, many of them showing up as 1899, a date that was used as a placeholder when the actual date wasn’t known.

Google admits that there are problems.  Google engineering manager Jon Orwant said  “Our providers have millions of errors like these, and we do what we can to eliminate them.  We have made substantial improvements over the past year, but I’m sure we can all agree there’s a great deal more to do.”

It’s a hard problem.  Getting and entering correct data is difficult, time-consuming and expensive.  Without good metadata searches can be highly inaccurate, either overlooking items that should be considered or targeting items that should be overlooked.  One suggestion for correcting the data is to open it up, much the same way as wikipedia does, to the public and letting them correct errors as they  come across them.

Digg This
Reddit This
Stumble Now!
Buzz This
Vote on DZone
Share on Facebook
Bookmark this on Delicious
Kick It on DotNetKicks.com
Shout it
Share on LinkedIn
Bookmark this on Technorati
Post on Twitter
Google Buzz (aka. Google Reader)

Leave a Reply

Your email address will not be published. Required fields are marked *

*