Access and Feeds

Technology: Pushing the Envelope on Database Sizes

By Dick Weisinger

It wasn’t long ago that talk of terabyte databases seemed unfathomable. It hasn’t taken long for terabyte-size data to become commonplace. So it’s not too much of a surprise that now companies are boasting of their petabyte capabilities — that’s a 1000 terabytes or 1,000,000,000,000,00 bytes.

In May Yahoo claimed the honor of running the first petabyte-plus database in production. “We’ve built it to scale to tens of petabytes”, said Waqar Hasan, VP of data at Yahoo. The aratchitecture is designed to be able to scale to tens of petabytes easily. Much of the data collected is usage information about how people interact and navigate through Yahoo web pages. Yahoo’s pages are stickier than those of Google or Microsoft, with people often staying within Yahoo’s page set two to three times longer than that of competitor sites.

What’s Yahoo’s secret in being able to process so many bytes of information? They say it’s proprietary technology based on clustering their servers together.

But maybe Yahoo wasn’t the first. In January, Google discussed how they are processing 20 petabytes of data daily. “Google currently processes over 20 petabytes of data per day through an average of 100,000 MapReduce jobs spread across its massive computing clusters. The average MapReduce job ran across approximately 400 machines in September 2007, crunching approximately 11,000 machine years in a single month.”

Digg This
Reddit This
Stumble Now!
Buzz This
Vote on DZone
Share on Facebook
Bookmark this on Delicious
Kick It on DotNetKicks.com
Shout it
Share on LinkedIn
Bookmark this on Technorati
Post on Twitter
Google Buzz (aka. Google Reader)

Leave a Reply

Your email address will not be published. Required fields are marked *

*