Access and Feeds

Big Data — Extreme Information Processing

By Dick Weisinger

We hear a lot about Big Data nowadays.  But just what is Big Data?

 
You’ll find that there is no lack of definitions for Big Data is, but then, there is no single standard and widely accepted definition.
 
Wikipedia defines Big data as “datasets that grow so large that they become awkward to work with using on-hand database management tools.”
 
When does data become ‘Big’ and when is the size of a data set just ‘Normal’ or ‘Small’?  Are data sets on the order of terabytes in size considered to be ‘Big’?  Or does a data set need to be on the order of petabytes of data before it is considered to be ‘Big’?  Just what is the threshold that a data set’s size must be before it can be called ‘Big’?  And if there such a threshold, does the threshold for ‘Big’ need to be a sliding one that gets continually bumped upwards as storage costs drop and as people and organizations continue to collect ever more data.
 
Because there is no mutually agreed on definition of Big Data, there are no right answers to these questions.  As of now — late 2011, the term Big Data has typically been applied to data sets that exceed 100 Terabytes in size.
 
The ‘Big’ in Big Data is actually a misnomer because it distorts how the technology is portrayed by focusing on just a single characteristic of the technology – the size, and there is really much more to Big Data than just collecting data sets that are huge in size.  Size doesn’t really adequately capture the whole picture of what people are trying to describe when they use the term Big Data.  An organization, for example, that passively collects hundreds of terabytes of data but which doesn’t attempt in any way to analyze or derive any insight from the data that they’ve collected isn’t practicing Big Data.
 
Research firms Gartner Inc. and Forrester Research both agree that the word ‘Big’ really doesn’t adequately describe what people mean by Big Data; ‘Big’ describes only a single parameter and is much more linear in description of something that is really very multi-dimensional.  Big Data is about processing huge data sets, but it’s also about performing complex analysis and visualization that can lead to insights and finding answers to difficult problems.
 
Brian Hopkins, Forrester Research analyst, says that the emphasis should not just be on the size or volume of data, which the descriptor ‘Big’ emphasizes, but should include other factors that better define what people are trying to describe when they use the term Big Data.  Both Forrester and IBM have begun to use additional adjectives to describe the characteristics of Big Data.  Forrester, for example, says that there are four V’s when it comes to Big Data.  Big Data is not just about large volumes, it is also characterized by velocity, variety of format, and variability of meaning.
 
Cisco’s Senior Vice President of Engineering and Chief Technology Officer, Padmasree Warrior,  identified Big Data in the future as being generated from two main areas: things or people.  In the area of things, enormous amounts of data sets are being generated from scientific research in areas like astrophysics, and from hardware in smart grids, and from machine to machine communications.  People-generated Big Data sets will come from things like social media, genetics, health care records, and general communication.
Digg This
Reddit This
Stumble Now!
Buzz This
Vote on DZone
Share on Facebook
Bookmark this on Delicious
Kick It on DotNetKicks.com
Shout it
Share on LinkedIn
Bookmark this on Technorati
Post on Twitter
Google Buzz (aka. Google Reader)

Leave a Reply

Your email address will not be published. Required fields are marked *

*