Access and Feeds

Technology: Dealing with 'Big Data'

By Dick Weisinger

The amount of data that companies and individuals are creating is growing incredibly quickly.  Some organizations have managed to collect massive data sets.  O’Reilly refers to these huge data sets as ‘Big Data’ — data sets that are in the hundreds of gigabytes, terabytes, or much much more.  Big data is being driven by data collection coming from web crowds, sensor networks, informatics and geodata.  O’Reilly says that companies that are able to tap into and effectively analyze data will be at a competitive advantage to other companies, and uses Google as an example.  Google’s success has been all about their ability of being able to make sense out of massive amounts of data.

Big data analysis techniques will especially benefit data analysis projects like predictive models and ratings engines.  When large amounts of data can be incorporated into the decision making process, better decisions will be made. And analysis of very large data sets will lead to revolutionary breakthroughs in commerce, science and society.

In the realm of business, profitability comes through customers, and the performance of business is controlled by these three dimensions:
– Increasing the number of custoemrs
– Increasing the profitability per customer
– Retaining customers for longer periods
Being able to better use and analyze customer data will let organizations improve performance — and analysis of Big Data collected from customers will help.  Forbes says that “‘Big Data’ is going to be ‘Big Business'”.

Businesses have grabbed onto the idea and are scrambling to get their hands on better tools for data analysis.  Hadoop, an open-source application, is one tool that is creating a lot of excitement. Hadoop is a framework for enabling the connection and scale-up of cheap hardware and tools for analyzing unstructured and structured data.  While Relational Databases are good at analyzing small subsets of data, Hadoop tries to examine and interpret all available data.  It uses an algorithm called MapReduce, and works by being able to distribute data for analysis to a large number of servers that can all work in parallel, collating the results and then presenting the final result.  The architecture allows Hadoop to operate at very high data availability and also to be able to process and analyze enormous amounts of information quickly.

Hadoop is one technique that has gained popularity, and there has been a cottage industry of tools and businesses that have been built on top of Hadoop.  But this field is in its infancy.  Expect to see a lot of new ideas and activity around collecting and analyzing ‘Big Data’.

Leave a Reply

Your email address will not be published. Required fields are marked *

*