Access and Feeds

Data Lakes: Recipe for Big Data Success or Roadblock to Data Analysis?

By Dick Weisinger

The Data Lakes Market is expected to grow from $2.53 billion to $8.81 billion over the next five years, at an average yearly growth rate of more than 28 percent, according to an estimate by analyst firm MarketsandMarkets.  The analysis breaks down the market into subsegments of: Data Discovery, Data Integration, Data Lakes Analytics, Data Visualization.

A data lake is a vast storage repository that holds raw data in its original format in a flat architecture.  A data lake contrasts to data warehouses which typically involves processing and curating of data that is then stored hierarchically in files and folders.

The term “data lake” was an idea first explained by  James Dixon, CTO at Pentaho.  He said that “If you think of a data mart as a store of bottled water – cleansed and packaged and structured for easy consumption – the data lake is a large body of water in a more natural state.”

Nick Heudecker, research director at Gartner, described the utility of data lakes, saying that “in broad terms, data lakes are marketed as enterprise-wide data management platforms for analyzing disparate sources of data in its native format.  The idea is simple: instead of placing data in a purpose-built data store, you move it into a data lake in its original format. This eliminates the upfront costs of data ingestion, like transformation. Once data is placed into the lake, it’s available for analysis by everyone in the organization.”

The popularity of data lakes can be attributed to businesses looking for better competitiveness and agility, increased adoption of the Internet of Things (IoT), and growing volumes of data.

But not everyone agrees that data lakes are good.    Andrew White, Gartner vice president, said that “data is simply dumped into the data lake… Getting value out of the data remains the responsibility of the business end user. Of course, technology could be applied or added to the lake to do this, but without at least some semblance of information governance, the lake will end up being a collection of disconnected data pools or information silos all in one place.”

Dan Woods, writer and technologist, wrote on Forbes that “with data lakes there’s no inherent way to prioritize what data is going into the supply chain and how it will eventually be used. The result is like a museum with a huge collection of art, but no curator with the eye to tell what is worth displaying and what’s not.”

 

Digg This
Reddit This
Stumble Now!
Buzz This
Vote on DZone
Share on Facebook
Bookmark this on Delicious
Kick It on DotNetKicks.com
Shout it
Share on LinkedIn
Bookmark this on Technorati
Post on Twitter
Google Buzz (aka. Google Reader)
One comment on “Data Lakes: Recipe for Big Data Success or Roadblock to Data Analysis?
  1. Wardah says:

    Interesting piece of information. Big data if properly managed can retrieve positive results beneficial for planning strategies for the future of a business. I also read an article https://www.rokittastra.com/winning-big-data-organization-operating-20-capacity/ which provides with interesting info about how organizations face failure in the strategies relevant to big data.

Leave a Reply

Your email address will not be published. Required fields are marked *

*