Access and Feeds

Apache Spark: Turbocharging Past Hadoop’s MapReduce

By Dick Weisinger

Apache Spark appears to be igniting Big Data projects in 2015.  Spark is an Apache open source engine built for Hadoop to perform sophisticated data analysis and analytics.  Compared to Hadoop-only applications, Spark applications are as much as 10 to 100 times faster, and they’re easier to program.  But Spark isn’t an alternative to Hadoop, it’s a tool that turbocharges what Hadoop does.

spark-logoRecently TypeSafe took a survey of how popular Spark has become and asked developers how they’re using it.  Some of the findings from that report are as follows:

  • Spark is very popular now.  71 percent of businesses surveyed by TypeSafe said that had done at least some amount of hands-on evaluation with Spark, and 35 percent of businesses are using it or planning to use it.
  • Spark is fast.  Spark has significantly faster performance and event stream processing compared to Hadoop’s standard MapReduce.
  • Few barriers to adoption.  Developers said the major barrier currently is just lack of deep experience with the technology and a lack of support for and integration with other middleware, like message queues and databases.

Dr. Dean Wampler, Big Data architect at Typesafe, said that “the need to process Big Data faster has largely fueled the intense developer interest in Spark.  Hadoop’s historic focus on batch processing of data was well supported by MapReduce, but there is an appetite for more flexible developer tools to support the larger market of ‘mid-size’ datasets and use cases that call for real-time processing.”

Digg This
Reddit This
Stumble Now!
Buzz This
Vote on DZone
Share on Facebook
Bookmark this on Delicious
Kick It on DotNetKicks.com
Shout it
Share on LinkedIn
Bookmark this on Technorati
Post on Twitter
Google Buzz (aka. Google Reader)

Leave a Reply

Your email address will not be published. Required fields are marked *

*