{"id":5332,"date":"2015-04-13T07:00:26","date_gmt":"2015-04-13T15:00:26","guid":{"rendered":"http:\/\/formtek.com\/blog\/?p=5332"},"modified":"2015-03-06T11:48:17","modified_gmt":"2015-03-06T19:48:17","slug":"big-data-speeding-up-big-data-processing-with-apache-spark","status":"publish","type":"post","link":"https:\/\/formtek.com\/blog\/big-data-speeding-up-big-data-processing-with-apache-spark\/","title":{"rendered":"Big Data: Speeding up Big Data Processing with Apache Spark"},"content":{"rendered":"<p><a title=\"Apache Spark\" href=\"https:\/\/spark.apache.org\/\" target=\"_blank\"><a href=\"http:\/\/formtek.com\/blog\/wp-content\/uploads\/2015\/01\/spark-logo.png\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-5241 alignleft\" alt=\"spark-logo\" src=\"http:\/\/formtek.com\/blog\/wp-content\/uploads\/2015\/01\/spark-logo.png\" width=\"155\" height=\"82\" \/><\/a>Apache Spark<\/a> is quickly gaining a strong following among Big Data users. \u00a0The big advantage of using Spark is that it allows in-memory processing which greatly speeds up the ingestion and processing of data. \u00a0Spark is also a bit more straight forward in terms of builds and the creation of data job workflows. \u00a0Apache Spark&#8217;s tagline is &#8220;Lightening Fast Clustered Computing&#8221;.<\/p>\n<p><a title=\"Ion Stoica bio\" href=\"http:\/\/www.cs.berkeley.edu\/~istoica\/\" target=\"_blank\">Ion Stoica<\/a>, Co-Founder and CEO of Databricks, <a title=\"Ion Stoica explains what Apache Spark is\" href=\"http:\/\/www.forbes.com\/sites\/brucerogers\/2015\/03\/05\/databricks-aims-to-become-the-platform-for-big-data\/\" target=\"_blank\">told Forbes that<\/a> \u00a0&#8220;Spark is a parallel execution engine that is better than Hadoop MapReduce in three dimensions: First, it\u2019s faster, because it is optimized to work efficiently with data stored both in memory and on disk. Spark holds the terabyte sort benchmark record, by beating the time of the previous record by 3x using 10x fewer machines. \u00a0The second advantage is that it provides a more powerful and flexible API than MapReduce, which makes it much easier for developers to write sophisticated applications. Typically, it takes between two to five times fewer lines of code to write the same application in Spark than in Hadoop MapReduce. Finally, Spark unifies a variety of computation models. It goes far beyond batch computation, and through a set of libraries, it supports many other workloads, including streaming, interactive queries, machine learning, and graph processing. \u00a0What makes all of these possible is a very flexible and powerful core engine which can execute large scale jobs in subseconds.&#8221;<\/p>\n<p>Stoica added that &#8220;Spark\u00a0also has a more general and easy-to-use API. So when you write applications in Spark, you don&#8217;t need to cast them as a bunch of maps and reduces. You can almost write arbitrary applications. \u00a0When we do regular surveys and ask people why they like Spark, half say it&#8217;s speed and half say it&#8217;s ease of use.&#8221;<\/p>\n<p>Databricks, a startup founded by the creators of Spark, has recently announced the availability of a cloud platform based around Apache Spark and a Databricks workspace. \u00a0It&#8217;s a way to simplify the interaction a user has with Big Data. \u00a0There&#8217;s no need to have to directly interact with an Hadoop cluster. \u00a0Once users upload data to the a project in the Databrick platform, they&#8217;re able to start interacting with it and begin creating visuals, like charts and dashboards. \u00a0It&#8217;s possible for the user to schedule jobs via the job launcher to ensure that Spark jobs get run at specific times.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<div class=\"lightsocial_container\"><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/digg.com\/submit?url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F&amp;title=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/digg.png\" alt=\"Digg This\" title=\"Digg This\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.reddit.com\/submit?url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F&amp;title=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/reddit.png\" alt=\"Reddit This\" title=\"Reddit This\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.stumbleupon.com\/submit?url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F&amp;title=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/stumbleupon.png\" alt=\"Stumble Now!\" title=\"Stumble Now!\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/buzz.yahoo.com\/buzz?targetUrl=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F&amp;headline=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/yahoo_buzz.png\" alt=\"Buzz This\" title=\"Buzz This\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.dzone.com\/links\/add.html?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/dzone.png\" alt=\"Vote on DZone\" title=\"Vote on DZone\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.facebook.com\/sharer.php?t=&amp;u=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/facebook.png\" alt=\"Share on Facebook\" title=\"Share on Facebook\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/delicious.com\/save?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/delicious.png\" alt=\"Bookmark this on Delicious\" title=\"Bookmark this on Delicious\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.dotnetkicks.com\/kick\/?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/dotnetkicks.png\" alt=\"Kick It on DotNetKicks.com\" title=\"Kick It on DotNetKicks.com\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/dotnetshoutout.com\/Submit?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/dotnetshoutout.png\" alt=\"Shout it\" title=\"Shout it\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.linkedin.com\/shareArticle?mini=true&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F&amp;title=&amp;summary=&amp;source=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/linkedin.png\" alt=\"Share on LinkedIn\" title=\"Share on LinkedIn\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.technorati.com\/faves?add=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/technorati.png\" alt=\"Bookmark this on Technorati\" title=\"Bookmark this on Technorati\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/twitter.com\/home?status=Reading+https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/twitter.png\" alt=\"Post on Twitter\" title=\"Post on Twitter\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.google.com\/buzz\/post?url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-speeding-up-big-data-processing-with-apache-spark%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/google_buzz.png\" alt=\"Google Buzz (aka. Google Reader)\" title=\"Google Buzz (aka. Google Reader)\" \/><\/a><\/div><\/div>","protected":false},"excerpt":{"rendered":"<p>Apache Spark is quickly gaining a strong following among Big Data users. \u00a0The big advantage of using Spark is that it allows in-memory processing which greatly speeds up the ingestion and processing of data. \u00a0Spark is also a bit more<span class=\"ellipsis\">&hellip;<\/span><\/p>\n<div class=\"read-more\"><a href=\"https:\/\/formtek.com\/blog\/big-data-speeding-up-big-data-processing-with-apache-spark\/\">Read more &#8250;<\/a><\/div>\n<p><!-- end of .read-more --><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[],"class_list":["post-5332","post","type-post","status-publish","format-standard","hentry","category-big-data"],"_links":{"self":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts\/5332","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/comments?post=5332"}],"version-history":[{"count":0,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts\/5332\/revisions"}],"wp:attachment":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/media?parent=5332"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/categories?post=5332"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/tags?post=5332"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}