{"id":4329,"date":"2013-11-19T07:00:08","date_gmt":"2013-11-19T15:00:08","guid":{"rendered":"http:\/\/formtek.com\/blog\/?p=4329"},"modified":"2013-11-15T20:43:12","modified_gmt":"2013-11-16T04:43:12","slug":"big-data-interacting-with-hadoop-2-data-using-standard-ansi-sql","status":"publish","type":"post","link":"https:\/\/formtek.com\/blog\/big-data-interacting-with-hadoop-2-data-using-standard-ansi-sql\/","title":{"rendered":"Big Data: Interacting with Hadoop 2 Data Using Standard ANSI SQL"},"content":{"rendered":"<p><a title=\"Cascading Java Application Framework\" href=\"http:\/\/www.concurrentinc.com\/cascading\/\" target=\"_blank\">Cascading is an Apache-licensed application framework\u00a0<\/a>\u00a0for building rich data processing and machine learning applications that run on Hadoop. \u00a0Cascading applications are built using a simple API that can be called from any\u00a0JVM-based language. \u00a0Over the past five years, it has been developed and supported commercially by\u00a0<a title=\"Concurrent Inc\" href=\"http:\/\/www.concurrentinc.com\/\" target=\"_blank\">Concurrent, Inc<\/a>. \u00a0The flow of\u00a0Cascading is to first capture data from &#8216;sources&#8217; , to then pass that data through &#8216;pipes&#8217; where it is processed, and finally to push the results into output files or &#8216;sinks&#8217;. \u00a0This flow of data is known as the &#8216;source-pipe-sink&#8217; paradigm. \u00a0Cascading runs as an abstraction at a higher level than MapReduce, so that while Cascading applications ultimately execute MapReduce jobs, when the Cascading application is written, no explicit interactions with MapReduce need to be programmed.<\/p>\n<p><a title=\"Chris Wensel bio\" href=\"www.linkedin.com\/in\/cwensel\" target=\"_blank\">Chris Wensel<\/a>, Founder and CTO of Concurrent, <a title=\"Chris Wensel quote no Concurrent Cascading\" href=\"http:\/\/www.concurrentinc.com\/posts\/2012\/06\/05\/concurrent-launches-cascading-2-0\/\" target=\"_blank\">said that <\/a>&#8220;building applications on Hadoop, despite its growing adoption in the enterprise, is notoriously difficult. We are driving the future of application development and management on Hadoop, by allowing enterprises to quickly extract meaningful information from large amounts of distributed data and better understand the business implications. We make it easy for developers to build powerful data processing applications for Hadoop, without requiring months spent learning about the intricacies of MapReduce.&#8221;<\/p>\n<p>More than 110,000 user downloads of Cascading are made every month. \u00a0Cascading is used by businesses like\u00a0<a title=\"Twitter on Cascading\" href=\"http:\/\/www.concurrentinc.com\/case-studies\/twitter\/\" target=\"_blank\">Twitter<\/a>, eBay,\u00a0<a title=\"The Climate Corporation on Cascading\" href=\"http:\/\/www.concurrentinc.com\/case-studies\/climate-corp\/\" target=\"_blank\">The Climate Corporation<\/a>, Square and\u00a0<a title=\"Etsy on Cascading\" href=\"http:\/\/www.concurrentinc.com\/case-studies\/etsy\/\" target=\"_blank\">Etsy<\/a>\u00a0for managing some or all of their Big Data requirements. \u00a0In fact, all of Twitter&#8217;s revenue-generating applications have been built with Cascading.<\/p>\n<p>Today, Cascading 2.5 is being introduced &#8212; a version-number jump from the previously available 2.2 point release. \u00a0The jump in numbering was intended to emphasise the significance of some of the new features in the release. Most significantly, the 2.5 release will include support for Hadoop 2 and <a title=\"Apache YARN\" href=\"http:\/\/hadoop.apache.org\/docs\/current\/hadoop-yarn\/hadoop-yarn-site\/YARN.html\" target=\"_blank\">YARN<\/a>. \u00a0Other highlights of the new 2.5 Cascading release include:<\/p>\n<ul>\n<li>Performance improvements for complex join operations and optimizations to dynamically partition and store processed data more efficiently on HDFS.<\/li>\n<li>Broad compatibility with other Hadoop vendors and Hadoop as a service providers, including Cloudera, Hortonworks, MapR, Intel, Altiscale, Qubole and Amazon EMR<\/li>\n<\/ul>\n<p>Coincident with the release of Cascading 2.5, Concurrent is also making another product, Cascading Lingual, generally available. \u00a0The Lingual product is an add-on to Cascading that enables a complete ANSI-SQL interface for interacting with Hadoop data. \u00a0Compatibility with standard SQL means that SQL developed in traditional relational databases can be brought over and used as-is within Lingual. \u00a0Like Cascading, Lingual comes with the Apache 2.0 license.<\/p>\n<p>Concurrent described the benefits of the Lingual product by saying that &#8220;Cascading Lingual provides out-of-the-box support for JDBC. Enterprises that have invested millions of dollars in business intelligence (BI) tools, such as Pentaho, Jaspersoft and Cognos, and training can now also access their data on Hadoop through standard SQL interface.&#8221;<\/p>\n<p><a title=\"Andre Kelpe bio \" href=\"http:\/\/www.linkedin.com\/in\/akelpe\">Andr\u00e9 Kelpe<\/a>, software engineer for Concurrent, <a title=\"Lingual design goals\" href=\"https:\/\/gist.github.com\/fs111\/7013230\">summarized three of the design goals<\/a> for the Cascading Lingual product:<\/p>\n<ul>\n<li>Enable immediate ANSI SQL query access to data<\/li>\n<li>Simplified System and Data Integration with read\/writes from hdfs, jdbc, memcached, HBase, and redshift<\/li>\n<li>Simplified migration of existing SQL within Cascading<\/li>\n<\/ul>\n<p>A YouTube on-line demo of Cascading Lingual can be found\u00a0<a title=\"Youtube demo of Lingual\" href=\"http:\/\/www.youtube.com\/watch?v=mV76u8avx6Y\" target=\"_blank\">here<\/a>.<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<div class=\"lightsocial_container\"><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/digg.com\/submit?url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F&amp;title=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/digg.png\" alt=\"Digg This\" title=\"Digg This\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.reddit.com\/submit?url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F&amp;title=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/reddit.png\" alt=\"Reddit This\" title=\"Reddit This\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.stumbleupon.com\/submit?url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F&amp;title=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/stumbleupon.png\" alt=\"Stumble Now!\" title=\"Stumble Now!\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/buzz.yahoo.com\/buzz?targetUrl=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F&amp;headline=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/yahoo_buzz.png\" alt=\"Buzz This\" title=\"Buzz This\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.dzone.com\/links\/add.html?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/dzone.png\" alt=\"Vote on DZone\" title=\"Vote on DZone\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.facebook.com\/sharer.php?t=&amp;u=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/facebook.png\" alt=\"Share on Facebook\" title=\"Share on Facebook\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/delicious.com\/save?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/delicious.png\" alt=\"Bookmark this on Delicious\" title=\"Bookmark this on Delicious\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.dotnetkicks.com\/kick\/?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/dotnetkicks.png\" alt=\"Kick It on DotNetKicks.com\" title=\"Kick It on DotNetKicks.com\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/dotnetshoutout.com\/Submit?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/dotnetshoutout.png\" alt=\"Shout it\" title=\"Shout it\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.linkedin.com\/shareArticle?mini=true&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F&amp;title=&amp;summary=&amp;source=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/linkedin.png\" alt=\"Share on LinkedIn\" title=\"Share on LinkedIn\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.technorati.com\/faves?add=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/technorati.png\" alt=\"Bookmark this on Technorati\" title=\"Bookmark this on Technorati\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/twitter.com\/home?status=Reading+https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/twitter.png\" alt=\"Post on Twitter\" title=\"Post on Twitter\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.google.com\/buzz\/post?url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-interacting-with-hadoop-2-data-using-standard-ansi-sql%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/google_buzz.png\" alt=\"Google Buzz (aka. Google Reader)\" title=\"Google Buzz (aka. Google Reader)\" \/><\/a><\/div><\/div>","protected":false},"excerpt":{"rendered":"<p>Cascading is an Apache-licensed application framework\u00a0\u00a0for building rich data processing and machine learning applications that run on Hadoop. \u00a0Cascading applications are built using a simple API that can be called from any\u00a0JVM-based language. \u00a0Over the past five years, it has<span class=\"ellipsis\">&hellip;<\/span><\/p>\n<div class=\"read-more\"><a href=\"https:\/\/formtek.com\/blog\/big-data-interacting-with-hadoop-2-data-using-standard-ansi-sql\/\">Read more &#8250;<\/a><\/div>\n<p><!-- end of .read-more --><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[],"class_list":["post-4329","post","type-post","status-publish","format-standard","hentry","category-big-data"],"_links":{"self":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts\/4329","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/comments?post=4329"}],"version-history":[{"count":0,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts\/4329\/revisions"}],"wp:attachment":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/media?parent=4329"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/categories?post=4329"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/tags?post=4329"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}