{"id":4938,"date":"2014-10-15T11:00:02","date_gmt":"2014-10-15T19:00:02","guid":{"rendered":"http:\/\/formtek.com\/blog\/?p=4938"},"modified":"2014-10-10T12:45:18","modified_gmt":"2014-10-10T20:45:18","slug":"big-data-real-time-stream-processing-with-datatorrent","status":"publish","type":"post","link":"https:\/\/formtek.com\/blog\/big-data-real-time-stream-processing-with-datatorrent\/","title":{"rendered":"Big Data: Real-Time Stream Processing with DataTorrent"},"content":{"rendered":"<p>Traditional computing techniques are challenged when it comes to handling large volumes of real-time data. \u00a0Stream processing is an evolving programming paradigm that attempts to take on the challenge of processing large real-time data. \u00a0While Hadoop is one technology that&#8217;s been used for working with large data sets, it hasn&#8217;t been until the recent introduction of Apache Hadoop\u00a0<a title=\"YARN\" href=\"http:\/\/hadoop.apache.org\/docs\/current\/hadoop-yarn\/hadoop-yarn-site\/YARN.html\" target=\"_blank\">YARN <\/a>and <a title=\"Apache Storm\" href=\"https:\/\/storm.incubator.apache.org\/\" target=\"_blank\">Apache Storm<\/a> that big data projects have been able to begin to do anything other than static batch processing.<\/p>\n<p>Stream processing is <a title=\"Stream Processing Applications\" href=\"http:\/\/www-03.ibm.com\/systems\/infrastructure\/us\/en\/technical-breakthroughs\/stream-processing.html\" target=\"_blank\">particularly suitable for applications<\/a> that have these kinds of requirements:<\/p>\n<ul>\n<li>Compute intensive [high ratio of operations to I\/O]<\/li>\n<li>Data Parallelism<\/li>\n<li>Data Pipelining where data moves continuously from producers to downstream consumers<\/li>\n<\/ul>\n<p>One company focused on building a platform for enabling large real-time applications is DataTorrent. \u00a0DataTorrent is a startup company that has created a data-stream processing platform built with Apache Hadoop and YARN. \u00a0The founders of the company come out of Yahoo! and were involved with the original development of Hadoop. \u00a0DataTorrent addresses markets that need to handle large volumes of real-time data, like the financial services industry, manufacturing, and the Internet of Things (IoT).<\/p>\n<p>The 1.0 version of the DataTorrent product focused on building a high performance real-time streaming platform. \u00a0With a 35-node cluster, DataTorrent was able benchmark 1.6 billion events\/second processing with sub-millisecond response times.\u00a0The platform is designed to be robust with stateful fault tolerance. \u00a0The focus of the 1.0 DataTorrent product was to drive the core technical aspects of the platform and to make the software available in as many ways, shapes and forms as possible.<\/p>\n<p>Barely now three months after the 1.0 release, DataTorrent is making available today the 2.0 release of their product.<\/p>\n<p><a title=\"John Fanelli bio\" href=\"https:\/\/www.datatorrent.com\/company\/\" target=\"_blank\">John Fanelli<\/a>,\u00a0VP of Marketing at DataTorrent, said that\u00a0&#8220;the story around our 2.0 release is really the customer and how to make it more accessible for them and to take it to the next level, allowing non-developers to participate in the creating of real-time streaming applications&#8230; When you talk to a customer, they&#8217;ll say &#8216;Can you do my data? Coming from this source. \u00a0It looks like this size. And coming in at this speed.&#8217; \u00a0And the speed can be millions of events per second to every five seconds for getting the data. \u00a0The data doesn&#8217;t always have to be coming in real time. What we really drive is faster insight and real time action&#8230; \u00a0this release also adds additional operators into our open source community&#8230; \u00a0You write your streaming app by connecting a number of operators that the data streams through. \u00a0We have over 500 of these operators in an open-source GitHub repository, and that really provides our customers a great starting block.&#8221;<\/p>\n<p>Coincident with the 2.0 DataTorrent release is also an announcement of private-beta availability of a graphical assembly application called DaVinci and a real-time visual dashboard designer, code-named Michelangelo.<\/p>\n<div class=\"lightsocial_container\"><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/digg.com\/submit?url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F&amp;title=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/digg.png\" alt=\"Digg This\" title=\"Digg This\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.reddit.com\/submit?url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F&amp;title=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/reddit.png\" alt=\"Reddit This\" title=\"Reddit This\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.stumbleupon.com\/submit?url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F&amp;title=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/stumbleupon.png\" alt=\"Stumble Now!\" title=\"Stumble Now!\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/buzz.yahoo.com\/buzz?targetUrl=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F&amp;headline=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/yahoo_buzz.png\" alt=\"Buzz This\" title=\"Buzz This\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.dzone.com\/links\/add.html?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/dzone.png\" alt=\"Vote on DZone\" title=\"Vote on DZone\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.facebook.com\/sharer.php?t=&amp;u=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/facebook.png\" alt=\"Share on Facebook\" title=\"Share on Facebook\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/delicious.com\/save?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/delicious.png\" alt=\"Bookmark this on Delicious\" title=\"Bookmark this on Delicious\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.dotnetkicks.com\/kick\/?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/dotnetkicks.png\" alt=\"Kick It on DotNetKicks.com\" title=\"Kick It on DotNetKicks.com\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/dotnetshoutout.com\/Submit?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/dotnetshoutout.png\" alt=\"Shout it\" title=\"Shout it\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.linkedin.com\/shareArticle?mini=true&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F&amp;title=&amp;summary=&amp;source=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/linkedin.png\" alt=\"Share on LinkedIn\" title=\"Share on LinkedIn\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.technorati.com\/faves?add=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/technorati.png\" alt=\"Bookmark this on Technorati\" title=\"Bookmark this on Technorati\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/twitter.com\/home?status=Reading+https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/twitter.png\" alt=\"Post on Twitter\" title=\"Post on Twitter\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.google.com\/buzz\/post?url=https%3A%2F%2Fformtek.com%2Fblog%2Fbig-data-real-time-stream-processing-with-datatorrent%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/google_buzz.png\" alt=\"Google Buzz (aka. Google Reader)\" title=\"Google Buzz (aka. Google Reader)\" \/><\/a><\/div><\/div>","protected":false},"excerpt":{"rendered":"<p>Traditional computing techniques are challenged when it comes to handling large volumes of real-time data. \u00a0Stream processing is an evolving programming paradigm that attempts to take on the challenge of processing large real-time data. \u00a0While Hadoop is one technology that&#8217;s<span class=\"ellipsis\">&hellip;<\/span><\/p>\n<div class=\"read-more\"><a href=\"https:\/\/formtek.com\/blog\/big-data-real-time-stream-processing-with-datatorrent\/\">Read more &#8250;<\/a><\/div>\n<p><!-- end of .read-more --><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3,1],"tags":[],"class_list":["post-4938","post","type-post","status-publish","format-standard","hentry","category-big-data","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts\/4938","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/comments?post=4938"}],"version-history":[{"count":0,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts\/4938\/revisions"}],"wp:attachment":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/media?parent=4938"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/categories?post=4938"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/tags?post=4938"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}