Access and Feeds

Cloud Computing: Lessons from the Amazon AWS Outage

By Dick Weisinger

It’s been about a month since Amazon suffered a major multi-day outage that caused many cloud-based sites significant grief.  Services took as long as four days until they were back up and running after the incident, many users reported continued problems and sluggish response for many days later.  Was the outage just a blip in the grand scheme of things or was it an event that will prove a setback to cloud technology?

For a small group of Amazon AWS users, the outage hit very hard.  The Amazon Health Dashboard reported that “We have completed our remaining recovery efforts and though we’ve recovered nearly all of the stuck volumes, we’ve determined that a small number of volumes (0.07% of the volumes in our US-East Region) will not be fully recoverable. We’re in the process of contacting these customers.”

For some vendors like SAP, the outage proved to be a relief.  SAP has been struggling to sell its SaaS-based Business ByDesign ERP software.  Sanjay Poonen, head of SAP’s global solutions business, said “We’ll have to work harder to make people comfortable with where cloud computing is.”

IDC analyst Matthew Eastwood called the outage a “wake-up call for cloud computing” and said “it will force a conversation in the industry.”

IDC analyst David Bradshaw said that “if this is proved an exceptional event that is not necessarily going to be duplicated by other service providers, there will be less damage to the cloud computing industry than if the outage is seen as a problem in the general cloud infrastructure.”

Lew Moorman, President of Rackspace, said that “computing equivalent of an airplane crash,” which is a major episode with widespread damage, but doesn’t change the fact that airline travel is still safer than traveling in a car.  The cost of a great number of small outages that occur every day behind closed firewalls easily outweighs the impact of this one event that happened to impact many users simultaneously.”

Dave Jilk, CEO at Standing Cloud, summed up lessons to be learnt from the AWS outage:

  • Amazon is not infallible, and the cloud is not magic.
  • Amazon is not the only IaaS provider, and your application should be able to run on more than one.
  • Cloud deployments must be automated and should take cloud server reliability characteristics into account
Digg This
Reddit This
Stumble Now!
Buzz This
Vote on DZone
Share on Facebook
Bookmark this on Delicious
Kick It on DotNetKicks.com
Shout it
Share on LinkedIn
Bookmark this on Technorati
Post on Twitter
Google Buzz (aka. Google Reader)

Leave a Reply

Your email address will not be published. Required fields are marked *

*