Showing posts with label SLA. Show all posts
Showing posts with label SLA. Show all posts

Thursday, October 03, 2013

Why cloud?

I have helped large organizations relocate enterprise applications to the cloud, helped small companies make cloud vendor selections and have been CTO of an IaaS cloud.   I also have experience running large IT organizations and have provided consultation to companies on many aspects of their internal processes and technology.   I have been asked many times "why cloud?"   While the question has been asked many times and there are many biased answers, my view through the fog provides me some specific answers.

While the answer will be different for every organization, but the parameters fall into a combination of five areas:

  • ROI - some organizations can show an ROI for moving infrastructure and applications to the cloud.   Typically, the cost savings come from reductions in employee costs (harder to justify since the downturn in the economy in 2009 when many companies cut IT resources and possibly reduced service levels).   There are some potential savings from CapEx avoidance on refreshes, potential cost reductions from colo-sites.   I have built a fairly elaborate model that always needs tweaking based on a specific business situation.
  • Agility - I believe this is the real the driver behind the long term growth of cloud.   More today than ever before, businesses are required to respond rapidly to competitive changes and business opportunities.   The legacy model of IT that is slowed by CapEx anchors of previous purchases and the availability of IT resources within the organization reduces agility.   Cloud is the newest method of outsourcing IT with portentially a new benefit of varying cost and changing vendors more easily.   There are immediate agility benefits from a move to cloud.   The ensuring that business remain agile with a transition to cloud requires planning and management.
  • Improved Service Levels - Cloud vendors will live or die based on their service levels.   Even AWS will need to have improved availability and MTTR going forward.   The ability of an internal IT organization to justify the cost of improved service levels has always been limited because it is managed as a cost center.   Regardless of the size of an organization, I find that there are never enough IT resources.   IT as a The shift to cloud will deliver the SLAs that are desired at an economy of scale that most businesses cannot afford.  Improved service levels will also come from cloud vendors having more resources, more specialized resources.
  • Improved Data Security - Yes, I said better data security in the cloud than in the enterprise.   As a CTO of an IaaS cloud, I knew that we had to be better than the internal IT organization at protecting data.   I also know, having been inside many businesses, they they actually struggle to achieve the data security that they demand of cloud vendors.   Transition to cloud is good for all in this area as well.
  • Improved BCP/DR - Many organizations are required to have business continuity plans, disaster recovery plans and an ability to execute on them.   Even some privately help companies that do not have compliance requirements driving them in this direction have decided that it is smart to have this insurance policy to ensure survivability.   Virtualization and cloud make BCP and DR much easier and provide many more options for controlling the cost of a BCP/DR program.



Tuesday, April 26, 2011

Amazon EC2 Outage and Cloud Strategy

Last Friday, Amazon experienced a partial outage of its cloud infrastructure.   Here the initial update and the closing updates:


Event Issue
"The problem started with a "networking event" that led to problems with how data is mirrored: We'd like to provide additional color on what were working on right now (please note that we always know more and understand issues better after we fully recover and dive deep into the post mortem). A networking event early this morning triggered a large amount of re-mirroring of EBS [Elastic Block Storage] volumes in US-EAST-1. This re-mirroring created a shortage of capacity in one of the US-EAST-1 Availability Zones, which impacted new EBS volume creation as well as the pace with which we could re-mirror and recover affected EBS volumes. Additionally, one of our internal control planes for EBS has become inundated such that it's difficult to create new EBS volumes and EBS backed instances. We are working as quickly as possible to add capacity to that one Availability Zone to speed up the re-mirroring, and working to restore the control plane issue. We're starting to see progress on these efforts, but are not there yet. We will continue to provide updates when we have them."

Closing update from Amazon:

As we posted last night, EBS (Elastic Block Store) is now operating normally for all APIs and recovered EBS volumes. The vast majority of affected volumes have now been recovered. We’re in the process of contacting a limited number of customers who have EBS volumes that have not yet recovered and will continue to work hard on restoring these remaining volumes…
We are digging deeply into the root causes of this event and will post a detailed post mortem.

One of the unfortunate realities of infrastructure and operations is that the goal will always be 100% uptime for all infrastructures but it cannot be achieved.   The SLAs for infrastructure and operations is very unlikely to be 100%.   The strategic question will always be what SLAs can be afforded, what is the impact to business agility for the target SLAs and what can be improved from a people, process and technology perspective to achieve the business goals and minimize cost. 

Because there are clear ties between performance, availability and security objectives and the success of outsource cloud infrastructure and operations, I believe that public cloud will outperform internal infrastructure over time.   This does not lessen the requirement for internal roles of architecture, end-to-end management of performance, availability and security, and vendor management.   These roles will increase in importance within organizations. 

The current Amazon issue re-emphasizes that a cloud strategy needs to include

  • Clear and continuous risk management program for IT
  • Enterprise change, incident, problem, release and configuration management process re-engineering
  • End-to-end SLA and systems management
  • Server provisioning process and technology
  • Patching process
  • Server configuration baselining and auditing
  • Repurposing of servers
  • Disaster recovery planning and testing

Sunday, February 21, 2010

The 3 foundation measurements of IT

There are 3 foundation measurements of IT.   These are the indicators of successful delivery to the business. 
  • Requirements delivered for changes to application functionality.   This can be measured as internal or external customer satisfaction, revenue growth attributed to application changes, function points, user stories, and use cases delivered.
  • Service level achievement.   This can be measured as achieving Service Level Agreements and the metrics will typically be in application performance, application availability, and time to close issues.
  • Cost.  Coupled with the first or second bullet point, it would be measured as ROI.   Otherwise, it is typically a historically trended set of numbers.   
The concepts are not complicated.   Execution is the challenge.