Showing posts with label Cloud. Show all posts
Showing posts with label Cloud. Show all posts

Thursday, October 03, 2013

Why cloud?

I have helped large organizations relocate enterprise applications to the cloud, helped small companies make cloud vendor selections and have been CTO of an IaaS cloud.   I also have experience running large IT organizations and have provided consultation to companies on many aspects of their internal processes and technology.   I have been asked many times "why cloud?"   While the question has been asked many times and there are many biased answers, my view through the fog provides me some specific answers.

While the answer will be different for every organization, but the parameters fall into a combination of five areas:

  • ROI - some organizations can show an ROI for moving infrastructure and applications to the cloud.   Typically, the cost savings come from reductions in employee costs (harder to justify since the downturn in the economy in 2009 when many companies cut IT resources and possibly reduced service levels).   There are some potential savings from CapEx avoidance on refreshes, potential cost reductions from colo-sites.   I have built a fairly elaborate model that always needs tweaking based on a specific business situation.
  • Agility - I believe this is the real the driver behind the long term growth of cloud.   More today than ever before, businesses are required to respond rapidly to competitive changes and business opportunities.   The legacy model of IT that is slowed by CapEx anchors of previous purchases and the availability of IT resources within the organization reduces agility.   Cloud is the newest method of outsourcing IT with portentially a new benefit of varying cost and changing vendors more easily.   There are immediate agility benefits from a move to cloud.   The ensuring that business remain agile with a transition to cloud requires planning and management.
  • Improved Service Levels - Cloud vendors will live or die based on their service levels.   Even AWS will need to have improved availability and MTTR going forward.   The ability of an internal IT organization to justify the cost of improved service levels has always been limited because it is managed as a cost center.   Regardless of the size of an organization, I find that there are never enough IT resources.   IT as a The shift to cloud will deliver the SLAs that are desired at an economy of scale that most businesses cannot afford.  Improved service levels will also come from cloud vendors having more resources, more specialized resources.
  • Improved Data Security - Yes, I said better data security in the cloud than in the enterprise.   As a CTO of an IaaS cloud, I knew that we had to be better than the internal IT organization at protecting data.   I also know, having been inside many businesses, they they actually struggle to achieve the data security that they demand of cloud vendors.   Transition to cloud is good for all in this area as well.
  • Improved BCP/DR - Many organizations are required to have business continuity plans, disaster recovery plans and an ability to execute on them.   Even some privately help companies that do not have compliance requirements driving them in this direction have decided that it is smart to have this insurance policy to ensure survivability.   Virtualization and cloud make BCP and DR much easier and provide many more options for controlling the cost of a BCP/DR program.



Tuesday, May 17, 2011

Feel like there is not enough IT to go around?

The good news is that Spring seems to have sprung for IT.   There are signs of activity in many areas, new job openings, technology sales and most importantly - new projects.   There appears to be a drive from the business to allow IT to execute on much of the demand that has been on hold for a couple of years.   The challenge is that IT has been on a significant diet for the same length of time.

Are you seeing any of these symptoms?
  • Projects that were expected to start near the beginning of the year, are just being kicked off
  • Even though development resources levels may be addressed, new projects may be constrained by infrastructure capacity that has resourcing issues and its own ramp up schedule
  • New projects appear to be affecting each other because they may have been initiated simultaneously and draw on the same pool of resources
  • You have a feeling that quality may be affected as IT pushes the envelope on risk to achieve delivery dates
If this is affecting your organization, what can you do as an IT leader?
  • Ensure that the project portfolio maps well to the business strategy – this is not the time to be working on any projects that aren’t the highest business priority.
  • Mitigate risk by recognizing the issues of resource constraints and interdependencies and focus on 100% execution through good planning – avoid the pressure for ready, fire, aim projects.
  • Leverage variable capacity opportunities offered through consulting resources and the new secret sauce: public cloud – consider short-term ramp up of consulting for burst capacity, moving more infrastructure resources to outsourced managed services and consider a long term shift to leveraging cloud for infrastructure where your organization can live with concerns for data security.
Taking on these issues now could help you live with the result in 12 months.  

Tuesday, April 26, 2011

Amazon EC2 Outage and Cloud Strategy

Last Friday, Amazon experienced a partial outage of its cloud infrastructure.   Here the initial update and the closing updates:


Event Issue
"The problem started with a "networking event" that led to problems with how data is mirrored: We'd like to provide additional color on what were working on right now (please note that we always know more and understand issues better after we fully recover and dive deep into the post mortem). A networking event early this morning triggered a large amount of re-mirroring of EBS [Elastic Block Storage] volumes in US-EAST-1. This re-mirroring created a shortage of capacity in one of the US-EAST-1 Availability Zones, which impacted new EBS volume creation as well as the pace with which we could re-mirror and recover affected EBS volumes. Additionally, one of our internal control planes for EBS has become inundated such that it's difficult to create new EBS volumes and EBS backed instances. We are working as quickly as possible to add capacity to that one Availability Zone to speed up the re-mirroring, and working to restore the control plane issue. We're starting to see progress on these efforts, but are not there yet. We will continue to provide updates when we have them."

Closing update from Amazon:

As we posted last night, EBS (Elastic Block Store) is now operating normally for all APIs and recovered EBS volumes. The vast majority of affected volumes have now been recovered. We’re in the process of contacting a limited number of customers who have EBS volumes that have not yet recovered and will continue to work hard on restoring these remaining volumes…
We are digging deeply into the root causes of this event and will post a detailed post mortem.

One of the unfortunate realities of infrastructure and operations is that the goal will always be 100% uptime for all infrastructures but it cannot be achieved.   The SLAs for infrastructure and operations is very unlikely to be 100%.   The strategic question will always be what SLAs can be afforded, what is the impact to business agility for the target SLAs and what can be improved from a people, process and technology perspective to achieve the business goals and minimize cost. 

Because there are clear ties between performance, availability and security objectives and the success of outsource cloud infrastructure and operations, I believe that public cloud will outperform internal infrastructure over time.   This does not lessen the requirement for internal roles of architecture, end-to-end management of performance, availability and security, and vendor management.   These roles will increase in importance within organizations. 

The current Amazon issue re-emphasizes that a cloud strategy needs to include

  • Clear and continuous risk management program for IT
  • Enterprise change, incident, problem, release and configuration management process re-engineering
  • End-to-end SLA and systems management
  • Server provisioning process and technology
  • Patching process
  • Server configuration baselining and auditing
  • Repurposing of servers
  • Disaster recovery planning and testing

Tuesday, March 22, 2011

Private Cloud - why and how?


There is an explosion of change occurring in infrastructure and operations.   While it took almost a decade for virtualization to become main stream, cloud options are evolving much more rapidly.   There are two major business drivers – variable cost for consumers of IT resources and a need for increased IT agility.   All cloud options are built upon shared physical network, virtualized server and storage resources.   Cloud takes virtualization to the next level.   On top of virtualization it layers automated self-provisioning, chargeback for resource utilization, and service level agreements for cloud services that are in the service catalog.

The primary cloud discussions today center on when an enterprise will use public cloud and if it needs to implement private cloud as a stepping stone along the way or as a step-sibling for a longer period of time.   The growth of public cloud is large.   IDC estimates that the total expenditure on public cloud to be $29.5 billion by 2014.   There are some issues affecting the speed of public cloud adoption.   These are compliance concerns, data security, and cost.   As a competitive public utility, cloud cost will eventually go away as a concern.

In the meantime, Gartner believes that the most enterprises over the next couple of years will focus their attention on implementation of private cloud.   Today, there are three options for private cloud.   Enterprises can build their own private cloud (in their data centers or colocation sites), they can contract with a public cloud provider to create a physically separate private cloud for the enterprise (in cloud provider data centers or contracted colocation sites), or they can contract with a managed services provider to manage a private cloud in the enterprise data center.  

Whether an enterprise chooses to move to a public cloud or implement a private cloud, the approach to developing a strategy and implementation plan needs to follow the same methodology.   At a high level, the methodology has a four steps:
  • Define an end-state that satisfies business requirements including the financial goals, service goals and resourcing/role goals.
  • Identify the transition actions including development of services, financial changes, skill/role changes, ITIL process changes, and infrastructure changes.
  • Plan and communicate the individual transition work streams.
  • Communicate the overall program frequently and execute.
While the overall transformation creates business value and opportunity, each of the transition actions will create resistance.   Call me if you want to discuss this further.

Tuesday, March 08, 2011

Planning for Cloud Implementation

We have done a good amount of consulting on moving to public cloud (especially for companies that are not happy about the cost of their existing managed hosting vender).   On the initial discussion, one of the first questions is “what do I need to think about and how do I choose a cloud vendor?”

Moving to a public IaaS cloud vendor and to a lesser extent, a SaaS vendor is a typical data center move or implementation with a few twists and the usual issues that are easily forgotten.   While it is amazing simple and fast to build a new environment in the cloud, caring for it will take some planning, and may require changes to existing technology and processes.   While it may not be as formalized, even small organizations need to think through the issues.   Here is a checklist of items to think about:

Resiliency and Availability
Adding another node to your infrastructure network requires that you think through network configuration and redundancy, as well as server resiliency for servers that are in the new cloud environment.

Data considerations
Do you have constraints because of compliance, performance or support that affects where your data needs to be located.   A compliance requirement may force you to keep data in an internal data center and use it from the cloud.   A performance requirement may suggest a hybrid cloud with database servers in the cloud managed environment and web and application servers in the self-service environment.   Your database vendor or your performance requirements may not support a virtualized database server.

Compliance and Security
Does your IT implementation require that you have an intrusion detection or intrusion prevention system?   Is there a requirement that your infrastructure be located in a SAS-70 certified environment?   Are there requirements in your security policy that require multi-factor authentication?   Will you need to extend your vulnerability and penetration testing activities for the new site?

Identity management
How will enforce the user authentication and control policies for the new environment, e.g., when an employee or consultant leaves the organization?   Will you need to create a new AD domain and build a trust?  
Managing capacity
Monitoring performance and availability

Change, Configuration and Release Management
Will you need to add roles or workflow changes to the change management process?   What changes to you need to make to ensure that your configuration management database is current as you add and remove CIs from the new cloud environment?  Will you need to modify your release management process to push changes to the cloud?

IT service management
If there is a bump in the night, do you need to modify your incident management process to deal with workflow or contacts associated with the new environment?   Are there new services that you need to add to your service catalog to support users of the new environment?

Licensing
Will you need to extend software licensing to cover the new environment from with your vendors or will you acquire licenses through the cloud vendor or SaaS provider?

Testing
If you are moving applications or portions of applications to the new cloud environment, how will you approach functional and performance testing?

Disaster recovery
Will you build a DR site for the cloud implementation?   How will you approach data synchronization?   How does this affect your change, configuration and release management processes?

Friday, February 11, 2011

Cloud economics may surprise you

The economics of Cloud may incentivize changes in architecture.   Here is an example:

Many companies aggregate security log files from servers and network devices into a single repository to facilitate alerting on events and to support forensic investigation of security events.   For some organizations, compliance requirements like PCI indirectly make this a requirement (it would be too onerous to satisfy Requirement 11 without implementing a SIEM). There are many Security Incident and Event Management systems (SIEM) that support this.   Security log aggregation can create a large amount of network traffic to the centralized database.

Many cloud providers allow an unlimited amount of inbound network traffic, but charge for outbound network traffic.   This could create a situation for a company, considering all costs, where it is less expensive to place the SIEM and other monitoring infrastructure in the Cloud rather than inside the walls of the organization’s data center.   This may become even more obvious as the company increases the number of servers it puts in the Cloud.

Let me know if you would like help analyzing the cost of cloud for your organization.

Tuesday, February 01, 2011

Cloud Computing will change IT Organizations

At its core, cloud computing is outsourced infrastructure or application services. There will continue to be increasing adoption to allow IT to improve service levels, lower costs (if it can dial down services during periods of lower demand) and respond faster to change required by the business. In some cases, especially smaller organizations, a move to outsourced services can ensure that the number of jobs within an IT organization does not need to grow. This is good for business.

In medium and large organizations, the new ala carte menu options provided by various options in cloud computing will create new challenges for the IT organization. This will initiate a shift in job functions. There will be more outsourcing of jobs that are single-focus technical specialists (either to managed services or to services bundled with cloud offerings), but there will also be growth in need for architects, designers, development integrators, security specialists, compliance officers and IT managers of outsourcers within the IT organization. This is also good for the business because the leverage and value for the funded job position grows. IT has always created and lived with change and transformation. The cloud transformation, like all change, creates opportunities and challenges, but from the perspective of jobs, I anticipate that there will be continued net growth because as a whole, IT enables business.

Tuesday, January 04, 2011

What happens when you run out of Cloud?

Cloud Computing is generating excitement in IT because it promises to improve responsiveness to the business and reducing the cost of infrastructure during non-peak periods.  Some of the excitement comes from an expectation that that cloud capacity is not limited.   It is a fallacy.   The simple matter is that cloud infrastructure, whether you manage it yourself as a private cloud or it is outsourced in a public cloud is still based on a finite number of servers, a network that requires care when changing and a finite amount of storage at any one time.   How do cloud vendors plan on dealing with this?   The same way that your infrastructure team would deal with it -  monitoring capacity utilization and forecasting the need to expand.   There will always be limits to the power, space and possibly network connectivity to data centers, so cloud vendors will need a strategy for managing this also.   While cloud provides a great solution to the speed of provisioning and cost reduction through improvements in capacity optimization, it still has some traditional underlying issues of capacity management.   

Here are some strategies for managing your risk of running out of Cloud.   They are not mutually exclusive

Negotiate to manage risk
Negotiate contracts that either give you guaranteed capacity or early notice of capacity limits.   Guaranteeing capacity will require that you be able to forecast how much capacity you will require in the future.   This option puts your capacity management gurus in the same situation that they have always been, except they now need to forecast yet another set of resources.   Early notice of capacity limits is an alternative, if you can get this agreement from your cloud vendor.   The notice period would need to be greater than the length of time it would take to negotiate alternative Cloud resources with your current vendor or another.

Use a 2 Vendor Approach
You could start with a strategy that you plan to have a primary and a secondary vendor for Cloud.   You could choose to be your own second vendor as an option.   The assumption of this approach is that your two vendors will not run out of capacity at the same time and you can either shift resources to the secondary vendor or you are always managing application capacity across your vendors.

Build for Cloud Flexibility
Build an application/infrastructure strategy that facilitates allocation to alternative Clouds.  Being indifferent about where servers or services are to be deployed requires some critical architecture decisions.   Some key requirements would be that the application can provide sufficient performance given the range of network latencies that are possible.   Another requirement would be that the release management process and toolset supports deployment to alternative clouds.   

Like earlier shifts in technology including mini-computers, desktop computers, client-server, and the rise of the Internet, there is no turning back from Cloud Computing.   Once you start using it, you will be hooked.   Now is a great time to start planning how you will implement it before you end up in a fog (sorry).  

T3 Dynamics has been focused on leveraging Cloud Computing for its own business and the challenges of monitoring hybrid environments.    

Tuesday, November 23, 2010

Would you send your mother to a generalist for open heart surgery?


I was speaking with a colleague last week who had recently lost a valued employee that managed the company’s SANs.    He felt exposed.   This wasn’t the first time he had lost employees who were specialists.   He was thinking that the solution was to hire only generalists and give them on the job training on the many disciplines that one must manage in an IT environment – network, security, SAN, virtualization, Windows admin, Linux admin, monitoring, some DBA skills, scripting, and each of the flavors of vendor solutions of each of these disciplines.  

Cost and transiency
The core of the issue that he faced is that a specialist may be the one that will save your mother’s life if she needs a heart valve replacement, but specialists they cost more than general practitioners.  And finding the right one that wants to live or work in your location may be difficult.   What’s more, many only want to do open heart surgery and if you want them to treat your mother’s arthritis, they may look for another job.

Tactical vs. Big bang outsourcing to the rescue?
While we may all need medical specialists from time to time, we do not consider hiring one as a full time employee.   We outsource.  Since we may not need a continuous relationship with a specific medical specialist, we don’t often interview his/her partners when we are looking for one.  This is tactical outsourcing.   It makes sense for one-off needs.   Hopefully, your mother doesn’t need more than one heart valve replacement.   In California, where we have Kaiser Permanente, we can consider outsourcing all of our medical needs to a single organization.   There are options to do this in IT also.   This Big Bang outsourcing  is a big decision with big risk and significant transition requirements.

The targeted managed service option
The Internet allows diligent workers or service providers with good process and tools to work anywhere at any time for anyone.   I can recommend great managed service providers for DBA services, for JD Edwards CNC services etc.   I think of this as targeted managed services.   The benefit of this approach is that one can get a guaranteed supply of specialized resource that is shared with other customers at a cost that can be less than a full time employee.   

At T3 Dynamics, we are launching a new SaaS and managed service offering for enterprise end-to-end monitoring.   Please let me know if you want to join the no-charge beta offering.

Thursday, August 26, 2010

Paying for too many drinks?

A friend and colleague, Sasha Gilenson is president of Evolven. Evolven has a product that takes detailed snapshots of server configuration across the enterprise that can be used to investigate and diagnose change. It can be used for daily audits of infrastructure to support stability and compliance. I was speaking with him today about the tool and I was thinking that it also may be very useful to take daily snapshots of cloud consumption to audit the financial statements of cloud vendors, i.e. Amazon or Rackspace. Here is the company website http://www.evolven.com/.   Do you think about this?

Tuesday, June 01, 2010

The future of the data center professional

I believe we are moving toward outsourcing of infrastructure vis a vis IaaS clouds including outsourced private clouds.   There are larger roles available in cloud utility providers.   At the same time, internal organizations will need to strengthen their architect/outsourcing manager roles.  

The new skills include more knowledge of various virtualization technologies, knowledge of APIs of specific clouds, negotiation skills with cloud providers, x-cloud architects, heterogenous infrastructure application/systems management knowledge and skills, new disaster recovery architects, and more specific consulting skills aligned with cloud independent managed service providers.

There is a bold new world that is rapidly evolving for data center professionals.   The biggest concern might be being left behind.

Friday, November 13, 2009

The promise of virtualization and cloud computing brings challenges

I find that like most new technologies that offer great promise, the gotchas of virtualization and cloud computing rest in people issues. There are two basic promises of virtualization and IaaS cloud computing.

  1. Virtualization - We can squeeze more out of underutilized CapEx in hardware by leveraging the ability to oversubscribe the resources across multiple virtual servers
  2. Infrastructure as a Service Cloud Computing - We can automate the provisioning and de-provisioning of virtual instances to fine tune the utilization.

Before you read on, please understand that I think that virtualization and cloud computing are incredibly important to the next phase of evolution for IT enablement of business. I believe the promises are real and important.

So where are the people gotchas? As a discipline, experienced professional IT leaders recognize that it is not only admirable, but required to take on the Sisyphean challenge of implementing standards and process. Standards and process are the only consistently proven approaches to increase predictability of effort and systems, reduce defects, and the resulting long-term cost of ownership. There are three pressures that impede the implementation of standards and processes. a) It is often difficult to show immediate value to the business (a people/strategy sale/trust issue); b) Most people in the IT organization don't think of implementation of standards and process as the "fun" stuff (another people issue); c) Progress toward implementation of standards and process is driven by indivituals and progress is affected by the natural changes in career and life paths of these individuals.

So what has this to do with Virtualization and Cloud Computing? I believe that virtualization and IaaS cloud computing will tend to act as the grease for all of the counter pressures to implementing standards and processes. The ability to rapidly solve a business problem by adding another server or 10 or 800 in minutes or hours without spending capital will drive people to make changes, just this one time, without implementing or following new standards and processes. The personal challenge of tackling the increased complexity offered by virtualized servers, virtualized storage and virtualized switches will drive some engineers to relax the fight against entropy. There are probably 10x more career paths in IT that will come from virtualization and cloud computing. Who would ever thought that a specialization on how to implement virtualized iSCSI to support a VMWare ESXi cluster would have existed even 5 years ago? So people in IT will move for good personal reasons toward new opportunities and away from tuning the machine.

So it is about people. I believe we have the equivalent of a new atomic energy in IT coming from virtualization and cloud computing. This is a great thing that we will eventually harness the old way, by thinking about how we manage and motivate people. Understanding will come from experience and pain.