Big Data and cloud projects can go horribly wrong. Don’t let this happen to you.

Feature
Apr 20, 20158 mins

IT projects are not bulletproof. They are as likely to fail or encounter obstacles before coming to completion than they are to go smoothly. But when it comes to cloud and Big Data projects, the failure rate is disturbingly high.

A 2012 McKinsey study found that on average, 45% of large IT projects run over budget and 7% run late, while delivering 56% less value than anticipated. Another 17% went so bad that the very existence of the company was threatened. Big ERP projects are the poster children for IT project failure, with failure rates of at least 25% commonly cited.

+ ALSO ON NETWORK WORLD: 4 ways to beat the Big Data talent shortage +

If you think that’s bad, cloud and Big Data projects are even worse. A disconcerting report from CapGemini says that only 13% of Big Data projects have achieved full-scale production. Just 27% of respondents described their Big Data initiatives as “successful,” and only 8% described them as “very successful.”

Gartner analyst Tom Bittman, who surveyed 140 Gartner clients and reported in a blog post that only 5% had their cloud deployment projects go off without a hitch. The other 95% had one of six different problems.

Why are these companies experiencing such a high failure rate? There are a number of reasons, but they do have an overlapping root cause: companies are engaging in cloud and Big Data projects because it’s cool and trendy and not bothering to ask if they actually need it.

Tom Bittman, Gartner

Tom Bittman, Gartner

“It starts in the beginning with a good business case,” says Bittman. “Did you define the services that would benefit? That’s where most fail. Second, more than technology, private cloud is about people and process. Too often organizations say ‘I want a private cloud, what do I buy?’ The hardware is the easiest part. The hardest part is process change and people. So focus on that first. You do those two things and focus it correctly, that’s going to solve the vast majority of these issues.

Gordon Haff, vice president of cloud strategy at Red Hat, agrees. “I chalk up a lot of the failures happening in the Big Data project space to projects not having a clear goal and a clear path to that goal,” he says.

“Many organizations are undertaking Big Data projects mostly because it’s something they think they ought to be doing even if they don’t know how or why. Of course they’re not going to succeed,” he adds.

Haff says it reminds him of the hype around data warehousing and open source in past decades. “There’s this feeling that with all the data out there, we must be able to do something with it even if we don’t know the right questions to ask or the right models to apply,” he says.

The first step of a Big Data or cloud project should be to ask, “Do we really need this?” There can be a number of reasons why companies won’t: not enough data to warrant a Big Data system, reliance on older systems like ERP that don’t fit in the cloud that well, regulations that require keeping data on-premises, and so on.

“Users say they are going cloud because it’s the next thing to do when it’s not the business case. They are not asking ‘where do I need to increase my agility even more than virtualization did and what workloads am I not delivering on?'” says Bittman.

Other problems

Not identifying a business need is one reason for cloud/Big Data failures. There are other reasons as well. They include ineffective coordination between the business and technology sides of the house, dealing with scattered silos of data, ineffective coordination of analytics initiatives, the lack of a clear business case for Big Data funding, and the dependence on legacy systems to process and analyze Big Data, according to Jeff Hunter, North American leader for business information management at CapGemini.

He said many times, he encounters clients who want to use Big Data one way, when he sees it best used to bridge those silos of data. “Do they need Big Data to power the next generation of analytics to power their business paradigm? The answer might be no, but can they use it for business intelligence and decision making?” he says.

So CapGemini tells them to flip their priorities and instead of using Big Data to create big data sets, they go the other way and use it to fix issues with existing data from ERP, CRM and other traditional data sources that have been siloed and thus kept apart.

“A client might have 50 instances of customer data around the world in different formats for different apps. Some times when you solve that issue first, it makes the conversation more meaningful and attractive to the customer,” says Hunter.

Then there’s the skills gap, which has been well documented. If your team members behind a cloud or Big Data project don’t have the skills needed to deploy the project, you can bet it’s going to fail.

“Technologies in Big Data are very different from most data platforms most people are used to working with,” says Yaniv Mor, CEO and co-founder of Xplenty, which does Big Data deployments for companies as a SaaS deployment.

“SQL is not prominent in Big Data but everyone has SQL skills. Also, they are very dependent on open source technologies, something completely new to a Microsoft guy. So you need to hire new people who are expensive and difficult to find or you need to train your employees,” Mor adds.

This leads into another problem. Enterprises often see the cloud or Big Data as an extension of existing technologies. A cloud project cannot simply be an extension of your existing virtualization infrastructure.

While clouds often use virtualization, they require new approaches and new technologies. Enterprise virtualization and cloud-native infrastructure are optimized for different workloads that provide availability through software, that scale out, and that are fundamentally based on a more dynamic and loosely-coupled distributed architecture. This is different from the traditional IT infrastructure, where it’s deploy-and-don’t-touch.

Also, firms don’t change their processes or operational model when moving to the cloud, which dovetails off the above problem. Eighty to 90% of what is deployed on AWS is not net new content, says Bittman. They are horizontally scalable loads that are short-lived.

“The average life span of a VM on-premises is a few years. Back in the day of physical apps it was a decade. A VM on Amazon has a life span of days or weeks,” he says. The problem is that many firms turn on a VM on AWS and forget to turn it off when they are done. You wind up with a bill for idle cycles. He estimates that anywhere from 30% to 50% of public cloud VM usage costs is waste because people forget to turn off the VM when they are done or not using it.

What to do

So what are companies to do to reduce the potential for failure? There are a number of steps to take and they won’t cost you very much, if at all.

“Ask if you need a Big Data project in the first place,” says Mor. “With Big Data, there’s so much hype I don’t think people understand what they can get out of a Big Data project today, so they don’t know how to define the metrics. They don’t know what they want to get out of it.”

The next step is to have a leader who can create and drive a vision for the project, says Hunter. “It’s the vision that’s more important than leadership. It can come from any level. If there is a vision that clearly articulates why we want to go after Big Data and how we will go forward, as long as it permeates the company and is accepted that makes it more successful,” he said.

Third, recognize that there are two basic modes of applications and infrastructures in the typical enterprises, traditional and cloud, and trying to straddle the two without recognizing the essential differences will cause problems.

“Cloud deployments should focus on new cloud-native workloads while bridging back to existing classic IT services, workflows, and datastores and providing unified management. They should not however try to be all things to all applications,” says Haff.

And on that note, organizations need to realize there is no one way for everything IT delivers. “We call this bimodal,” says Bittman. “You have to get used to the idea of different architectures on-premises and different providers off prem. So rather than putting everything into one big architecture, they need to think managing heterogeneous architectures and sources.”

Also, start as small as you can. Don’t try to solve all your data problems when you start a Big Data project. “Just pick a business case, make sure data sources are limited from just a few data sources. Define exactly what you want to get out of this project. You need to start small because when you think about Big Data, you should not try to solve all your organizations data problems at once,” adds Mor.

Patrizio is a freelance writer. He can be reached at andypatrizio@gmail.com.

Andy Patrizio is a freelance journalist based in southern California who has covered the computer industry for 20 years and has built every x86 PC he’s ever owned, laptops not included.

Andy writes the Data Center Explorer blog for Network World. His work has appeared in a variety of publications, including Tom's Guide, Wired, Dr. Dobbs Journal, Tech Target, Business Insider, and Data Center Knowledge. Earlier in his career, he held editorial positions at IT publications like InternetNews, PC Week and InformationWeek.

Andy holds a BA in Journalism from the University of Rhode Island.

More from this author