mary_brandel
Contributing Writer

Tier factor

How-To
Oct 11, 200410 mins

Users take baby steps on road to information life-cycle management.

Before embarking on the implementation of a new storage architecture at North Bronx Health Network, there were certain things Dan Morreale knew. The CIO at NBNH, one of six regional hospital networks in New York, knew he was always running out of storage space. He also knew that government rules such as the Health Insurance Portability and Accountability Act were changing how hospitals needed to retain data.

What he didn’t know was that his response to solving these problems would set him on the road to information life-cycle management (ILM). That is, rather than put all 100T bytes of data on an EMC Symmetrix-based storage-area network (SAN), Morreale opted to purchase two additional tiers of less-expensive and slower EMC-based storage, a Centerra and a Clariion system. That way, hospital data can move among the tiers based on its criticality, how long it needs to be retained and how quickly users need to access it. “It was a real process of discovery,” Morreale says.

ILM is tough to ignore in the storage industry these days. Every storage vendor, it seems, is touting an ILM strategy. A lot of users have jumped on the bandwagon, as well, using the term for just about anything they’ve done to gain control over their ever-growing data, from centralized backup to database archiving.

In its purest form, ILM is a combination of processes, policies and technologies that classify information according to the corporation’s policy, store it in a tiered architecture and transparently move the information among those tiers based on the information’s value, business process needs, user access needs and retention/deletion requirements.

Deployed correctly, an ILM system will ensure that information moves to the right place at the right time as it changes in value, from the moment it’s created to when it needs to be deleted.

ILM might sound a lot like hierarchical storage management (HSM) from the mainframe world. However, HSM manages only files, while an ILM strategy manages structured, semi-structured and unstructured data in a heterogeneous, networked environment. While HSM moves data based on objective measures, such as how often users access it, an ILM strategy is supposed to consider the value of the information, using parameters such as age, access frequency, date of last access, size of file, type of file and other tags the administrator attaches.

ILM: It’s a long march

For all the talk, no user has yet implemented a full-blown enterprise-wide ILM strategy, according to analysts. Despite all the data migration, storage resource management and SAN management software on the market, as well as the policy engines, document management systems and archival tools for databases, e-mail and files, the technology pieces to support full-blown ILM are not yet available.

Developing an ILM strategy is a five- to seven-year endeavor requiring the cooperation of an entire corporation and backing from the CEO, according to Randy Kerns, a senior partner at Evaluator Group, a storage analysis firm.

“Even the most sophisticated users are just at the whiteboard phase,” says Steve Duplessie, founder and senior analyst at Enterprise Strategy Group. “Others have not even gone to Staples to buy a whiteboard.”

What is ILM?

Information life-cycle management is comprised of the policies, processes, practices and tools used to align the business value of information with the most appropriate and cost-effective IT infrastructure from the time information is conceived through its final disposition. Information is aligned with business processes through management of policies and service levels associated with applications, metadata, information and data.

Source: SNIA Data Management Forum

Morreale says his system is far from complete. His group uses EMC ControlCenter to manually move data; but automated, policy-based data movement is at least a year out. Morreale also would prefer to use just one tool set that sits under one control element rather than manage the variety of SQL datastreams and other scripts the group currently uses to augment ControlCenter’s capabilities. He also would like a system that could manage data on a more-detailed level, rather than as large blocks.

So why all the hullabaloo? Well, it’s amazing what soaring data growth, the high cost of managing data and increasingly strict regulations will do. “Five years ago, we could afford to have everything stored on an EMC Symmetrix system,” Duplessie says. Since then, data has grown fivefold, faster than the cost of storing it. “Treating data like we did in 1975 is economic stupidity,” he says.

But rather than swallow the entire ILM elephant in one bite, analysts say IT groups should analyze their information storage and retention needs – with lots of interdepartmental discussion – to determine each data set’s retention, access and speed-of-retrieval needs. In the shorter term, companies should address storage pain points with technology and policies that fit within that larger strategy.

Pain points might include out-of-control e-mail growth or a slow-performing database. E-mail archiving is a hot area, according to Ray Paquet, an analyst at Gartner. “People do e-mail archiving to improve performance or conserve disk space or for regulatory purposes,” he says. These smaller bites will keep the project manageable and will drive ROI. “If a project takes more than six months, it’s three times less likely to achieve ROI,” Paquet points out.

ILM: It’s largely methodology

Doing the upfront work of developing an enterprise-wide ILM strategy is a low-tech but challenging job. Kerns says many users will need to turn to professional services for this strategy work. Although it might be tempting to skip this step – which analysts admit is painful and tedious – it is a crucial one.

Paquet advises approaching your data as three separate categories: unstructured data, such as files; semi-structured data, such as e-mail; and structured data, such as databases. “These are three distinct problems that need three distinct technologies and three distinct tools,” he says.

You can further classify your data into three hierarchical buckets, Duplessie advises: gold, silver and bronze. For each bucket, define the data’s needs in terms of reliability, disaster recovery, backup, retention, availability and performance, and then map those buckets to your storage infrastructure.

But this can’t be a one-time exercise. Data values change, so you need to create business policies that support the movement of data onto higher or lower storage tiers as needed. You also need to maintain a central repository of metadata, or “data about your data.” Policy engines, discovery tools and other systems are being developed to recommend what data should be moved based on subjective analysis and then automate data movement.

As straightforward as that sounds, this strategy work treads on thorny territory. “It’s daunting at best because the information can touch many people within the organization,” says Dave Johnson, director of IT at Grant Thornton, a global accounting firm. The Chicago company recently formed a task force to study information flow throughout its 50 U.S.-based offices. This includes paper-based and electronic documents – including what’s stored on people’s notebooks, which Johnson says poses the biggest challenge.

“Our most valuable information resides on our notebooks,” Johnson says.

And the stakes are high: “In the world of Sarbanes-Oxley, you have to retain information related to any engagement that’s significant,” Johnson says. Although he has implemented a centralized back-up strategy for portable computers, he also needs to create a post-engagement archive for completed projects, and this archive needs to be searchable in case of litigation. Furthermore, he wants automated tools that would retire the data when it no longer needs to be retained.

Although, before he looks into the necessary technology to do this, the task force needs to create a methodology for storing and archiving data based on individual projects rather than on the geographical location where they were created. “This is a people- and process-focused, not a tech-focused, solution,” Johnson says. “The technology is just a methodology of storing data once you get your arms around it.”

ILM: Implement logical migrations

It is possible to resolve your most dire storage problems using an ILM-like solution while keeping your eye on a more complete solution. At Motorola in Schaumburg, Ill., the biggest storage concern within the Personal Communications Sector (PCS) revolved around Sarbanes-Oxley and the performance of its Oracle database, says Bill Brewer, global IT configuration manager.

Brewer chose database archiving software from OuterBay Technologies to migrate customer account information that is more than 15 months old from production databases to an EMC Symmetrix SAN. The solution has increased database performance by 68% and kept required data online and easily accessible without consuming production server space. Brewer says he hopes to eventually incorporate less-expensive Clariion-based storage to further reduce storage costs while maintaining seamless end-user access. His current outsourcing contract lets him use only EMC Symmetrix systems.

“We hope to get to the point where the archive data would move onto the cheapest storage and the high-powered data would stay on the production ERP schema,” Brewer says.

Before implementing the OuterBay software, Brewer needed to define the business rules for the PCS business unit in the U.S. “Our ERP system is in multiple countries, and each has compliance regulations,” he says. For instance, the retention period in Mexico is 10 years, while in China it’s 20 years. The OuterBay system also let him define other rules, such as to not archive if a transaction is open or if its closing date hasn’t passed.

The bad news about ILM is that the upfront work can be daunting, the technology is immature and expensive, and there are plenty of internal politics to deal with. For instance, if IT can’t communicate with business leaders or even get database administrators talking with storage managers, they won’t get anywhere.

The good news is any action you take to analyze your corporate information or improve data management is a step in the right direction. “If all you do is just classify your data and understand your usage and access patterns, that’s a remarkably powerful thing,” Paquet says. “Even if you just put your new data in a tiered infrastructure, you can see significant cost savings.”

There’s plenty to do even if you’re not ready to make a technology selection. “Without retention rules and a well thought-out policy of the aging of data over time, the technology won’t do anything for you,” Morreale says.

For users like Morreale, the important thing is to get started. “I’ll start something even if I know the pieces aren’t in place – that’s how you innovate,” he says.

Timing is everything

The key to ILM is to classify data by its age and business value.
PRIMARY (ONLINE) STORAGE
Readily available
Enterprise class disk
High-performance
Mission-critical OLTP and database applications
Mirroring
Stores active data

SECONDARY

(NEARLINE) STORAGE
Accessible without operator activity
SATA disk, virtual tape
Fixed content, backup, reference data
Replication

Stores reference data,

many times read-only
PRIMARY (ONLINE) STORAGE
Removable media
Fixed content

Sarbanes-Oxley, HIPAA, government

regulations
Tape libraries, deep archive
Stores archived data that is not expected to be needed
SOURCE: EVALUATOR GROUP