mary_brandel
Contributing Writer

Time to shred

News
Oct 20, 20088 mins

A funny thing happened on East Carolina University’s journey to creating a data-retention strategy: It developed a data-disposal policy.

As part of a compliance project launched about a year and a half ago, Brent Zimmer, systems specialist at the university, was working with attorneys and archivists to determine which data was most important to keep and for how long.

But it soon became clear that it was just as vital to identify which data should be thrown away.

Zimmer was aware of the importance of being able to quickly produce required information during litigation, “but the thing we never thought about was keeping data too long,” he says.

The risk is that by keeping data you wouldn’t normally be required to produce in a lawsuit, you could open the door for it to ultimately be used as evidence against you.

The university had its share of data that was overdue for purging. “We never made anyone throw away anything unless they ran out of space,” Zimmer says. The result: Some users had e-mail dating back to 1996.

East Carolina University is not unusual; many organizations hang on to more data than they need for much longer than they should, according to John Merryman, services director at GlassHouse Technologies Inc., a storage services provider in Framingham, Mass.

One reason is fear. “Companies are really sensitive because there’s a perceived underhandedness to purging data,” he says. “People might wonder, ‘Why aren’t you keeping all your records?’ “

Another is the low cost of storage. Organizations have historically preferred to simply buy more disks rather than spend time and resources sorting through what they do and don’t need. “Many people would prefer to throw technology at the problem [rather] than address it at a business level by making changes in policies and processes,” says Kevin Beaver, founder of Principle Logic LLC, an information security services firm in Acworth, Ga.

But thanks to the e-discovery risk and burgeoning data volumes — a 20% to 50% compound annual growth rate for some companies — the tide is starting to turn, Merryman says.

Good thing. The average cost companies incur for electronic data discovery ranges from US$1 million to $3 million per terabyte of data, according to GlassHouse. And although you need to pay attention to retaining data, “all indications are that you need to be keeping less,” says Merryman.

A recent report from Gartner Inc. concurs. It states that the current explosion of data is outpacing the decline in storage prices, even before the resource costs for maintaining data are taken into account.

Estimating that the average employee generates 10GB of data per year, at a cost of $5 per gigabyte to back it up, Gartner says a 5,000-worker company would face annual costs of $1.25 million for five years of storage.

And considering that many companies maintain multiple copies of data (such as test data, operational data and disaster recovery copies, not to mention backups), “there’s an explosion of data in most companies,” Merryman says.

Aside from the costs, all those records, if kept indefinitely, can become a gold mine for attorneys looking for evidence, he adds.

Policy Points

One way to address this problem is to set retention policies that reduce exposure to legal problems.

But don’t try to boil the ocean, Merryman advises. Rather than looking across the whole data landscape, create policies about what to save and how long to save it, from the application or business level down. Also, create black-and-white rules that are easy to understand and follow.

For instance, roll all data types — such as e-mail, application and file data — into 10 to 30 categories of big-picture policies rather than hundreds of granular ones. “You need broader rules, like ‘Accounting data needs to be retained six years,’ not ‘This annual report needs to be retained five years,’ ” Merryman says.

And how long to save is another consideration. According to research from Enterprise Strategy Group Inc. in Milford, Mass., the average required retention period for files, e-mails and databases is on the rise. Most companies retain such data for four to 10 years, says Brian Babineau, an analyst at ESG.

East Carolina University started with the low-hanging fruit, setting retention and purging policies for e-mail, medical records and security video footage. It archived that data on a new system that uses Symantec Corp.’s Enterprise Vault storage management software and EMC Corp.‘s Centera content-addressed storage array. E-mails from the chancellor and the dean are saved for seven years, Zimmer says; faculty and staff e-mail is purged after three years.

Moving data to the Centera devices has reduced primary storage costs by 40% to 50%, Zimmer says.

He adds that before East Carolina University used a Centera disk array, it relied on tape backups for data retention. But since backups collect data in daily snapshots, there was always the potential for data to be missing if it got to the server after the snapshot was taken. And even if the data could be found on tape, Zimmer says, the cost of restoring it would be extremely high, especially if the information that was needed was a year or more old.

“You could potentially be working on gathering that information for a week or two, just to get to a certain piece of e-mail to restore to tape for the test lab to extract,” he says.

In fact, while researching the return on investment of Enterprise Vault, Zimmer estimated that it would take 80 staff-hours to recover all the e-mail generated by one employee for one year if it had to be restored from every monthly backup tape. With the archive system, it would take just 15 to 20 minutes, and the company is guaranteed to get every piece of e-mail, he says.

The university’s security video is archived for only 30 days — a good thing, since university police collect a terabyte per day. Patient records from the medical school need to be kept for 20 years after the patient is deceased, but East Carolina University now uses EMC’s Rainfinity to take that data off primary storage and archive it to the Centera device so it’s out of the backup environment.

Beyond those policies, the job of determining rules is getting more difficult, Zimmer acknowledges. “There’s a lot of other stuff that we don’t know the retention [requirements] for, so that will be more tricky,” he says.

Gartner offers a hint. It says the key to reducing data volumes is a process called “content valuation,” which examines factors such as usage patterns, nature of content and business purpose.

The simplest way to reduce data volumes is to delete the data you don’t need. But that is much more easily said than done. In fact, outside of e-mail, most data is never dumped, Merryman says. “Most legacy applications have never purged data, and new applications are rarely designed to accommodate purging,” he says. Moreover, he adds, deleting production data is complicated.

Also, the issues associated with legal, compliance and operational risks are often ambiguous, and few organizations have a process to accommodate a web of requirements for data retention.

“If you look at legacy data outside the application world, a lot of people have no idea what it is, but they’re scared of getting rid of it,” Merryman says.

At one large bank in New York, he ran across hundreds of file extensions that no one knew about, as well as data that had been kept even though it was inaccessible by currently maintained applications or interfaces.

Another difficulty with purging is the lack of a guarantee that you’ve deleted all instances of a data set. You might think you deleted all your old e-mail, but it may be stored on tape from two years ago, so it still exists. “Some companies figure if you can’t delete it consistently, don’t delete it at all because it’s probably somewhere that no one knows about,” says ESG’s Babineau.

Given the current state of data retention in many companies, the task of purging old data could seem too daunting to even begin. But don’t let that stop you, Merryman says. Start setting purging policies now rather than trying to apply them to old data. “If you address high-risk, high-volume applications and databases, you’ll address 90% of the risk,” he says. “If you target all 700 applications in your environment, you’ll never get it done.”

And remember that business logic is with you. In a tiered storage environment, Merryman says, the business case is much stronger when you purge data rather than simply archive it on lower-cost disk. “The cost of perpetually managing and refreshing huge amounts of data that’s never been culled or purged is extremely high,” he says.

Unfortunately, he adds, most companies that develop tiering strategies figure they’ll purge sometime in the future. “But that’s the problem with purge,” he says. “It’s always ‘later,’ like cleaning out the basement.”

Still, he says, “if you invest in technology that helps you retain data, why not invest in technology that helps expire data when you don’t need it anymore?”

Brandel is a Computerworld contributing writer in Newton, Mass. Contact her at marybrandel@verizon.net.

Got something to add? Let us know in the article comments.