Where does all the time go?

Feature
Aug 21, 20064 mins

One researcher hopes to give IT executives a way to answer this irksome question.

Inside this Article
Proudest accomplishments
Early results
Related Links
Carnegie Mellon’s storage guru
Video: Carnegie Mellon’s storage research center
Sophisticated SAN change-management

The New Data Center model promises that one day your IT infrastructure will be self-managing. As executive director of Carnegie Mellon University’s Parallel Data Laboratory and its state-of-the-art Data Center Observatory (DCO), Bill Courtright is working to make that promise a reality. Under his management, researchers at this Pittsburgh institution are exploring automated storage management. Their test bed, the DCO, also is a live data center serving university users. They call their work the Self-* (self star) Storage project, with the asterisk a placeholder for the words configuring, organizing, tuning, healing and, of course, managing. Courtright first must understand which management tasks consume the most time.

What are your goals for the Self-* Storage project?

Self-* Storage is the recognition that storage, especially as you get into distributed systems, is complex and difficult to understand. You don’t build the storage and then after the fact jam a bunch of storage management on top of it and think it’s going to work; you build the storage management in from Day One. So architecturally you make the system self-tuning, self-healing and so on – a part of the fabric of the design implementation.

How do the Self-* project and the DCO dovetail?

Self-* is being instrumented and is having control systems cooked in from the beginning that will really help us in the DCO research. And there’s another piece to the DCO, which is managing the applications and how they get mapped onto the machines. So the [DCO IT and environmental systems], holistically when we look at them, will give us instrumentation and control over all hardware and software components – hardware meaning computers, valves, whatever, and software meaning file systems, operating systems, applications. Then we can look through that data and begin to infer trends, find opportunities for efficiencies that we wouldn’t get in isolation looking at a file system or a server [in the IT infrastructure] or the air-conditioning system [in the physical infrastructure].

You’ve obviously got grand plans for the project and the DCO. What accomplishment would make you the most proud?

That would have to be gaining an understanding of the human administration problem, because it’s so poorly understood now. Everybody faces the problem. If we came up with some results there, even in just framing the problem, that will have a big impact. I want to understand this problem at the highest level, at the data center. So if you’re running a data center of a given size, how many people do you need to run it, and most importantly, why? What skills do they need and where does the time go? Because if you have that model – you’re going to have to hire three people with these skills, and they’re going to spend their time roughly doing 20% this and 30% that, it’s going to become more of a tractable problem. If we can start just by getting that model, that is a large result in and of itself. Then we can [attack other problems]. The human-cost factor piques our interest because everyone complains about it. It’s been a difficult problem to study. This is a unique opportunity we have here.

Do you have early results to share?

At first, we did our time sheets at the end of the week. But doing that, we realized that we could account for half our time at best. We sort of laughed because all we did was duplicate what we knew the problem was. So now we all log our time diligently, daily. I’ve got several hundred data points just after six weeks, and I’m trying to look at the data, infer from that and map out [how time is spent], but it’s not that obvious. This speaks of – even in just a short period of time – the breadth of things we have to do to run the DCO. And that will evolve over time as the room matures. Late this year we’ll have some pretty good inferences about what that model looks like, and over the next one to two years, we’ll be

Inside this Article
Proudest accomplishments
Early results
Related Links
Carnegie Mellon’s storage guru
Video: Carnegie Mellon’s storage research center
Sophisticated SAN change-management

The New Data Center model promises that one day your IT infrastructure will be self-managing. As executive director of Carnegie Mellon University’s Parallel Data Laboratory and its state-of-the-art Data Center Observatory (DCO), Bill Courtright is working to make that promise a reality. Under his management, researchers at this Pittsburgh institution are exploring automated storage management. Their test bed, the DCO, also is a live data center serving university users. They call their work the Self-* (self star) Storage project, with the asterisk a placeholder for the words configuring, organizing, tuning, healing and, of course, managing. Courtright first must understand which management tasks consume the most time.

What are your goals for the Self-* Storage project?

Self-* Storage is the recognition that storage, especially as you get into distributed systems, is complex and difficult to understand. You don’t build the storage and then after the fact jam a bunch of storage management on top of it and think it’s going to work; you build the storage management in from Day One. So architecturally you make the system self-tuning, self-healing and so on – a part of the fabric of the design implementation.

How do the Self-* project and the DCO dovetail?

Self-* is being instrumented and is having control systems cooked in from the beginning that will really help us in the DCO research. And there’s another piece to the DCO, which is managing the applications and how they get mapped onto the machines. So the [DCO IT and environmental systems], holistically when we look at them, will give us instrumentation and control over all hardware and software components – hardware meaning computers, valves, whatever, and software meaning file systems, operating systems, applications. Then we can look through that data and begin to infer trends, find opportunities for efficiencies that we wouldn’t get in isolation looking at a file system or a server [in the IT infrastructure] or the air-conditioning system [in the physical infrastructure].

You’ve obviously got grand plans for the project and the DCO. What accomplishment would make you the most proud?

That would have to be gaining an understanding of the human administration problem, because it’s so poorly understood now. Everybody faces the problem. If we came up with some results there, even in just framing the problem, that will have a big impact. I want to understand this problem at the highest level, at the data center. So if you’re running a data center of a given size, how many people do you need to run it, and most importantly, why? What skills do they need and where does the time go? Because if you have that model – you’re going to have to hire three people with these skills, and they’re going to spend their time roughly doing 20% this and 30% that, it’s going to become more of a tractable problem. If we can start just by getting that model, that is a large result in and of itself. Then we can [attack other problems]. The human-cost factor piques our interest because everyone complains about it. It’s been a difficult problem to study. This is a unique opportunity we have here.

Do you have early results to share?

At first, we did our time sheets at the end of the week. But doing that, we realized that we could account for half our time at best. We sort of laughed because all we did was duplicate what we knew the problem was. So now we all log our time diligently, daily. I’ve got several hundred data points just after six weeks, and I’m trying to look at the data, infer from that and map out [how time is spent], but it’s not that obvious. This speaks of – even in just a short period of time – the breadth of things we have to do to run the DCO. And that will evolve over time as the room matures. Late this year we’ll have some pretty good inferences about what that model looks like, and over the next one to two years, we’ll be validating that model every day.

A virtualization breeze | Next story: All things virtual >