Bill Randall, director of MIS infrastructure at restaurant chain Red Robin, dishes up the scoop on automated management.
| |||||||||||||||||||||||||
Restaurateur Red Robin is devouring new data center technology. Use of automated management, for example, has led to impressive performance and financial gains, such as a reduction of network downtime by 50% and a resulting $130,000 saved in 2004. Now the staff has discovered that security information management (SIM) can greatly cool the sting of Sarbanes-Oxley audits. Bill Randall, director of MIS infrastructure at the Denver company, shares his insights on the payback of proactive management, the realities of using today’s automation tools and more in an interview with Signature Series Executive Editor Julie Bort.
Red Robin has been hopping with cool, new technology since building your new data center. How are you managing it all?
We started out with NetIQ AppManager Suite two years ago just to help us keep a grip on everything. We have over 40 servers sitting on a rack, and I don’t want to pay an engineer for logging in and checking on them everyday. My engineers are paid to think and work on projects. So it made sense to have AppManager monitor our systems and check that the processors, temperature, memory and so on look fine. If it sees that, hey, memory utilization has been over 50% now for over eight hours, it notifies us so we can take the appropriate action.
Are you using advanced automation – where processes a human used to do by hand are now done automatically?
There are things that we let NetIQ do. For instance, if utilization gets to a certain point on our NAS box and stays there, we know that it just needs a restart. Somehow memory utilization creeps up on Windows boxes until you reboot. So it monitors for that and reboots the box at an off-hour and sends us a notification.
I don’t know that we’re ready to turn the keys entirely over to any automation program. Remember that Matthew Broderick movie – “WarGames” – where the computer was ready to declare World War III? I don’t think that would happen in our network, but we’re still careful. Automation technology lacks is the ability to test before it acts. We need to find out what the impact is going to be. If you change the way SNMP is done, for example, it may affect the backup server and several other environments. We look at automation from a standpoint of wanting [the monitoring system] to collect information, make recommendations, and let us use our human and engineering perspectives to determine whether that recommendation actually applies. Once we’ve tested it and know the right course to take, we can script and automate.
You attribute a 50% reduction in downtime, which saved some $130,000 last year, to automation. Can you give me an example?
We looked at what our downtime was a year prior vs. a year after implementing AppManager. Outages that happen with a network infrastructure are often due to silly things – no one notices the print server is turned on until it finally runs out of disk space and print services go down because it crashes the operating system. Then you’ve got to rebuild it from the [backup] image. We now take a proactive approach and we’ve alleviated most types of errors like memory or disk space [problems] that give you page fault errors. With a proactive approach, we’ll replace that memory right away, before it can corrupt the application.
A few months ago you rolled out SIM. Why?
We expanded our relationship with NetIQ and deployed Security Manager in March – it was compliance-related. Sarbanes-Oxley is a big challenge for everybody. When auditors come in for SOX, there are many things they don’t necessarily have an understanding of from an IT perspective. Logins, expiration dates and logs are what auditors can really wrap their arms around. To meet the requirements of what our auditors wanted us to do with our logs, we were going to have to spend the equivalent of one person’s time for a full two days a week just reviewing logs. Security Manager collects all that information for us.
Now we’re taking Security Manager to the next step and creating correlation rules. One warning popping up on a log doesn’t always have much meaning. If you start seeing that warning tied to other events that are happening in different parts of the network, then you need to respond to it. If we see something on our TippingPoint IPS [intrusion-prevention system], at the firewall and at the network, then we know we need to respond.
Does your SIM monitor every device that produces a log?
Yes – routers, firewalls, IPS. Then it spits out only the things that are relevant to us.
|
For the Cisco pieces, it has built-in agents for fine-tuning the monitoring. From TippingPoint, we get the raw log, and there’s a ton of stuff there, since it’s at the edge. [With the help of systems integrator Performance Enhancements Inc. (PEI) of Boulder, Colo.,] we tuned Security Manager to get us some real value. For example, NetIQ has dozens of scripts for Dell. You can certainly deploy all those and have tons of alerts. Does that necessarily have any value? Probably not. PEI showed us which ones to use for Dell, for Exchange, to get the key information and then as we want to drill down, we know other scripts exist to dig deeper.
What kinds of tasks does your SIM do for you that you couldn’t do before?
Security Manager tests for best practices related to [system configurations]. It digs into how SNMP is configured on every machine. We’d like to think that we caught everywhere there could possibly be an anonymous login allowed, but this product is out there and it’s not ignoring anything.
What’s up next for your new data center?
We’re putting some infrastructure in place right now to support hierarchical storage management – first for our Exchange environment and eventually our whole network environment. So we just brought on a new NetApp filer for our mid-tier storage. We currently run on an EMC SAN, and we have an Aberdeen AberNAS box just for raw storage – it’s cheap, effective and reliable.
Next year, we’ll be looking at server virtualization [for our production environment]. We’ve used VMWare for sometime, for our test/dev environment. We’ve been really happy with it for that, but it seems to me you can also use it to create [an architecture] where redundancy doesn’t become a big a hardware burden. We’ve been in a new building for over a year and half, and we’re already running tight on space – we don’t have two empty racks anymore. When you talk about that automation piece, virtualization starts becoming very practical. You can literally have two of everything, where one of them can sit unused until it needs to be accessed for maintenance.
| ||||||||||||||||||||||
Restaurateur Red Robin is devouring new data center technology. Use of automated management, for example, has led to impressive performance and financial gains, such as a reduction of network downtime by 50% and a resulting $130,000 saved in 2004. Now the staff has discovered that security information management (SIM) can greatly cool the sting of Sarbanes-Oxley audits. Bill Randall, director of MIS infrastructure at the Denver company, shares his insights on the payback of proactive management, the realities of using today’s automation tools and more in an interview with Signature Series Executive Editor Julie Bort.
Red Robin has been hopping with cool, new technology since building your new data center. How are you managing it all?
We started out with NetIQ AppManager Suite two years ago just to help us keep a grip on everything. We have over 40 servers sitting on a rack, and I don’t want to pay an engineer for logging in and checking on them everyday. My engineers are paid to think and work on projects. So it made sense to have AppManager monitor our systems and check that the processors, temperature, memory and so on look fine. If it sees that, hey, memory utilization has been over 50% now for over eight hours, it notifies us so we can take the appropriate action.
Are you using advanced automation – where processes a human used to do by hand are now done automatically?
There are things that we let NetIQ do. For instance, if utilization gets to a certain point on our NAS box and stays there, we know that it just needs a restart. Somehow memory utilization creeps up on Windows boxes until you reboot. So it monitors for that and reboots the box at an off-hour and sends us a notification.
I don’t know that we’re ready to turn the keys entirely over to any automation program. Remember that Matthew Broderick movie – “WarGames” – where the computer was ready to declare World War III? I don’t think that would happen in our network, but we’re still careful. Automation technology lacks is the ability to test before it acts. We need to find out what the impact is going to be. If you change the way SNMP is done, for example, it may affect the backup server and several other environments. We look at automation from a standpoint of wanting [the monitoring system] to collect information, make recommendations, and let us use our human and engineering perspectives to determine whether that recommendation actually applies. Once we’ve tested it and know the right course to take, we can script and automate.
You attribute a 50% reduction in downtime, which saved some $130,000 last year, to automation. Can you give me an example?
We looked at what our downtime was a year prior vs. a year after implementing AppManager. Outages that happen with a network infrastructure are often due to silly things – no one notices the print server is turned on until it finally runs out of disk space and print services go down because it crashes the operating system. Then you’ve got to rebuild it from the [backup] image. We now take a proactive approach and we’ve alleviated most types of errors like memory or disk space [problems] that give you page fault errors. With a proactive approach, we’ll replace that memory right away, before it can corrupt the application.
A few months ago you rolled out SIM. Why?
We expanded our relationship with NetIQ and deployed Security Manager in March – it was compliance-related. Sarbanes-Oxley is a big challenge for everybody. When auditors come in for SOX, there are many things they don’t necessarily have an understanding of from an IT perspective. Logins, expiration dates and logs are what auditors can really wrap their arms around. To meet the requirements of what our auditors wanted us to do with our logs, we were going to have to spend the equivalent of one person’s time for a full two days a week just reviewing logs. Security Manager collects all that information for us.
Now we’re taking Security Manager to the next step and creating correlation rules. One warning popping up on a log doesn’t always have much meaning. If you start seeing that warning tied to other events that are happening in different parts of the network, then you need to respond to it. If we see something on our TippingPoint IPS [intrusion-prevention system], at the firewall and at the network, then we know we need to respond.
Does your SIM monitor every device that produces a log?
Yes – routers, firewalls, IPS. Then it spits out only the things that are relevant to us.
|
For the Cisco pieces, it has built-in agents for fine-tuning the monitoring. From TippingPoint, we get the raw log, and there’s a ton of stuff there, since it’s at the edge. [With the help of systems integrator Performance Enhancements Inc. (PEI) of Boulder, Colo.,] we tuned Security Manager to get us some real value. For example, NetIQ has dozens of scripts for Dell. You can certainly deploy all those and have tons of alerts. Does that necessarily have any value? Probably not. PEI showed us which ones to use for Dell, for Exchange, to get the key information and then as we want to drill down, we know other scripts exist to dig deeper.
What kinds of tasks does your SIM do for you that you couldn’t do before?
Security Manager tests for best practices related to [system configurations]. It digs into how SNMP is configured on every machine. We’d like to think that we caught everywhere there could possibly be an anonymous login allowed, but this product is out there and it’s not ignoring anything.
What’s up next for your new data center?
We’re putting some infrastructure in place right now to support hierarchical storage management – first for our Exchange environment and eventually our whole network environment. So we just brought on a new NetApp filer for our mid-tier storage. We currently run on an EMC SAN, and we have an Aberdeen AberNAS box just for raw storage – it’s cheap, effective and reliable.
Next year, we’ll be looking at server virtualization [for our production environment]. We’ve used VMWare for sometime, for our test/dev environment. We’ve been really happy with it for that, but it seems to me you can also use it to create [an architecture] where redundancy doesn’t become a big a hardware burden. We’ve been in a new building for over a year and half, and we’re already running tight on space – we don’t have two empty racks anymore. When you talk about that automation piece, virtualization starts becoming very practical. You can literally have two of everything, where one of them can sit unused until it needs to be accessed for maintenance.
|




