by Mark Settle, Chief Information Officer, BMC Software, special to Network World

Peeking Under the Hood of Your Data Center

How-To
Dec 23, 201010 mins

In many ways, managing a data center is like maintaining a car. You need to periodically peek under the hood to make sure that everything is in good working order. Here are the questions to ask to gauge the health of your center.

This vendor-written tech primer has been edited by Network World to eliminate product promotion, but readers should note it will likely favor the submitter’s approach.

In many ways, managing a data center is like maintaining a car. You need to periodically peek under the hood to make sure that everything is in good working order.

What does it mean for a data center to be in top condition? You should be able to meet service expectations and commitments and maintain control by defining and managing risk. The data center should help your company get the most out of every person, asset, project, and activity. In addition, your data center should enable your IT organization to operate at the speed of the business.

Taking inventory

Is your data center in need of a tune-up? Here are some questions to help you assess its condition.

Checking under the hood of your data center begins with finding the answer to this fundamental question: Do I have an accurate asset inventory of my hardware and software?

Having information about your current equipment gives you leverage with your vendors. Imagine, for example, that you’ve purchased blade servers from multiple vendors. One of your existing vendors may be suggesting you could save money by standardizing on his blade offering. If you know the book value of blade servers purchased in the past from other vendors, you might challenge the vendor proposing standardization to assist you in writing down the outstanding value of the legacy servers.

In many companies, finding the current book value of the existing blades could tie up operations staff for days, manually checking what’s out there, getting into the fixed asset records, researching other peoples’ spreadsheets, and so on. Rather than waiting until you need the information, ask yourself these questions today:

    * What do I own?

    * How much of it do I have?

    * Where is it located?

    * What is the financial value of the technology I own?  What’s still on my balance sheet that I’m depreciating?  What am I leasing?

    * How many generations of technology have I fallen behind my vendors’ current offerings?

Not knowing what technology you own can cost money and time, and cause frustration. For example, a large company was in the process of consolidating its data centers from five to one. It had picked one data center as the consolidation location and purchased extra WAN circuits to migrate applications and data from the satellite centers.  When one of those centers was emptied a network engineer discovered there was still WAN traffic emanating from that center.

The operations team ultimately tore up the raised floor and found servers underneath that were still operating. Yet there was no record of them. The operators had no knowledge of what the servers were doing. And they didn’t know if it was safe to disconnect them or what might happen next if the servers were disconnected.

This company had a lack of information about what they really owned. A complete, up-to-date inventory of all their equipment would have prevented this situation from occurring.

Take a Fresh Look at Capacity Management

Virtualization in the data center and the move to cloud computing has brought renewed interest in capacity management. With the shift from mainframes to distributed environments, companies bought boatloads of servers, with many ending up operating at 20% capacity or less. Servers were so inexpensive they tended to purchase more and more of them, even though they might use those servers only 15% of the time.

Now the pendulum is swinging back, and we want to get more out of the boxes. The servers also take up space and increase energy expenses and labor costs. Ask yourself these questions to help determine the true cost of your investment in servers and storage equipment:

    * How many devices do I own?

    * How much heat are they generating?

    * How many kilowatts do they use, and what’s the cost of cooling my data center?

    * How much time is my operations staff spending on maintenance and wiring?

    * What is the return on my investment in that hardware?

Focus on Efficiency

Labor costs can represent one of the most significant aspects of your operational budget. Labor efficiency is a key measure of data center efficiency. To get control over labor costs, ask yourself the following:

    * What are my admin-per-device ratios for servers, storage, and network equipment?

    * How much time does my infrastructure and operations team spend on production support versus projects that drive efficiency and effectiveness?

    * What repetitive, manual tasks can be eliminated, streamlined, or automated?

    * How much busy work are we creating for ourselves by not following disciplined procedures for root cause analysis, change control, release management, and so on?

Ensure Electronic Security

The physical security of data centers is generally well managed. It’s more likely that security could be jeopardized by an electronic invasion. As hackers develop new technologies and look for different kinds of targets, it’s important to develop ways to protect your infrastructure.

To verify that you have the right security safeguards in place, ask yourself these questions:

    * Do I have a third party performing unannounced penetration testing? (The CIO and the VP infrastructure would likely know about these tests.)

    * Do they test my internal defense systems so that problems can get caught and/or reported?

    * Am I able to test various applications randomly at different times?

    * Is IT able to intercept viruses?

    * Is my company protected against a hacker trying to get into the systems?

Manage Patches Effectively

Software vendors publish patches frequently. It’s a good idea to trickle patch releases over time so as not to create contention or performance issues on your network. If you are not patching effectively, the service desk will likely receive repetitive calls reporting patch inconsistencies.  Ask yourself these questions about your internal ticket flow:

    * Am I doing more reactive work or proactive work?

    * Which incidents have disrupted user productivity?

    * Which incidents are simply user requests for increased IT capabilities and are not related to service disruptions?

Pay Attention to the Right Metrics

According to Gartner, “Through 2015, 80% of mission critical outages will be caused by people and process issues, and more than 50% of those outages are caused by change/configuration/release integration and handoff issues.” That’s why it’s so important to measure the availability of your Tier 1 systems. Begin by tracking these metrics, and have reporting information available as well. Answer the following questions:

    * Have I properly identified all of the critical configuration items (CIs) supporting my Tier 1 business systems?        

    * Have I instrumented these CIs appropriately, to ensure that I am getting early warning bulletins about performance issues and/or outages?

    * What are my resolution times for lower-severity problems on my Tier 1 business systems?

A lower-severity problem would be an outage of a component or service that wouldn’t necessarily derail the entire service. A classic example is a high-availability server cluster where the loss of a single server won’t cause the outage of the entire application. The resulting problem is more related to a degradation of service quality than of wholesale failure. The users might not see a performance issue, but if you track this information, you would know that you’re running a riskier configuration because you lost one of the four servers. Address the following questions:

    * What is the outage time around my Tier 1 applications, and is it increasing or decreasing over time?

    * How am I monitoring those outages?

    * Do I have an effective “inside-out” strategy for monitoring the critical CIs supporting each Tier 1 application and forecasting disruptions to my end users?

    * Do I have an equally effective “outside-in” strategy for tracking the experience of my end users and detecting service degradation/disruptions, even when all my CIs appear normal and healthy?

Internal monitoring is similar to the warning lights on the dashboard of an automobile — such as the oil-pressure, temperature, low-fuel, and check-engine lights — that let you know when your car needs urgent attention. The system monitors the oil pressure, temperature, fuel levels, and so on, and when the data reaches a certain threshold, the related warning light is illuminated.

Look Closely at Lower Severity Incidents

Another principal, and often overlooked, diagnostic tool is the running tally of lower-severity incidents that may impact significant numbers of people (five to ten people or more). These should be automatically reported as part of your monitoring strategy. For example, you shouldn’t have to wait for a user call to the service desk to discover that the Amsterdam switch has gone down.

The problem with many monitoring tools is that they generate a tremendous number of false signals. You don’t know which signal is valid and which one is not. Address these questions:

    * What is the monitoring strategy, particularly for lower-severity issues related to Tier 1 business systems?

    * How effectively do I use that data stream to identify problems in advance of my users?

    * Am I keeping statistics regarding my ability to resolve those problems — the true mean time to repair (MTTR)?

Put IT in the Driver’s Seat

So many IT shops are overwhelmed because they don’t have the bandwidth, skill sets, tools, or cultural model to do things proactively. And they’re sometimes rewarded for tactically solving problems in a crisis. There’s almost an expectation that they’ll lose one of their Tier 1 applications each quarter, because every quarter some new disaster happens. This quarter, the problem might be related to the ERP system. The next quarter, it might be related to the supply chain, customer support, or another area.

This cycle of disasters becomes expected, but it shouldn’t be. You should be building quality into the front end. This seems counterintuitive when people are sometimes rewarded for solving problems after they occur. But that kind of thinking is akin to a mechanic heroically rescuing a stranded driver after his car has broken down on the side of the road, rather than performing regular, preventative maintenance such as oil changes, belt replacements, tire rotation, and so on, so that the customer doesn’t become stranded.

Automation can keep your data center running smoothly, improve efficiency, and increase your data center’s overall performance. People can only scale so much, and automation helps your organization reach a level of process maturity and reliability that can’t be achieved even by the best managers.

Automation also plays a key role in keeping humans “out of the equation” when it comes to performing critical activities on an ad hoc basis, like rolling over your Exchange environment in the event of an outage. There is a wide variety of tasks in the data center that frankly you would like to have performed on an automated basis, particularly under stressful conditions that are a departure from normal operations. These are the moments in which your susceptibility to inadvertent human error is the highest.

Repeatable processes bring an immense amount of value, ultimately affecting how well you satisfy the needs of the business. By addressing the questions discussed in this article, you can help ensure that your data center is in peak condition and well equipped to manage physical and virtual environments. You will be able to expect top performance from your data center — today and into the future.