by Bernie Lubitz

Biggest lie in the enterprise: ‘The network is down’

Opinion
Jul 5, 20075 mins

Stop blaming the network for outages -- applications and Microsoft operating systems are at fault over 99% of the time

If you are considering doubling the size of your flexible-spending medical account to buy more Ibuprofen, antacid and sleep aids, I’ve got news for you. None of these is a cure for ills caused by this message: “The network is down.” Knowing that the message is not true 99.9x% of the time is not the cure either.

This article was contributed by a reader. If you have an opinion or technology experience you would like to share, contact Online Community Editor Julie Bort (jbort@nww.com).

Enterprise network professionals take note: If you are considering doubling the size of your flexible-spending medical account to buy more Ibuprofen, antacid and sleep aids, I’ve got news for you. None of these is a cure for the curdled stomach and throbbing head caused by this dreaded message: “The network is down.” Knowing that the message is not true 99.9x% of the time is not the cure either.

As the director responsible for the enterprise network of a midsized healthcare organization, I can sympathize. Our telecommunication technology department services a converged network for data, voice, video and physical security access control. The network spans 25 buildings in a 40-mile circumference linked over private fiber. During the hurricane seasons of 2005 and 2006, our small town took three direct hits. But we experienced no loss of network services.

Our network is completing the fourth major upgrade in 15 years. Yes, we have more sites and more users, but that’s not the only reason driving us to move from 10M to 10G. We also are upgrading because inefficient applications have helped create a fivefold increase in chatty broadcast traffic, which now comprises 25% of all network traffic. While network devices have become more efficient at processing packets, applications and operating systems have not.

Now, I ask: Should network departments continue to increase bandwidth or take on network application optimization to compensate for poorly written code? I believe the best solution is for IT departments to demand that developers of operating systems, applications, hosts and end devices provide more efficient products.

With the move to Intel-based server/hosts and personal computers, somehow the network got stuck with the expectation of compensating for Microsoft’s operating systems. The MS OS is used in the mainstream while bugs continue to be found and patched. Before the fix process is complete, a new OS release will force the process to begin again. Is there any enterprise network that has not been down because of a DOS attack propagated by poorly patched PCs or Servers? A manufacturer of network equipment would not survive with a level of bug patching equal to Microsoft’s.

But when a customer cannot make a phone call, connect to a server, access a database, send a document to a printer or browse a Web site, in the user’s mind, “the network is down.” The true cause more often than not had nothing to do network failure. A VoIP phone may have lost a file and cannot authenticate, a specific server on the intranet could have been down, or the Web site that was browsed was not available. No matter. As users see it, that “the network is down.”

We, the keepers of the network, know for a fact that network availability is consistently 99.9X%, meaning that only a few thousandths of a percent of the time is the physical network at fault when a user experiences trouble. Gigabit Ethernet switching runs on redundant equipment and redundant fiber; network-management systems monitor traffic in real time; network analyzers, network-access control and intelligent-network event logs continue to enhance reliable network service.

But all these improvements have done little to improve customer perception. So maybe network support groups are doing something else wrong. Somehow we have not communicated clearly and educated our users. I submit to you that the enterprise network is the victim of poor public relations. Much of the bad PR comes from other disciplines within IT. Some bad PR originates from the public and industry news media. Yet often, we are our own worst enemy. Network guys are not known for their participation in corporate politics or their savvy PR skills.

That needs to change. Enterprise network professionals need to make an all-out effort to let people know the network is not down. We need to use easy-to-understand explanations of the services we provide, the challenges encountered and the solutions in place that nullify those challenges. We need to help our users see how well we balance financial expenditure to create an efficient, reliable network. We need to talk on a regular basis with our corporate executives and our users to instill confidence in our work. Network professionals need to talk the talk, not just walk the walk.

Lubitz is the director of telecommunications technology for Martin Memorial Health Systems, Stuart, Fla.