The greatest risk to your network is not hardware, software, or link failures. It is the human being legitimately logged in to a router or server, making authorized changes.
I said that a couple of posts ago, repeated something like it in the last post, and I say it regularly in front of audiences. When I do, I invariably get nods of agreement from most of the audience. They know; it reflects their own experience.
Now and then, I get someone who disagrees. A little questioning usually reveals one of two circumstances: Either their network is small enough and simple enough that not that many opportunities for screwups present themselves, or they operate under strict change management policies and procedures.
Assuming that you are not working with an overly vulnerable or downright boneheaded network design (and there’s plenty of that out there, but that’s a different story), there is no other measure you can take that will reduce outages in your network as much as implementing strong operational change management.
For some, change management simply means keeping track of configurations and other variables. That’s a part of it, but far from all of it. Change management is the establishment of methodologies for making changes to the network, and the enforcement of rules for adhering to the methodologies—that is, policies and procedures.
The complexity of your operational policies and procedures should be proportional to the size and complexity of your network. They also, like good security policies and procedures, must be strong enough to protect your network without becoming so inconvenient that users actively try to circumvent them.
And like good security practices change management must have executive buy-in. Without that, you have neither the funding to make it practical nor the teeth to make it enforceable.
The methodologies, or procedures, or methods of procedure, or whatever you want to call them are documented step-by-step instructions for making a specific change to your network, whether it is replacing an interface, modifying a routing policy, or upgrading the operating system. Sometimes your senior engineers write the methodologies, sometimes they are written by whoever is making the change. As more network changes are executed more methods of procedure are created; as the same changes are executed multiple times, the relevant methods of procedure are modified and improved. After a while, your library of methodologies becomes a textbook for how to operate your network efficiently and safely.
Good change management should also include standardized procedures for proposing, evaluating, and approving any changes. The change proposal should always be a standardized form explaining the change, reasons, risks, back-out procedures, and listing responsible parties. Evaluation and approval might be the responsibility of a senior engineer, or it might be the responsibility of a change committee meeting regularly and chaired by a full-time change officer. Again, complexity should be proportional to the size and complexity of your network.
And while speaking of standardization, don’t forget configuration standards. Every network is unique, with its own design philosophy, mission objectives, and operational tolerances. Your configuration standards should reflect that uniqueness, but should be consistent throughout your unique network. I have worked with amazingly few networks in which all routers in service for more than a year were configured consistently. Yet variations in configurations can contribute heavily to the potential for operator errors. Choose a configuration standard and implement a policy for insuring that it is followed consistently.
Again, strong change management is the single most important way for you to reduce outages in an otherwise well-designed and well-run network. Design a change management plan, write a business plan, and take it to your executives.




