I’ve written about network risk in this blog before, and have talked about the fact that the most common cause of network outages is simple human error. And that’s just for day-to-day operations.
Few would argue that by far the highest risk exposure comes during network changes: Take the risks associated with everyday operations and multiply them by – depending on the scope of the change – several times to several orders of magnitude. As a result, many networks are not quite – or are not nearly – what their CIO/CTOs would like them to be, because of the adversity to the disruptions caused by upgrades, expansions, consolidations, implementation of new applications, or other changes.
The approach to implementing change projects in most networks is some variation of the following:
– Carefully inventory, assess, and baseline the steady-state network.
– Create an incremental implementation plan, based on the best experience you have available.
– Factor in reasonable assumptions that unexpected things are going to happen.
– Extend the project timeframe to account for those unexpected events.
– Create detailed backout plans for when things go wrong. (And things always go wrong.)
The reason this traditional method is so fraught with risk is that you are taking your best assumptions and then executing them on your production network.
I’ve been preaching for years the extreme importance of maintaining a lab where you can test your assumptions and qualify new hardware and software before taking them live. You can pretty well judge the importance a network operator puts on risk reduction by what his lab looks like: The better the lab, the lower the risks.
But also, the better the lab, the higher the capital expense.
And even the labs of those network operators who see them as wise investments rather than expensive luxuries are unlikely to be extensive enough to simulate more than a few PoPs and perhaps the key elements of the data center. As a result, even the best of labs cannot give you a completely clear picture of how a change will impact your entire network.
That’s where offline network modeling applications come in. For a fraction of the cost of a good lab you can simulate your entire network. A precise model allows you to run multiple “what if” scenarios, refining your designs and implementation plans, and yielding a deep understanding of just what your changes will do to your network capacity, routing, security, and performance.
So should you forget about building a lab and rely on modeling instead?
No.
A network model is only as good as the data – hardware and software configurations, software versions, link loads, and application flows – that you use to build it. If you make a mistake, your model might not accurately reflect the effects of your changes.
But combine modeling with a lab and you can drastically reduce the risk of network changes. Project timelines and expenses can also be reduced.
Here’s how the two work together:
– Create an accurate baseline model of your network.
– Implement and refine your changes in the model.
– Once you have the results you want, model your lab and implement your changes again.
– Then, implement your changes in the lab. If your data and assumptions are good, your lab should give you the same results your model gave you. You can then project with reasonable confidence that the model is correctly telling you what impact your changes will have on your entire network.
Only after these phases are completed are you ready to execute your changes, and only then do you actually touch your production network.
These days many businesses simply cannot function without their networks; as a result what was an acceptable risk five years ago is no longer acceptable. Labs and modeling applications are no longer nice-to-have luxuries; they are essential tools for reducing risk and giving you the confidence to keep your network at its most beneficial.
I’ll have a bit more to say about offline modeling in my next post.




