CIRT management: Learning from emergencies

Opinion
Feb 1, 20072 mins

* The postmortem

This is another in an occasional series of articles looking at computer incident response team (CIRT) management.

One of the most important principles of management in general and operations management in particular is that fixing a problem has two aspects: the short term and the long term. One must be able to solve problems quickly enough to be effective; that is, the speed of solution must be appropriate to the consequential costs of delay. However, we should not figuratively wipe our hands in satisfaction and walk away from the problem resolution without thinking about why it happened, how we fixed it, and whether we can do better to avoid repeats and to improve our response.

As a matter of standard operating procedure, every technical support and CIRT must schedule time to analyze the underlying factors that led to the problem they have just resolved. This analysis will likely involve operational staff outside the CIRT; these are the people with line expertise who will be able to contribute their intimate knowledge of technical details that contributed to this security breach.

These discussions can often lead to practical recommendations for improvement of our security architecture such as topology or firewall placement, operational procedures such as monitoring standards or vulnerability patching, and technical details such as configurations or parameter settings.

Similarly, it is commonplace in discussions of disaster recovery and business continuity planning that every practice run or real-life incident should be analyzed to see where we have made errors or have achieved less than our goals in performance. Managers must ensure that these analyses are not perceived as finger-pointing exercises for a apportioning blame.

In a previous column, I have explained the concepts of “egoless work”; the postmortem analysis of an incident must be ego-free.

Managers can set the tone by responding positively to what might otherwise be perceived as criticism; “That’s a good point” and “Very good observation” are examples of positive, encouraging responses to observations such as “We were too slow in getting back to the initial caller given that she clearly stated that the entire department was off-line.” The meeting should focus on ways to improve the response given the insights resulting from detailed analysis of successes and failures during the recent incident.

In my next article on this subject, I’ll look at analyzing underlying causes of security incidents.