[url=http://thinkingproblemmanagement.blogspot.com/2007/11/bmx-taking-shape.html]BMX[/url] is a draft [url=http://en.wikipedia.org/wiki/Problem_solving]problem solving[/url] management methodology that I have been devising in 10 steps. 10 steps being the natural process count limit which makes sense. Problem solving methodologies exist like [url=http://thinkingproblemmanagement.blogspot.com/2007/10/kepner-tregoe-houston-we-have-problem.html]Kepner-Tregoe[/url], [url=http://thinkingproblemmanagement.blogspot.com/2007/11/root-cause-analysis-fault-tree-analysis.html]Fault tree analysis[/url], [url=http://en.wikipedia.org/wiki/Ishikawa_diagram]Ishikawa[/url], [url=http://www.google.com/search?hl=en&q=technical+observation+post&meta=]Technical observation post[/url], [url=http://thinkingproblemmanagement.blogspot.com/search?q=after+action+review]After action review[/url], [url=http://thinkingproblemmanagement.blogspot.com/2007/10/major-incident-process.html]Systems outage analysis[/url] and [url=http://en.wikipedia.org/wiki/Single_point_of_failure]Single Point of failure[/url]. However, BMX is my way and I have built in associations to make it easier to remember. The methodology assists in troubleshooting Information Technology problems in a structured manner. The method has checkpoints along to way to determine if it is worth proceeding. [img=320×305]http://lh4.google.com/ronaldxbartels/Rzl7k7efNZI/AAAAAAAACX0/ODjSdaay-YM/s144/BMX.gif[/img] The methodology is as follows: 1. Construct a [url=http://thinkingproblemmanagement.blogspot.com/2007/11/tiger-team-step-1-in-bmx.html]Tiger Team[/url]. A Tiger team is an expert problem solving team. Checkpoint: Do we have all the right skills to deal with the problem? The Tiger Team does not have hierarchy (all are equal)! 2. Use the [url=http://thinkingproblemmanagement.blogspot.com/2007/10/risk-management-as-taught-by-meerkat.html]Meerkat bolthole[/url]. Identify at a high level which of the following entities are involved in the problem: People, process, partners and/or products. It is important to document the visible and/or immediate causes. Also document the exact conditions (environment) in which this problem occured. After having listed the entities above, conduct a risk assessment on them using CRAMM Lite. Checkpoint: Is it worth moving on? Are the risks low and mitigated? Is there justification in investing further time and resources or should it be documented and parked. 3. Do you have a [url=http://thinkingproblemmanagement.blogspot.com/2007/10/ultimate-test-pilots.html]pilot’s checklist[/url] that covers the area in which this problem has occurred? Conduct an investigation using the checklist. Note any positive hits on the checklist. If there are no hits the note the five most likely hits on the checklist from lessons learned. Another alternative is to query a knowledge base. A large proportion of problems are no unique and have occurred before and have been documented. If a knowledge base does not exist then it it even suitable to type a description of the problem in Google and review the search results! Both Cisco and Microsoft have technical knowledge bases. Here is an example checklist to use for [url=http://msexchangeteam.com/archive/category/3306.aspx]Exchange[/url]. Checkpoint: Has the issue been concluded with a positive hit on the checklist or is there a requirement for further root cause analysis? 4. We now have a small sample of potential causes but need to expand the list. However, some groundwork needs to occur. We need to create a full and detailed inventory of all people, partners, processes and products. Take the high level list that has been created in step 1 and expand it to as much details as possible. As described in the methods of [url=http://thinkingproblemmanagement.blogspot.com/2007/08/roald-amundsen-cmdb-expert.html]Roald Amundsen[/url], it is important that the correct equipment and tools exist and that it is tested. As part of this process a full set of diagrams of the setup involved in the problem needs to be made available or else created. It serves no purpose to roll out a tool for the first time while dealing with a problem. Familiarity with the tool should already exist before it is used (which is the methods that Amundsen used). Although problems are not always network related, the network is a good place to further investigations. The tool sets available are network management tools like [url=http://thinkingproblemmanagement.blogspot.com/2007/10/riders-on-storm.html]Eye of the Storm[/url] or Crannog’s [url=http://www.crannog-software.com/index.php?go=Product.ShowDetail&ProductID=1]Netflow Tracker[/url]. The tools use underlying technologies like [url=http://en.wikipedia.org/wiki/IP_SLAs]IPSLA[/url] and [url=http://en.wikipedia.org/wiki/Netflow]Netflow[/url], to highlight areas of potential causes. This will either confirm a network related issue or discount connectivity problems as a cause. Note any issues that are detected that could be a secondary influence to the problem. Microsoft provides an exceelant set of Amundsen tools for Exchange. An example is ExTrA which is available [url=http://technet.microsoft.com/en-us/exchange/bb288481.aspx]here[/url]. A checklist to resolve Exchange issues highlighted by ExTrA is available [url=http://technet.microsoft.com/en-us/library/aa997574.aspx]here[/url]. Checkpoint: Has a network related issue been discovered and is it resolvable. 5. We know what is involved from the above steps, so now we need to dig deeper! This is in the spirit of [url=http://thinkingproblemmanagement.blogspot.com/2007/11/dart-fossil-excavation-step-5-in-bmx.html]Dart[/url]. The items related to the problem should be recorded and available in a CMDB. Upon interrogating this CMDB we should be able to extract a list of related incidents (especially [url=http://thinkingproblemmanagement.blogspot.com/2007/10/major-incident-process.html]major incidents[/url]), work requests, changes and associated problems. Note down a short list of about five of each of the above types. Investigate whether there is an item in the list that is directly related to the problem. Especially focus on new changes as these are often candidate causes. Try an determine if there is multiple failures related to a specific component? Is there a reliability issue related to those components? Checkpoint: Has a change or related request or incident been the cause for the problem? Is the problem related to component failure? 6. By now the problem is becoming more difficult and we need to find the patterns and break the code, as [url=http://thinkingproblemmanagement.blogspot.com/2007/11/turing-recognize-pattern-and-break-code.html]Turing[/url] would have done. Investigate the version of software and hardware being used (this information should be recorded in the CMDB). Review the release notes for the latest hardware and software versions. Investigate what software upgrades, patches and bug fixes have been applied. Often a fix for one problem causes another. Note down any deviations and references to issues that match the problem. Checkpoint: Is the problem related to a change in revisions? 7. It is time for [url=http://thinkingproblemmanagement.blogspot.com/2007/11/dambuster-step-7-in-bmx.html]Dambusters![/url] It is useful to work backwards from the problem and the way to do this is to create a timeline. Time lines are often used in FTA. Checkpoint: Is the problem time dependant? 8. Turn the fucos on production and bring in [url=http://thinkingproblemmanagement.blogspot.com/2007/11/henry-ford-step-8-in-bmx.html]Ford[/url]. Review the SOPs for the services that are involved in this problem. Perform a gap analysis on what is happening in the LIVE production environment and what has been document in the SOP. Note these differences. Checkpoint: Are these deviations in the SOP a cause of the problem? Is the SOP itself incorrect? 9. When you come to this step, the problem is a real dog, so ask for help from [url=http://thinkingproblemmanagement.blogspot.com/2007/11/pavlov-step-9-in-bmx.html]Pavlov[/url]. Investigate the people aspect of the problem both from a vendor/supplier and customer/user view. Talk to the vendors and users about the problem. List and document any of the follow issues: perception, interpretation, decision-making (knowledge-based, rule-based), and execution (skills). Checkpoint: Are we dealing with human error? 10. Time to prioritize ([url=http://thinkingproblemmanagement.blogspot.com/2007/11/pareto-step-10-in-bmx.html]Parento[/url]). We have either sufficiently determined the cause or have a list of potential causes. The causes include items from the checklists, knowledge base, network issues, change issues, component failures, release issues, SOP gaps, and human errors. Brainstorm this list and produce a list in order of most likely cause to least likely cause. Following the Pareto principle concentrate the analysis of causes to the top 20% of this list. Don’t discount the rest of the list but use it as a stimulant to the analysis process should the process come to a dead end. Checkpoint: Has the needle in the haystack been discovered? If not don’t throw in the towel as the process can be repeated at a later stage with more data and information as time is not an issue.
BMX – Problem solving methodology
Opinion
Nov 13, 20078 mins




