Power outage lead to VMware HA toggling on and off repetitively and prevented proper booting of DNS, and other crucial VMs. Redundancy is required to survive such events. My office, is not as redundant as I would want. Mainly because of lack of funds. Within my book I claim you should always have redundant networks, however on this cluster, we have just enough ports to support the networks required. This was a conscious decision based on the cost of more ports. This lack of redundancy bit us last night, specifically due to a flakey switch. It would toggle on and off causing HA to toggle on and off and lack of access to the other hosts failed using this switch. The solution was to cannibalize a gigabit switch from another rack and put in place a spare 10/100 switch to suffice for the other nodes. However, even if I had multiple switches, could we have prevented this from happening? I am not so sure about this. Could we have lost both due to basic flaws within the switches? The other part of the switch was to recreate the VLANs, and to patch direct to the main switch. We have a complex cabling setup to allow multiple networks to be used by multiple hosts. Once fixed, we were able to safely bring up VMs and get things running cobbled together as they are. While VMware HA does work, without a network, it may be impossible for it to do its magic. Remember it only works if there is no lack of isolation for any of the hosts involved. Isolation implies the link is inactive not that it can not reach a host. A flakey switch that toggles switch ports to active and inactive causes VMware HA to think it is working and then not. Redundancy is required for Virtualization Servers. Specifically network redundancy. Tied into this would be power redundancy.
Blue Gears – VMware HA and Losing a Switch
Analysis
Nov 5, 20082 mins




