Diverse WAN connections, continuous monitoring, sub-second response enables RAID-like revolution in enterprise WAN architecture
endif; ?>Last time, we looked at the analogy between WAN Virtualization and RAID at the business benefits level. Here, we examine the parallels from the technical point of view.
In this post, I will refer mostly to the WAN Virtualization technology of the company I founded, Talari Networks, as obviously that is the one with which I am most familiar. Ipanema Technologies‘ implementation is similar in many regards, while also including WAN Optimization capabilities as well. Mushroom Networks‘ solution shares some of these characteristics, but not others. As with any newer technology going through rapid innovation, it’s important to check with prospective vendors about the capabilities of their solution.
Wrapping hardware and intelligent software around core enabling technology
Seagate 5.25-inch hard disk technology, originally targeted at the nascent personal computer market, was not nearly as reliable as the existing mainframe/minicomputer disks, having nowhere near the MTBF (Mean Time Between Failures) or seek times. But RAID – Redundant Array of Inexpensive Disks – took advantage of the hugely better price/bit of the Seagate hard disk to revolutionize the enterprise storage market. By combining a layer of hardware and intelligent software with multiple inexpensive disks, RAID delivered a storage system with higher capacity, lower cost, competitive – and ultimately superior – access times, and greater reliability than the older-generation storage solutions.
For WAN Virtualization, the analogous enabling technology is the public Internet. By wrapping a two-ended system of appliance-based hardware running intelligent software around multiple WAN connections – most or all of which are Internet connections, but can include existing expensive MPLS connections as well – WAN Virtualization creates an enterprise WAN which is lower-cost, massively better cost/bit, higher-capacity, and more reliable than the best single vendor MPLS WAN.
Redundancy to deliver application continuity
The basic idea behind RAID – and behind WAN Virtualization as well – is that while two devices (network connections) operating in series which each have 99% reliability will deliver a system with only .99 *.99 = 98% reliability, a properly designed system with the same two devices (connections) operating in parallel will deliver 1 – (1 – .99) * (1 – .99) = 99.99% reliability.
The key phrase, of course, is proper design.
The first premise of RAID is that the loss of any single disk ensures not only that no data is lost, but also that the application – data reads and writes – continues to function normally, without meaningful performance degradation. This is essentially what RAID Level 1 delivered.
For WANs, existing traditional routed networks with appropriate link and device redundancy at each location provide network availability in the face of any hard single link failure or router failure. But this “no loss of network connectivity” is merely the equivalent of “no data loss” in the storage world. Even a no-single-point-of-failure routed network does not provide for application continuity in all cases, as a routed network can take upwards of 30 seconds at times for router convergence in the face of a given link or especially router failure. More importantly, routing does not handle the case where packet loss or excessive latency causes significant problems with application performance. Yet these “soft failures,” due to congestion on shared IP networks, especially shared WANs, occur with far greater frequency than hard link or device failures.

In fact, it is precisely because of congestion-based packet loss and jitter, which occurs most frequently at Internet peering points between ISPs, that the public Internet has earned its “works pretty well most of the time” reputation. Of course, “pretty well” isn’t good enough for most WAN managers, and “most of the time” isn’t good enough for anyone.
Handling “soft failures” – and doing it quickly!
WAN Virtualization, by leveraging multiple paths across the network between locations, ensures that no single hard or soft failure of a network link, device or peering point in the middle of the Internet will cause a loss of connectivity or application performance predictability.
WAN Virtualization does its equivalent of RAID Level 1 by continuous measurement of network path performance (loss, jitter, latency, bandwidth) combined with sub-second reaction to problems with any network path. The best WAN Virtualization implementations can move traffic off of a path experiencing high loss or excessive jitter in typically fewer than 3 round-trip times (RTTs) from problem occurrence.
For TCP applications, some WAN Virtualization solutions also deliver further application performance predictability by buffering packets from flows and retransmitting them in the face of loss, making the WAN look to the applications like a zero-loss network with occasional bouts of jitter.
For real-time application flows, WAN Virtualization technology can replicate the traffic on multiple paths, suppressing the duplicate packets at the receiving appliance, providing almost an exact equivalent mechanism to what RAID 1 does. For real-time apps like VoIP or VDI, application predictability is much more important than efficient capacity utilization, and avoiding loss and minimizing jitter by doing replication delivers exactly that. By using Internet connections, the cost of that extra bandwidth consumed is minimal. For VoIP and VDI, it’s usually trivial. Even for videoconferencing, it’s fairly small where the bandwidth capacity exists. [Note that WAN Virtualization also allows “conditional” replication of such flows, based on availability of bandwidth at the time.]
Striping = greater throughput
RAID Levels 2 – 5 (and up) allow improved data access performance by doing bit, byte and/or block level striping across multiple disks, allowing simultaneous disk seeks. With some WAN Virtualization implementations, an individual TCP flow can be striped across multiple paths/links to deliver greater aggregate throughput. The buffering and retransmission of packets, in conjunction with buffering and packet reordering at the receiving WAN Virtualization appliance, and packet forwarding logic armed with the relative latencies of each of the network paths involved, enables high TCP flow throughput even in the face of packet loss.
Automating the solution to MTBF and MTTR issues
Network availability and predictability are quite naturally the top concerns of any enterprise WAN manager. Just as RAID changed how storage systems were designed, WAN Virtualization turns the historic importance of, and so emphasis on, certain metrics on their head.
Before RAID, the MTBF of a disk subsystem was a critical factor in designing a highly reliable and available IT system. With RAID, this concern pretty much vanished. Before RAID, MTTR (Mean Time To Repair) of the storage system was a big deal as well. RAID solutions allowed much faster system MTTR (just swap out the hard disk and insert a relatively cheap replacement disk). The defective disk itself is almost never actually “repaired” any longer. And with proper redundant system-wide design and automatic synchronization and backup processes in place, if in fact the MTTR of even making that disk swap ends up being 24 or 48 hours, no one is particularly bothered.
Similarly, prior to WAN Virtualization, for private WANs the union of network availability and predictability of packet delivery, usually somewhere in the 99.95% – 99.99% range, has quite correctly been of huge importance to the enterprise network manager. And the private WAN provider’s SLA for a 4-hour MTTR, say, when a WAN problem does occur has also been extremely important.
With WAN Virtualization, however, the need for each WAN “subsystem” to be 99.9%+ predictable and reliable is greatly reduced. Given multiple diverse WAN connections – and diversity is key here – combined with sub-second switchover in the case of packet delivery problems with any of them, the need for a 4-hour MTTR guarantee goes away as well. Who cares if the MTTR is 24 or even 48 hours for a broadband connection at a branch office, say, if the network continues to work, and users notice no loss in connectivity and minimal loss in application performance and predictability? And if delivering application predictability even during those rare times when a given link is down completely is important for your application or location, by having 3 diverse connections at that location rather than 2, you will have superior application performance as well as network and application availability versus a private WAN actually delivering 99.99% uptime with an SLA committing to a 1 hour MTTR!
With WAN Virtualization, the paradigm of “monitor the network, get actionable alerts, fix it yourself” is replaced with “monitor, find and fix the problem (sub-second), tell you about it afterwards.”
RAID revolutionized storage economics, improving storage reliability and capacity while radically reducing costs. In a very similar way, WAN Virtualization is revolutionizing enterprise WAN economics and enabling the Next-generation Enterprise WAN (NEW) architecture, improving WAN reliability and application predictability and delivering enormous WAN bandwidth increases while simultaneously allowing network managers for the first time in more than a decade to not just have control over their WAN costs and their WAN providers, but to radically reduce that huge budget item as well.
A twenty-five year data networking veteran, Andy founded Talari Networks, a pioneer in WAN Virtualization technology, and served as its first CEO. Andy is the author of an upcoming book on Next-generation Enterprise WANs.




