* London Internet Exchange CTO Mike Hughes on Ethernet standards efforts
Earlier this week I highlighted a message from a reader, which in turn pointed to a presentation given at a NANOG meeting. Through the magic of the Internet, the newsletter got back to that presenter, and he got back to me. I’d like to share his thoughtful comments here.
The presenter was Mike Hughes, CTO for the London Internet Exchange. He writes further about the need for something beyond 10 Gigabit Ethernet, and the rest of this newsletter is in his words:
“The demand for a faster interconnect is definitely real. My organisation has already deployed 4x10GE trunked links in a metro environment, and we’re planning already for 8x10GE. Some people are already using 8x10GE. ISP networks are deploying similar parallel links inside their own networks in the U.S., using either OC-192 or OC-768 – especially between ‘big traffic’ city pairs such as between the Bay Area and L.A., or between N.Y. and Washington.
“I’ve not reached the traffic levels on my ‘Scary Doom Curve’ from the slideset, but they continue to march along – though with a respite for the summer months while the kids are out and people are on holiday.
“The good news is that since February, things are starting to move on the standards front – maybe slowly at first, but they are moving. From what I hear, it looks like 100G (and common sense) will prevail over 40G. Building the 100G MAC, I’m told, is the simple bit. Optics remain the hard bit, especially if you want to retain the ‘plug and play’ behaviour of Ethernet, which is one of the reasons that Metro Ethernet has been so successful.
“However, I’m still planning for 100GE not to ship until around 2009, so putting my operator’s hat on, what do we do between now and then?
“In the meantime, one of the things which will help operators build large enough Ethernet networks is a switch with the highest non-blocking capacity, and we’re talking something in the region of an 8-10 slot chassis with 16 wire-speed, non-oversubscribed 10GE ports on each slot.
“The biggest on sale today is Foundry’s RX-16, a 16-slot chassis with 4x10G ports per slot for a maximum of 64 10G ports in a chassis. It can forward approximately 50Gbps full-duplex to each slot.
“Now take that RX-16 and put it in a ring-type metro environment. Imagine 100G hasn’t come, you’re already filling 80G, so that you have to build your backbone from 16x10GE link-aggregated trunks (so far, trunk size increases by ‘powers of 2’). You need two of these, one westbound, the other eastbound, for redundancy. At that point, you’ve consumed 32 ports – half the chassis – just for inter-connectivity to other ‘core’ chassis in your network, and that’s before you’ve connected customers or aggregation layers.
“A bigger switch deals with this scaling problem – backplane bandwidth is cheaper than ‘interswitch’ bandwidth, as there’s no optics involved – and if you size your slots appropriately, you’ve also built yourself a ‘100G ready’ switch with a fairly good (about 5-year) shelf-life.
“Another tool in the operator’s arsenal would be support for CWDM/DWDM XFP/Xenpak modules (the pluggable 10G optics, for the uninitiated), which allows a metro or campus operator such as myself to ‘sweat’ existing dark fibre resources by providing a cost-effective means of getting multiple 10G channels onto a single fibre pair.
“Unfortunately, while such optics exist, a number of vendors ‘lock down’ their hardware to only support optic modules they have qualified and stuck one of their logos on. Some vendors don’t offer qualified CWDM/DWDM 10G optics, and if they ‘lock,’ third-party optics won’t work. The whole ‘optic locking’ issue is quite contentious in itself – see this link http://www.toad.com/gnu/sysadmin/sfp-lockin.html for a backgrounder. I won’t open that can of worms here.
“Finally, remember we’re talking Ethernet here, so to some extent, we’re stuck with imperfect redundancy protocols such as RSTP [Rapid Spanning Tree Protocol] to manage network redundancy. Imperfect in that they are wasteful of network bandwidth; you can end up with many blocked links which aren’t useable when the network is intact and stable. What if we could use some of that bandwidth, rather than leave it sitting idle, maybe we can out share the network load? (Please, don’t suggest I deploy MPLS, it’s simply not worth the hassle in this environment, and I’m already bald.)
“Enter the cavalry… Two sets, in fact, as both the IEEE and the IETF are coming to our rescue here, maybe. The question is, which is most attractive?
“The IEEE are working on something called Shortest Path Bridging (a.k.a. SPB, 802.1aq), which builds on some existing standards (RSTP and VLAN tagging) to allow multiple trees to exist in the same Ethernet network, and still avoid loops.
“The IETF-originated proposal is ‘rbridge’ (routing bridges), which plan to use a ‘shim’ (an additional header/wrapper) and ISIS to ‘route’ within a L2 domain based on the destination MAC address.
“Both have good things and bad things going for them, and the network operator such as myself is in a bit of a ‘wait and see’ position. Rbridge will require more modifications to forwarding hardware to be able to do things at line rate (you’ve got to be able to generate, add, remove and examine shims at 10G line rate for it to be useful), when compared to SPB, which is trying to make cunning reuse of existing capabilities that ship in most high-performance switching silicon. However, Rbridge may have a better loop-avoidance strategy… while SPB is weakened by remaining dependant on RSTP.”




