Campus Design, Troubleshooting for config errors, and a big whoppinng pool
endif; ?>I hope to leave a few posts this week from the Cisco Live aka Networkers show. It’s still early in the show, so no big announcements yet, and no answers yet on the questions wer’re collecting with the previous post. But I’ve seen a few things of some interest, some networking, some not, so I’ll start this week’s posts with a random thoughts kind of post.
The show is at the Mandalay Bay hotel/casino/conference center. John C (not that John C) told me when I picked up my badge that they were told the number registered was 12K so far. The hotel can handle it. 1 big impression? The pool area accommodates 4000 people. Wow.
The show started Sunday with Techtorials and Labtorials, and the main show gets started today with 3 timeslots for breakout sessions, plus more x-torials, vendor area opening, book store opening, and all that.
First networking-centric random thought: what percentage of problems in your network can be tracked to a configuration problem? My all-day session yesterday quotes some research from Gartner and from The Yankee Group regarding percentages of reasons for problems in the IT world. I’ll drag out the book and quote some numbers next time we do some discussion of troubleshooting on the exams, but it was interesting in that the thing that can be tested easily on the exams – a small config mistake – wasn’t the biggest item in the stats. I’d think that being able to troubleshoot and find a config mistake is avaluable skill, since you still have to isolate the problem, but how often is the fix a config change? I figured I’d get some opinions from ya’ll if you’re interested to take a swag.
My favorite thing at the hotel so far: a headless 25 foot tall statue of Lenin. I asked what the deal was, and the folks here tell me that the statue has been there since at least the late 90’s. It’s next to a Russian-themed Vodka bar. The boss at the time thought that it seemed to be celebrating Communism, so rather than take it down, Vlad lost his head. Weird.
The most interesting network techie thing that the show has sparked so far is the future of campus LAN design. My Sunday session was about High Availability campus design. That’s not something I look at every day in my current job, but I’ve had a ton of experience with it in years past. The rest of today I’ll ramble about it, and you can chime in if you like.
Once upon a time, the campus design was built on a distribution block, with a pair of distribution switches for redundancy, and at least 1 portchannel from each dist switch to each access layer switch. Layer 3 at the distribution layer, layer 2 to the access layer. As much as possible, you tried to constrain each VLAN to exist on on the dist switches plus only 1 access layer switch, which let you constrain STP for a VLAN to just those 3 switches.
The theory suggested in the session was that one reasonable goal was to design such that an outage discarded traffic for 200 milliseconds or less. The number comes from the idea that the human ear/brain can work around voice loss of up to 200ms with relative ease. Well, to make that kind of number, you have to do a lot, like:
- Avoid STP convergence
- Avoid Routing Protocol convergence
- Tune STP and RP to converge quickly
I’m not going to attempt to summarize the whole session here, but you can find the ideas in one of the Cisco design guides at www.cisco.com/go/srnd. However, I asked the question of one of the experts, and asked what he thought they campus LAN would look like by say 3-4 years from now. Here’s a synopsis with some of the major aspects listed. Where do you think your shop will end up in the next big campus redesign?
Layer 3 to the access layer. Cisco’s adding small amounts of layer 3 processing into some of the typical access layer switches, so that the switches could do L3 forwarding and EIGRP and OSPF stub routing. That’s in the base software (more work to do to figure out which models.) Effectively, this removes STP convergence from the equation. It means that you must make the RP converge quickly/instantly when it has to, but more importantly, you still try and avoid convergence events. EG, use Equal-Cost Multipath routing, use Multichassis EtherChannel, and of course, implement tools like NSF/SSO so that the individual boxes don’t impact traffic forwarding when rebooted. (Yep, avoiding convergence events was a popular theme.)
Virtual Switching System (VSS) at the distribution layer. Remember the age old model with 2 distribution switches? Make them a pair of 6500’s, use VSS, and the two switches act as 1 switch from both L2 and L3. The Etherchannel to the access layer can actually terminate on both 6500’s with multichassis Etherchannel. One whole 6500 could fail physically, and the convergence event is avoided. Plus if you keep the layer 2 to the access layer part of the design, the STP topo is simpler. Many of you doing this already for the campus?
Making he distribution block act like 1 switch. The idea is for the campus switch products to be developed to work somewhat like the Nexus 5000/2000. Sorry Nexus experts, I will likely butcher this, cause I’ve not looked at these boxes much yet. (I have a session on it today.) The 2000’s act like line cards for the 5000, but you locate them around the campus. Together, the 5000 and associated 2000’s act like a single switch from L2/L3 perspective. The interfaces on the distributed 2000’s appear as if on the 5000. Translated: imagine your typical distribution switch pair, and bunches of access switches, and presto-chango, it all acts like one switch. No more STP to worry about. The products to do this in the campus will take time to emerge, maybe a couple of product cycles to get there, but I found the idea intriguing.
That’s it for today. I should be able to get started on the list of questions today, now that the show floor and other stuff opens up today. Later…
W




