by Jeffrey Fritz, Director Of Enterprise Network Services, UCSF

UCSF’s up-tempo upgrade

Feature
Aug 11, 200311 mins

Network director wins funding for major upgrade of San Francisco campus network. Now his team faces tough design issues and tight deadlines.

Editor’s note: First in a multi-part series. The second installment of Jeff Fritz’s account of his network upgrade has been delayed until the fall due to amazing circumstances that you’ll be able to read about when Fritz completes the RFP process.

By any measure, the enterprise network at the University of California, San Francisco, is a big one. It has 20,000 nodes, 15,000 switch ports, three Class B licenses (192,000 IP addresses), and encompasses three medical campuses and four hospitals in San Francisco, plus more than 100 remote campuses and regional clinics throughout California.


See the design plan (pop-up)

The network is heterogeneous, with devices from 3Com, Avaya, Cisco, Enterasys Networks, Foundry Networks and Nortel. The network backbone is multi-protocol (IP, IPX/SPX, AppleTalk), and there are multiple routing protocols (Open Shortest Path First [OSPF], Enhanced Interior Gateway Routing Protocol [EIGRP] and Routing Information Protocol). The core of the network is a SONET ring at the Parnassus Heights main campus, and most San Francisco sites connect to it over ATM.

The complexity of this infrastructure makes life difficult for the six-person network operations center staff. They must have an intimate knowledge of multiple switch/router operating systems, multiple protocols and multiple network monitoring applications. The way things stand today, it could take hours to detect an intrusion attack and days to react to it.

Not only that, the network infrastructure is becoming obsolete. Some devices are 11 years old – well exceeding their three- to five-year life spans. Some of the cabling is Category 3 or older.

And from a design perspective, the network topology no longer makes sense. Devices were reverse-engineered into the network over time on an as-needed basis, causing the network to have no “flow” in its topology and little rhyme or reason in its design.

Instead of a well-groomed lawn, the network looks like weeds in a garden. While the university briefly considered upgrading the network slowly over time, it soon became obvious that nothing less than a full-blown, immediate network upgrade would do.

Building the business case

Inventory of network devices as of June 2003

Total network devices

1146

By vendor

Cabletron/Enterasys
910
Cisco152
Foundry60
Other vendors24

By device type

Layer 2 switches
916
Hubs57
Routers103
Remote access devices22
ATM switches9
Wireless access points39

Network architects are reluctant salespeople, but we realized that we were going to have to put our technologist hats aside and become marketing gurus, spreading the word about the need and the vision for a UCSF Next Generation metropolitan-area network (NGMAN).

Enterprise Network Services put together the business case, which was vetted by Ken Orgill, CIO and assistant vice chancellor for IT Services. Before a Chancellors’ retreat in May, a briefing was sent out describing the NGMAN plan. We also put together a PowerPoint presentation describing the current network, the need for the NGMAN, and the estimated costs and benefits to the campus. One key point was that having a new network with far fewer routers and switches likely would reduce operating costs.

Enterprise Network Services (ENS) members also spoke with influential campus users, gained support from the campus Computer Support Coordinators and actively promoted the NGMAN to key campus technology groups such as UCSF IT Governance Network Subcommittee.

Orgill made a formal presentation at the Chancellor’s Retreat, and all the work paid off handsomely with a multimillion-dollar allocation spread over three years. Chancellor J. Michael Bishop approved our proposal in June, saying, “UCSF is a first-rate medical institution with a third-rate network. Now let’s do something about this.”

Determining the network applications

We were elated to have received quick approval for our proposal. But our joy was short-lived when we noticed the tight time frames that had been handed to us. The go-live date and migration of the first users had to occur within 14 months of project approval. Migration of all users to the new network had to be completed in 12 to 18 months after the NGMAN went live.

Fourteen months isn’t a long time to implement a top-down network redesign, especially with a four- to five-month RFP process ahead of us. We needed to determine the major requirements and key network applications  within a few weeks.

We began by examining the current network applications, talking to end users and doing a little judicious forecasting. The medical center, School of Medicine, medical researchers and library IT staffs helped identify potential applications, such as medical imaging distribution (MRI, CT scans), and remote clinician consultation/diagnosis.

IP telephony, distance learning and high-definition video distribution were also obvious applications. Secure e-mail and transfer of research, patient, doctor and student information also were pegged as key applications. On top of that, all network-based medical applications had to comply with the Health Insurance Portability and Accountability Act.

Culling this together gave us a pretty good base from which to forecast the network applications that would need to be supported over the NGMAN.

The design team

Once we had established the underlying network applications, and thus the nature of the network, it was time to put together a design team.

No one knows better how a network design should operate than the people charged with installing and operating the network. Therefore, all four ENS units (Enterprise Architecture Design, Network Operations, Enterprise Project Management and Enterprise Production Services) were included.

From the beginning, we determined that the best way to get user needs incorporated into the design and achieve user buy-in was to involve technical members of the user community. Consequently, representatives from the medical center, School of Medicine, School of Nursing, medical researchers and library IT staffs joined the design team.

The team also included volunteers from the technical staffs of three major vendors: Cisco, Foundry and SBC. It was a bit tricky to orchestrate their participation because of vendors’ conflicting interests. But the vendor technical staffs promised to play fair, and for the most part they did. We made sure that they did not participate in the RFP creation or award decision process – and their involvement had to cease before a California law limiting vendor involvement in design efforts took effect on July 1.

However, their participation was worth the effort. They called on technical resources unavailable to us, making them a valuable addition to the team.

There was one other important team member – Material Management, the UCSF procurement unit. These are the folks who know how to purchase the equipment and services needed for a project of this nature.

Altogether there were 16 people on the team. The size of the group meant that we had to be especially considerate of each other and work to keep the design moving ahead expeditiously.

The sky is the limit

The best designs come from a combination of blue-sky dreaming and slice-and-dice reality. During the blue-sky process, there are no practical limitations, cost is no issue, and the resources are infinite. No one is critical of any idea. This lets the designers be as creative as possible. Once the list is as comprehensive as it is likely to be, the slice-and-dice phase begins. Here the designers become critics, hacking away at impractical or inordinately expensive items on the list. Because everyone is allowed to be a contributor to the blue-sky process and a critic during the slice-and-dice phase, no one’s ego is on the line. If conducted properly, the result is the best possible design elements.

Before the design process could begin in full force, several decisions had to be made. The previous network used a SONET transport. There’s nothing inherently wrong with SONET, but in many ways, SONET is showing its age. It is somewhat limited in scalability, only supports one lambda (light wave) and usually requires a managed service from the provider.

The option of converting from SONET to coarse wavelength division multiplexing or even dense WDM  was extremely attractive. It would let the UCSF network staff manage its own Layer 1 transport, increasing or decreasing the bandwidth and number of lambdas (optical channels) as required. And it would provide multiple virtual fiber pathways over one pair of fiber strands, meaning far fewer fiber pulls in the future.

The existing network uses an ATM backbone. While the technology still makes sense in a service provider’s network, the complexity of ATM makes it unwieldy in an enterprise network. The simplicity of Gigabit Ethernet, combined with its ever-growing bandwidth (1G and 10G bit/sec today, and the promise of 40G bit/sec tomorrow) made it a natural. That made the decision to go with Gigabit Ethernet in the core network a no-brainer.

There was a plan underway to take the existing network from multi-protocol to single protocol. We saw no reason to deviate from this plan. The new network would be all IP.

Furthermore, Layer 3 ASIC-based switching offers the speed advantages of bridging at Layer 2 with the intelligence of routing, so we made the core network Layer 3 switched. As for the core protocol, we divorced ourselves from Cisco’s proprietary EIGRP and went with OSPF.

To manage or not to manage

We had settled on Gigabit Ethernet riding over some flavor of WDM. Now we had to refocus on Layer 1. Would it be managed or unmanaged?

UCSF doesn’t have the staff or expertise to descend into the depths of the city’s manholes to splice or pull fiber. Therefore, everyone agreed that a service provider should manage the physical fiber infrastructure. However, the design team was divided on whether the lambdas should be managed.

Some team members favored unmanaged (dark) fiber. They said managed lambdas would be limiting, because changes generally have to be submitted ahead of time to the service providers. They argued that adding capacity could require contract and service-level agreement renegotiation in addition to long lead times.

Other team members said they believed managed lambdas would make our task easier because we could leave it up to the service provider to handle the light waves. They were concerned about the additional training required if we went the unmanaged route. They noted that it wasn’t clear whether we could even obtain dark fiber at every location in San Francisco that we needed to connect to the NGMAN.

We agonized over this decision for several weeks and finally decided to offer vendors the opportunity to bid on both managed and unmanaged fiber solutions. This affords us a closer look at the pluses and minuses of each approach as part of the RFP responses.

Topology selection was nearly as controversial as the managed lambda decision. For the ultimate in reliability, some team members preferred a full-mesh core with dual homing of the Building Distribution Facilities (BDF) that serve as connection points between the campus buildings and the core network. Others preferred a protected ring configuration. They said a protected ring was nearly as redundant, was less complex and required less fiber than a full mesh. The discussions got quite heated at times.

To resolve the issue, the group was arbitrarily split into two teams. Each team of eight people was given a week to create a design.

A week later both groups put their designs side by side on white boards. Each argued the relative advantages of their configuration and pointed out weaknesses in the other team’s topology. Both teams created strong designs, and, frankly, either would have sufficed as a decent backbone design. But in listening to the arguments it quickly became clear to that the best solution was a hybrid design that included elements of both topologies. The final topology selected was a protected ring core with dual-homed connections from the BDFs.

What’s next

In the second part of this series – to be published in about three months when the RFP process is completed – we will talk about the implementation of the RFP, vendor selection, pre-installation testing and the issues we encounter as we take the design from whiteboard to procurement.

Next Generation MAN design criteria

Optical core

Redundancy
Self-healing
Multiple lambdas (DWDM)

L2, L3

IP only (single protocol) backbone
Standards-based protocols only
1 G byte Ethernet backbone (10G byte Ethernet ready)
1 G byte Ethernet to building distribution points (10G byte Ethernet ready)
Stateful awareness, re-routable backbone
 System
Low network latency (end-to-end)
Minimal jitter
Scalability
COS
Modularity (no forklift upgrades)
Robust network management including Layer 1 monitoring
Out of band management
UPS providing a minimum of 20 minutes holding time for all network devices.
24-7 technical support with four hours on-site part delivery every day of the year.
Note: Core is defined as the optical L1 infrastructure. Backbone is defined as the L2-L3 IP network.