Guaranteeing uptime at world’s largest particle physics lab

News
Jun 5, 20073 mins

CERN improves monitoring and control of technical infrastructure

As the European agency CERN was gearing up to build the world’s largest particle accelerator, officials there knew they could not afford to have problems in their technical infrastructure cause any downtime.

Back in 2003, a central messaging program that carried data among systems and applications was crashing far too often, according to Peter Sollander, head of technical infrastructure operations. CERN was preparing to build the Large Hadron Collider (LHC), the massive particle accelerator expected to begin operating less than one year from today, and knew it would need continuous availability.

“Because we’re monitoring all the services, electricity, cooling and so on, the system must work all the time,” Sollander says. “We cannot afford to lose a message, an event. Our understanding of the state of the infrastructure depends on all messages coming through.”

Sollander’s monitoring responsibility extends from systems like electricity distribution, ventilation and safety monitoring to cryogenics systems and particle accelerators. This involves facilitating the transfer of information among 115 data sources across a facility that straddles the border between Switzerland and France and houses the 27-kilometer tunnel making up the LHC. More than 1 million data updates, many of them alerts about system failures, must be processed each day, and that number will rise when the LHC begins operating.

“Because CERN is an old place, we have different generations of equipment ranging up to 40 years old to brand new stuff we’re installing now for this new accelerator,” Sollander says. “We needed to be able to streamline this information into a single format.”

By June 2005, CERN fully deployed a new solution using Progress Software’s SonicMQ, an enterprise messaging system, and an Oracle application server called OC4J.

Records of events – or problems – recorded in data sources are sent to a central server. If the central server fails or is undergoing scheduled maintenance, a backup server creates redundancy that CERN did not previously have, allowing the system to continue operating.

In the first year after deployment, the system had 99.8% availability, and most of the downtime was scheduled. Sollander expects the availability numbers to be higher once statistics are in for the second year.

CERN’s previous central messaging system, which used Talarian SmartSockets, lacked redundancy and was often unavailable, Sollander says.

With the LHC coming online, “we needed more throughput for more data coming in. The previous system was outdated, it would not support putting in a lot more data,” he says.

The previous system also had trouble finding the cause of problems when individual data sources were unavailable. Now, the Oracle middleware figures out that something is wrong when it doesn’t receive certain “watchdog messages,” according to Sollander.

“If it doesn’t get one, it will deduce that something’s wrong,” he says. “If we still have the network connection to the data source, it can send us a message to tell us why it’s not feeling so well. By analyzing this information from the watchdog piece and the data source, we can figure out what’s going on.”

By smashing particles into each other at unimaginable speeds, the Large Hadron Collider may be able to answer some of the most basic yet least understood physical processes that determine the shape of our universe.

The collisions might determine why elementary particles have mass, and whether gravity and the strong force can be unified into a single theoretical framework.