Pleiades HPC cluster helps NASA complete scientific mission
NASA’s biggest supercomputer fell out of the top 10 list, but it gets the job done for NASA’s scientists.
NASA’s biggest supercomputer seems to have gotten a little smaller. Ranked the sixth-most powerful HPC cluster in the world by the June 2010 Top 500 supercomputers list, NASA’s Pleiades fell to 11th place in the most recent rankingreleased this week.
But in reality, Pleiades is growing and becoming more useful to the giant cadre of NASA scientists who need one of the world’s largest supercomputers to plan space missions, discover new planets and perform critical research here on Earth.
Microsoft breaks petaflop barrier, loses Top 500 spot to Linux
NASA could have probably cracked the top 10 of the Top 500 ranking again if it had submitted a new score for the recently expanded system, but the agency’s priority is science, not rankings glory (although NASA’s had plenty of that, too).
Top500.org requires supercomputing researchers to submit scores under the Linpack Benchmark, which measures both the maximum achieved performance and the theoretical peak performance. While the theoretical peak can be determined with a calculator, the max speed must be found by running the Linpack software on a supercomputer.
With recent upgrades, NASA has pushed Pleiades over a petaflop — one thousand trillion calculations per second — in theoretical peak performance, and probably would reach more than 800 teraflops in actual performance. That could be good for 8th or 9th place in the Top 500, but running Linpack every time NASA upgrades the system would waste valuable time that should be devoted to scientific problems.
“In order to run Linpack, we have to take a weekend,” William Thigpen, deputy project manager for NASA’s high-end computing project, said in an interview at the SC10 supercomputing conference in New Orleans. “We don’t have time. We’re really heavily utilized, and our goal is to meet our users’ needs.”
Pleiades, based at the NASA’s Ames Research Center in the San Francisco Bay Area, is based on Linux and uses Intel’s Harpertown, Nehalem and Westmere processors across 148 racks and 84,992 cores. Pleiades also sports what Thigpen says is “the largest InfiniBand network in the world,” and connects to a separate cluster called “hyperwall” that uses AMD Opteron chips and Nvidia graphics processing units. This lets scientists access a wide range of chips, depending on the problem they’re trying to solve, in what vendors like to call “heterogeneous” computing.
Pleiades debuted in 2008, and was upgraded in both 2009 and 2010 with Nehalem and Westmere processors. Thigpen is expecting delivery of more processors in December and beyond and will likely submit a new Linpack run before June 2011, at which time Pleiades could hit measured speeds of a petaflop or at least come close.
NASA has publicly stated that it aims to hit theoretical peak performance of 10 petaflops by 2012. However, NASA had also planned to reach one petaflop in peak theoretical performance by 2009 and only just achieved that goal a few months ago.
Regardless of the timeline, Thigpen says the real goal is to give scientists the computing capacity needed to do their research. “I hate it when people can’t do their jobs, or when they have to wait,” he says.
Pleiades runs at 80% of computing capacity on average and can hit about 90% during busy times. At any given time, 300 to 400 jobs are running on the system, with the largest jobs typically requiring 25,000 to 35,000 cores, and the most commonly sized jobs at 1,000 to 2,000 cores.
However, a recent Kepler project designed to search the universe for planets required 73,000 cores, a large majority of Pleiades’ total processing power. NASA is seeing more and more scientific jobs that require 16,000 to 32,000 cores, one reason the agency is continually boosting processing power.
NASA could, theoretically, run Pleiades at 100% utilization, but it would be difficult to accommodate all the differently sized computational projects, and still make room for new ones.
NASA science that requires use of supercomputing runs the gamut, including launch simulations, planning for Mars missions, climate modeling, the study of changes on the Sun and their effect on the Earth, study of ice formation, ocean activity, earthquakes, rainfall and many other research areas.
While some supercomputers ranked highly in the Top 500 are purpose-built for a narrow range of applications, “NASA can’t do that because we have to support everyone,” Thigpen says. “We have to have a general-purpose computer.”
Creating a general-purpose supercomputer that can run a broad range of codes requires balance between “the operating system, the memory and the chip,” he says. Different applications need different amounts of memory, processing cores and I/O speeds.
“You’re always going to give up something,” Thigpen says. “Nobody has a job that’s going to fit in perfectly. You’re either going to not use all the cores, or all the memory, or all the I/O bandwidth. You’re going to be bound someplace. What you try to do is have a balanced system that’s not optimized for one specific feature.”
Linpack is not a representative measure for NASA’s supercomputer because, according to Thigpen, “systems with poor memory performance would not be detected with Linpack.” Thigpen notes that Intel’s Harpertown and Nehalem processors performed similarly in Linpack, but in terms of “real work,” the Nehalems were 2.4 times faster. That’s because Nehalem chips are more efficient in performing frequent memory operations.
Linpack, however, is useful in helping Thigpen and his team identify problems in the Pleiades cluster. While Linpack doesn’t necessarily predict performance of real scientific computing, it can determine “what type of performance you should be getting across the whole system,” Thigpen says. “It helps you find things like processors that are running slow, or memory that’s not working right, or links that aren’t working right. To us it’s really helpful as a diagnostic. And I think people like to rank things and it’s a way of ranking them that’s easy to run.”
When it comes to efficiency, Pleiades ranks a respectable No. 54 on the Green500 list, a measure of performance per watt, though Thigpen says “we’re well behind things like a Blue Gene and a Roadrunner,” both of which are IBM systems. Again, though, the Green 500 uses the Linpack performance data, and so Thigpen doesn’t consider it the be-all and end-all of efficiency ratings.
NASA has identified and turned off inefficient systems within its supercomputing infrastructure. Thigpen says NASA recently turned off parts of Columbia, once the second fastest supercomputer in the world, decommissioning 264 racks that used over 1,500 kilowatts of power. The same amount of work can now be achieved with just seven racks and 112 kilowatts, representing power savings of $1 million per year.
The entire complex that includes Pleiades costs more than $2 million a year for power, but Thigpen says his team is “constantly looking for ways to reduce that number.”
Follow Jon Brodkin on Twitter: www.twitter.com/jbrodkin




