SQL Server 2008 at 1.1 petabytes, proof is in the pudding

Analysis
Nov 8, 20084 mins

Earth
Microsoft is bragging in this press release that SQL Server is being used to create some of the largest databases ever. Actually, they are taking it a little over the top saying that database software (theirs) may have what it takes to “Cure Alzheimer’s and Save the Earth.”

All this doesn’t mean that the research projects using SQL Server aren’t huge or cool, they most definitely are both. For instance, the Panoramic Survey Telescope and Rapid Response System (Pan-STARRS) is a wide-field celestial imaging facility being built at the University of Hawaii’s Institute for Astronomy. It will be used to photograph the sky — all of the sky — several times a month with the goal of trying to detect asteroids and comets, particularly ones that could pose a danger to Earth. It will rack up unspeakably huge volumes of images that will be made available to other scientists for other reasons.

How much is unspeakably huge? The facility eventually plans to have four telescopes, each outfitted with a 1.4-gigapixel resolution camera, Microsoft says. With just one telescope camera operational now, the facility generates 1.4 terabytes of image data per night. The facility is installing a storage system capable of hosting 1.1 petabytes (quadrillion bytes). And the database it uses is SQL Server.

Meanwhile, at the University of Washington near the Microsoft campus, protein researchers are trying to identify malformed proteins that may cause disseases such as mad cow, Alzheimer’s, Parkinson’s, emphysema and cystic firosis. The lab has created over 64 terabytes of data and is stockpiling bytes at a rate of 15 terabytes annually. This too, in SQL Server, with applications that mine this data via the database’s OLAP capabilities.

The point Microsoft is making is that SQL Server has outgrown its image as an SMB database. And that’s true. The part that Microsoft leaves out, however, is that the capabilities its competitors have for handling big scientific databases have also matured. And, given that discounts for software licenses for education/research can make Microsoft software downright free to use, even on monsterously large systems, universities have a lot of little greenback reasons to stretch SQL to new breadths.

Data warehouse analyst and Network World blogger Curt Monash says that when it comes to scientific data warehouses, Microsoft’s competition needs to be considered, too. He points out that Kognitio has an astronomical database too, at Cambridge University, adding 1/2 a terabyte of data per night. Oracle is used for a McGill University proteonomics database called CellMapBase. Meanwhile, Netezza also claims the be a good choice for giant collections of images. Monash’s overall take:

“Long-term, I imagine that the most suitable DBMS for these purposes will be MPP systems with strong datatype extensibility — e.g., DB2, PostgreSQL-based Greenplum, PostgreSQL-based Aster nCluster, or maybe Oracle.”

Which is not to say Microsoft shouldn’t strive for the stars (or the cells) as a proving ground for how much SQL Server has grown. But the operational — and financial — considerations of the typical enterprise may not be in close enough alignment with an scientific database for a “proof is in the puding” comparison. This may be a short-term problem. In October Microsoft acquired MPP provider DATAllegro. Still, it should help SQL Server to grow into a true MPP player, and while that might not save the planet, it could make Microsoft a bigger data warehouse star.

Also see: Curt Monash: A World of Bytes

 

Visit the Microsoft Subnet home page for more news, blogs, podcasts:

Brian Egler’s SQL Server Strategies7 keys to cleaning up Windows with Windows 717 job-hunting resources for Windows prosGlenn Weadock on Windows Server 2008Library of Windows management tools from A Better Windows Worldall Microsoft Subnet bloggers.bi-weekly Microsoft newsletter. (Click on News/Microsoft News Alert.)

Subscribe to

Sign up for the