Imagine the consequences of this blog post not being saved!

Opinion
Jul 7, 20112 mins

Old Dominion University researchers getting a handle on how much of the Web is archived

We’ve all seen reports on the zettabytes of digital data now in existence, but the really imposing challenge is how to keep that stuff (or at least the most useful stuff) safe for later reading, including the oodles of data on the Web and Internet.

Old Dominion University professors and students in the Web Science and Digital Libraries Research Group have recently published a paper titled “How much of the Web is archived?” that attempts to get a handle on how much of this precious information is being socked away in some fashion.

The researchers use tools such as the Memento browser plug-in to spot old versions of web content. And to keep things manageable they sample uniform-resource identifiers (URI) from 4,000 web pages across the Open Directory Project, URL shortening service Bitly, bookmarking app Delicious and search engine caches from Google, Yahoo and Bing. What they found was that quite a bit of data is archived when you include the search engine caches (88% to 97% — excluding the Bitly content, only 35% of which is archived).

A Chronicle of Higher Education report on the research quotes an Internet Archive rep who found the results interesting but commented that Web data is such a moving target that it’s hard to really understand what it all means. However, he did say that such research is needed to strategize over how libraries and others should go about archiving online materials.

The Internet Archive made headlines — and heads turn — recently when it said it plans to store a hard copy of every published book in the world.

“LIKE” our Alpha Doggs Facebook page for more on network research