Sphere, ever more refined topical searching

Opinion
Dec 20, 20064 mins

* Blog content indexing and ranking

Until the rise of blogging the only way to find “stuff” on the Internet was to use a search engine such as Google or Yahoo!

Blogging changed that. The incredibly broad range of blogging topics and the huge number of active blogs means that humans are filtering out the “good stuff” making it easier to find what you want as long as you can search through a gazillion blogs. The obvious answer is a new strategy for finding what’s important in blogs: Blog content indexing and ranking, what we’ll call generically “blog search.”

Blog search is a relatively new and rapidly evolving field that can tell us all sorts of interesting things such as what do people care about right now, what is the consensus on a topic or issue, what resources (Web sites and blogs) do people consider valuable, and what is related to what we’re looking for.

An interesting company in this area is Sphere. Launched in May after having been in stealth mode for the previous 12 months, Sphere is arguably the most in-depth blog spidering, indexing, content ranking, and searching system on the Web.

Sphere indexes tens of thousands of blogs and produces what, according to a pre-launch posting on Om Malik’s blog, you might think of as “blog rank,” analogous to Google’s PageRank.

Allow me to digress: If you want a serious analysis of the math behind Google’s PageRank algorithm see “How Google Finds Your Needle in the Web’s Haystack” published by the American Mathematical Society. Sphere presumably uses similar techniques to rank blog content.

The result of Sphere’s ranking of blog content is a focused view of the blogosphere’s take on whatever you search for. The results returned by Sphere can be filtered for the last 24 hours, the last 7 days, or the last 4 months and sorted in temporal or relevance order interpreted in English or “in any language” (i.e. ignoring English ordering).

Some topics return a huge pile of results (“Web applications” scored a total of 2,287 posts) while pop culture topics seem to be generally more “diffuse” (“Nicole Richie” scored a measly 398 posts that were more varied in focus than those of the other search – who knew tech was the more important than fame?}.

A reason for pop culture topics being more diffuse might be that they are intrinsically more personal. Malik noted that “[Sphere] has also taken a few steps to outsmart the spammers, and tends to push what seems like spam-blog way down the page. Not censuring but bringing up relevant content first. They have pronoun checker. Too many I’s could mean a personal blog, with less focused information.”

Where Sphere will make its money is in its deals with the likes of Time.com and TechCrunch. On these sites at the end of articles you’ll find a button labeled “Sphere It!” which pops up an AJAX-driven panel listing related content from the publisher and from blogs indexed by Sphere.

There’s no doubt that this is a very slick service and smart publishers will understand that the system improves site stickiness as well as increasing the visibility of the publisher’s own blogs.

Interestingly some publishers using the Sphere service appear to have created blogs specifically to mirror their content that they normally publish only on their Web site.

What we are seeing is the evolution of search relevance that started with Web sites indexed and ranked by Web search engines, progressed to vertical search services, and has now reached blogs (that in turn reference Web sites and blogs) indexed and ranked by blog search engines (the publishers I mentioned echoing their Web content in blogs bridge the hierarchy).

You have to wonder when Google and Yahoo! will start using blog indexing and ranking too …