Belgians want REP replaced with ACAP

Opinion
Nov 1, 20064 mins

* Belgians looking to replace the Robot Exclusion Protocol with the Automated Content Access Protocol

When you have a de facto Internet standard that works well why would you want to replace it? Such is the case with the World Association of Newspapers, the European Publishers Council, the International Publishers Association, and the European Newspapers Association who want to dump the now well-understood and extensively used Robot Exclusion Protocol.

REP has been around since 1994 and is obeyed by all of the major search engines as well as most of the legitimate minor ones. What the REP does is to tell the likes of search engine spiders where they can and or cannot go in the content of a Web site.

REP is actually very simple and consists of a file located in a Web site’s root that is named robots.txt which contains a specification of which user agents (i.e. spiders) are allowed to go to which subdirectories. Obviously if a spider is “badly behaved” it can just ignore the REP specification but that’s the rule on the Internet, if someone makes their use requirements clear then honorable users will abide by those restrictions.

Well it turns out the aforementioned publishers were incited to consider creating their own robots exclusion system after a publisher in Belgium, Copiepresse, which represents a number of publishers, started legal proceedings against Google over its inclusion of Belgian news sources without explicit permission.

According to an article on SearchEngineWatch: “A hearing was held in Belgium on Sept. 5, then the ruling came out last Friday, Sept. 15. Google didn’t take part in the hearings, for reasons it says it is still investigating.”

The ruling required that Google do two things:

“1. Remove French and German-language content from the publishers from Google Belgium’s Web sites or pay a fine of €1 million per day.”

“2. Publish the ruling on Google Belgium and Google News Belgium or pay a fine of €500,000 per day.”

This is a truly crazy ruling as you would think that the publishers would want to be indexed so they and their products can be found, but rationality on this case seems to be a rare commodity.

The end result was that Google completely excluded these publishers from their index and the publishers decided to reinvent the wheel and introduce a completely new protocol which they think will prevent “conflicts” with companies like Google.

What the whole issue comes down to is apparently the idea that there should be a legal framework in which search engines and other indexing services have to get permission before they scan and index a site. Therefore, the rule being denial by default rather than today’s assumption of permission by default (which is consistent with, if you don’t explicitly restrict access then your content is by default public – nope, that’s not going to work as it obviously is not going to keep the lawyers employed).

According to a recent SearchEngineWatch article the new system is to be called ACAP or Automated Content Access Protocol and will be “a system by which the owners of content published on the World Wide Web can provide permissions information (relating to access and use of their content) in a form in which it can be recognised and where necessary interpreted by a search engine ‘crawler’, so that the search engine operator (and perhaps, ultimately, any other user) is enabled systematically to comply with such a policy or licence.”

While one can sort of see the publisher’s viewpoint it can only be seen as valid in an “old school” way of thinking about where the value of content lies. The ACAP supporters are doomed to become an historical footnote.