How we tested the various enterprise anti-spam products.
We installed the enterprise anti-spam products by putting them in a production e-mail environment for about one month. To accommodate the large number of participants, we used two high-end, dual-Xeon servers loaned to us by HP, along with EMC’s VMware ESX server. We installed 19 products within 19 separate VMware images on the HP servers, using whatever operating system the vendor recommended. For the nine appliances, vendors shipped us their hardware, which we installed in our machine room. To handle these systems, we used an AMX5100 KVM system from Avocent. Finally, 10 products were service-based, so we forwarded our mail through a dual-DS3 Internet connection to those providers.
Main index: Spam in the Wild, The Sequel
Because of the scope of the test, we installed the products and attempted to debug any installation problems with the vendors over a two-month period. Etherpeek, provided to us by WildPackets, was invaluable in debugging problems with several vendors’ SMTP implementations. We also were indebted to Microsoft for providing copies of Windows 2000 and Windows 2003 needed to test several products. During this time, we sent a subset of the Opus One incoming mail stream to each product to aid in testing and tuning. All vendors who requested it had full access to their test systems for tuning, configuration and debugging purposes.
One week before testing started, we reconfigured Opus One’s corporate mail servers so that mail would come directly from the Internet to Opus One’s mail systems. Because this was an anti-spam test, we ran the message through a virus scanner and discarded messages with viruses in them.
Pre-scanning the e-mail for viruses was important. Because nearly 100% of the virus-infected e-mail is actually created by mass-mailing worms (rather than true messages from people with infected attachments), many anti-spam products are beginning to treat mass-mail worms as spam. We didn’t want the complication of dealing with differing interpretations of certain kinds of viruses as spam, so we pulled them out of the mail stream where possible.
We also found that some anti-spam products consider the double-bounces that these viruses can cause as spam, even through they are not virus-infected, and not mass-mailed. To head off disagreements on whether those messages are spam, we dropped approximately 100 obvious double-bounces from our test data and didn’t count them in the final statistics. But we did discover that some products are too aggressive in deleting all non-delivery report (NDR) mail, including notifications from our own mail servers about real messages that that had failed delivery. Those clearly erroneous false positives were counted against products that tagged them, incorrectly, as spam.
Once we had cleaned the incoming mail stream, we turned it around and simultaneously re-fed the stream to each of the participants in very close to real time. An important part of our test was that each of the products in the final test was seeing the mail as close to the time we received it (when possible). Later testing with some products where we re-fed the same stream to them days or weeks later turned in dramatically different scores, showing how important spam updates can be for products that depend on signature-based technology. Each product was connected to the Internet, and got signature and software updates as often as the vendor recommended.
Our goal with the test was to get between 10,000 and 30,000 messages over a one-month period. However, we discovered vendor irregularities in the middle of the test, and stopped the mail stream after two weeks, and approximately 11,000 messages.
We then went through every single message, classifying it as spam, not spam or unknown. We defined 8,027 messages as spam, for which there was no conceivable business or personal relationship between the sender and receiver, and which was obviously bulk in nature. In the not spam category, 2,386 messages that may or may not have been solicited, but which had a clear business or personal relationship between sender and receiver, or was obviously a one-to-one message, even if unsolicited and unwanted. All mailing lists that had legitimate subscriptions were considered not spam. We didn’t make any mailing list changes during the term of the test.
In the unknown category were several messages that either were the result of virus double-bounces, messages that we couldn’t put into one category or the other definitively, and some messages that were so malformed that we didn’t know whether they were spam, viruses or just software acting up. We also took messages with duplicate message IDs and deleted them from the data set. In theory, a duplicate message ID is impossible, but spammers don’t follow the RFCs, and we found more than 100 of those.
We then asked CentosPrime database guru Chris Janton to build us a database application to hold all the results of the tests. Janton loaded the classifications of all the vendors as well as our definitive results, and created a GUI to let us massage and report on the data. We spent some time sanity-checking our data by looking at messages that many vendors disagreed with our classification, which identified about a half-dozen errors in our first classification pass.
Janton then ran reports giving us the false positives and false negatives for each vendor. We walked through every single false positive for vendors who had 100 or fewer errors to verify our statistics and see if any re-evaluations were needed. After a final few adjustments, Janton ran us the final reports we used to turn into false-positive rates and sensitivity levels. In a perfect world, the false-positive rate of any product would be zero: It measures the level at which one of these products marks non-spam messages as spam. Similarly, the ideal sensitivity level of each product would be 100%: All spam properly marked as spam and diverted away from the end user. Since a corporate mail system would likely be much more sensitive to false positives, we asked each vendor for advice in tuning their product to reach a false-positive rate of less than 1% (see “What is a false positive?”).
Four vendors (Process Software, Ciphertrust, MailFrontier and Privacy Networks) had irregularities in their statistics, and results had to be re-evaluated. With Process Software and MailFrontier, the problem was caused on our end, so we worked with the vendors to re-run messages. Ciphertrust’s statistics were calculated on a subset of the total message stream, about 30% of the final total. In the case of Privacy Networks, we could not re-test their system fairly, so we had to drop them from the test.
Once we knew how well each product caught spam, we tried to figure out how fast they were. We turned to Spirent Communications’ WebAvalanche and WebReflector product lines. These test products can generate thousands of SMTP transactions per second. We also used an Apcon Physical Layer Switch to help manage the process of connecting, disconnecting and re-connecting the various anti-spam products to our test network and the Internet. We analyzed the more than 35,000 messages we collected during the test and pre-test period to determine which message sizes would be most appropriate. We broke down the data into 10-percentile chunks of 2.5K bytes (including all headers), 3K, 3.5K, 4K, 5K, 6K, 7K, 8K, 9K, and 10K. For example, 10% of messages are 2.5K bytes or smaller, the next 10% are 3K bytes or smaller, etc. As part of the analysis, we determined that more than 95% of the messages sent were sent in their own SMTP session. We generated non-spam messages for each of these buckets and loaded them into the WebAvalanche system.
We then created a test profile designed to reproduce the flow of message sizes that we saw in the real world. The WebReflector was programmed to accept messages without delay or loss, and simply log statistics. We gradually ramped the number of simultaneous SMTP connections to each product up to 100 simultaneous connections, and ran for a total of 4 minutes. Because the WebAvalanche and WebReflector had capacity far beyond what any of the products we tested could handle, we let the products themselves determine how much mail we would send. For each SMTP connection, the WebAvalanche would send one message, then close the connection. Then it would open a new connection to replace that, and send another message, all as fast as the receiving system could take them. In our tests of WebAvalanche talking directly to WebReflector, this turned out to be a rate in excess of 2,000 messages per second.
Because of security problems encountered in the spam-catch portion, we did not let vendors tune them for the performance portion. However, we sent out our detailed test methodology, and invited vendors to provide any detailed tuning instructions needed to optimize performance, which we then applied. For the VMware systems, we shut down all VMware images, except for the implementation under test and gave the IUT 2G bytes of memory to play with. Where possible, we disabled virus scanning and any policy-based rules to try to isolate the test to only message passing and spam scanning.
We ran each test at least three times to eliminate effects of any DNS or other lookups that devices might be doing. Because none of the messages were spam, we also hoped to eliminate the effects of any messaging quarantining that might occur.
Because of the numbers of products tested, we did not performance-test the service-based, anti-spam vendors this time around. We also found several products that weren’t compatible with the Spirent test gear, and were unable to report numbers for them. Finally, one of the appliance vendors, MailByDesign, sent us a low-performance system for the spam-catching part, and planned to send a more typical enterprise appliance for performance testing. Later, MailByDesign decided not to provide us with equipment for the performance part.
Each test generated a raft of statistics, which we boiled down to two simple numbers: the rate of message acceptance, and the rate of message delivery.
Although we wanted to see how every product submitted to our test would work when it came to identifying spam, we didn’t want to evaluate products that obviously weren’t making top grades at their main job: catching spam. Because there really are differences between products when it comes to false positives and spam-catch rates, we took that out of the equation by focusing on the top products that achieved a spam-catch rate of more than 90% and a false-positive rate of less than 1%. This let us focus on other differences between products that would sway a buying decision.




