The blogosphere debates what’s causing iPhones to bring down parts of Duke University’s wireless LAN.
The iPhone Wi-Fi puzzle continues to baffle both the Duke University IT team and readers at Networkworld.com and other tech sites.
As reported Monday, a few iPhones on the campus-wide wireless LAN occasionally and unpredictably are triggering floods of Address Resolution Protocol (ARP) requests, as many as 18,000 per second. The flood overwhelms groups of Cisco WLAN access points, anywhere from a dozen to 30 at a time, taking them offline for 10-15 minutes or so.
So far, the Duke troubleshooters are pretty sure it is the iPhone, and not the university’s Cisco infrastructure, that is causing the disruptions. The phones are calling for the MAC address of an invalid router, and in effect, won’t shut up.
The speculation is that the flood may be triggered when the iPhone user moves from one part of the campus WLAN to another, disconnecting along the way. For some reason, when the iPhone tries to re-associate at the new location, it may be calling for an address that previously worked: the wireless router in the user’s home. It’s not a valid device on the Duke WLAN, so the iPhone inexplicably keeps jabbering for it.
But there’s no want of other explanations from tech-savvy (and some perhaps not so tech-savvy) readers on several sites and listservs. You can contribute to the networkworld.com discussion here, and cast your vote on who’s to blame in our online poll. Early Tuesday evening, Apple was leading with 54%, followed by Duke’s IT group, with 24%.
For a few Network World readers, there’s no question who’s to blame. It’s Apple. Unless it’s Cisco.
“Jeez, sounds like these guys at Apple are having their hands full when it comes to their new phone. I personally wouldn’t buy the piece of crap!” wrote ‘hpv.’
Another reader, more reasonably, observed that, “The article doesn’t specify if this is a single device malfunctioning, or if it is a defect in the iPhone itself. Reviewing the logs should help to narrow that a bit and if it can be isolated to a single device or two.”
In fact, as of Monday, July 16, Duke’s traffic analysis indicated that at least two separate iPhones had triggered ARP floods. IT staff talked with the school’s iPhone users, who didn’t seem to have done anything out of the ordinary, except that the users of the two separate iPhones had moved from one building to another. It was after this second connection was made to the WLAN that the flood occurred. That still leaves open the question of whether it’s a malfunction in one or two iPhones or a genuine iPhone bug, or no bug at all but something exposing a flaw or misconfiguration in Duke’s net.
Another reader, also anonymous, contends the problem is caused by a poor Apple implementation of a proposed IETF standard, a set of steps called Detecting Network Attachment for IPv4 (DNAv4). From the Request for Comments document, “DNAv4 optimizes the (common) case of reattachment to a network that one has been connected to previously by attempting to re-use a previous (but still valid) configuration, reducing the re-attachment time on LANs to a few milliseconds.”
This reader posted this excerpt from Section 2.1 of the document, seeming to suggest that the iPhone is “aggressively retransmitting” when it should adopt an alternative strategy (also specified in the RFC document) to minimize competition with parallel DHCP retransmissions:
“Where the reachability test does not return an answer, this is typically because the host is not attached to the network whose configuration is being tested. In such circumstances, there is typically little value in aggressively retransmitting reachability tests that do not elicit a response.
“Where DNAv4 and DHCP are tried in parallel, one strategy is to forsake reachability test retransmissions and to allow only the DHCP client to retransmit. In order to reduce competition between DNAv4 and DHCP retransmissions, a DNAv4 implementation that retransmits may utilize the retransmission strategy described in Section 4.1 of the DHCP specification [RFC2131], scheduling DNAv4 retransmission between DHCP retransmissions.”
Despite the Duke team’s consensus that this was not a Cisco problem, several readers cautioned Duke not to dismiss that possibility. “This isn’t the first time there’s a problem between a Cisco router and a specific device,” wrote ‘js.’ “For example, I know of cases where autonegotiation just cannot be allowed on 10/100/1000 Ethernet between some Cisco routers and Sun Fire series servers.” His recommendation: talk to Cisco and walk through the router configuration.
Duke has been in close touch with Cisco, but as of yesterday, nothing had turned up yet to implicate the Cisco gear.
One anonymous poster wondered if, “Maybe there is a way to apply a filter to suppress iPhone ARP [requests] down to a manageable level when an ARP storm happens.” His option: “Hell, I’d probably BAN the offending iPhones via a MAC filter from production WLAN use. After that, I’d consider sending out a campus bulletin telling users about the problem and ask them to find their MAC (with instruction as to how) for a match with the malfunctioning device(s) so that it/they can be troubleshot and or replaced (if needed).”
That process might require a lot of manual work for a campus and user population (about 13,000 students alone, according to one report) the size of Duke’s, even if the actual number of offending iPhones is small. Currently, Duke’s IT group discovered about 150 iPhone registered on its WLAN. Only a handful at most, as of Monday, seems to be involved in the ARP floods. One concern on campus is that the number of iPhones will jump dramatically as students return in late August.




