Jitter and the Cloud

Opinion
Jan 19, 20102 mins

Latency variability is a significant issue

Over the past week a number of bloggers have been discussing performance issues in the public cloud, most notably Amazon EC2. For example, since mid-December, Cloudkick has been seeing average latency spikes up to 1000 msec versus a typical average of 50 msec. A look at Amazon EWS status history shows no service disruptions during this timeframe. What’s striking to me is Cloudkick’s latency blow out began exactly when Amazon announced EC2 spot instances. Cloudclick’s experience may be a symptom of Amazon’s success, indicating the significant latency management challenges in a highly dynamic, over-subscribed, loaded, virtualized infrastructure—all characteristics of public and private clouds. This also may be a symptom of network control and of data planes running on general purpose servers rather than purpose-built switches. Regardless, this raises a significant concern as clouds scale: jitter.

Latency variability (jitter) is a significant issue, particularly for real-time applications. Jitter is a symptom of network and CPU congestion due to oversubscription and inadequate capacity management. It’s virtualization that makes oversubscription possible, but jitter in the private cloud is low right now since enterprises are only running 44% of their workloads on virtual servers with relatively low stacking depths. Low jitter encourages enterprises to consider virtualized real-time communication servers such as virtual PBXs from Mitel and Avaya. If Cloudkick’s experience is a symptom of scaling issues, jitter will become a significant management challenge as enterprises scale private clouds. Virtualization planners must work upfront to monitor jitter and implement procedures and processes to manage it in real time.