Technology is only part of the solution. For turning large amounts of data into useful information in real time, staffing and organization are key.
Imagine this: You collect vast amounts of data from more than 20 million Wi-Fi locations with more than 315 million hotspots to which tens of millions of users connect every day. You need to turn all of those mounds of data into useful information for your clients, notably the firms that own those hotspots — in real-time.
The problem isn’t just that you’re looking for a needle in a haystack. You don’t even know if you’re looking for a needle.
Those are the kinds of problems that Devicescape, which develops software for wireless networking, faces every day. They’ve been living for years with the problems of acquiring and making use of the flood of IoT data, and so have solid advice for other companies facing similar challenges — more about how to staff and organize an organization than about the hardware-and-software nuts and bolts.
One of the first issues they faced was where to host their data, and how to manage processing their extremely large data streams. Included in those data streams is the quality of service at each hotspot in real time, how people use the hotspots, large amounts of DNS information, and much more. It’s important that the data be processed in real-time, because it’s used to track the changing conditions at each hotspot.
Cedar Milazzo, Devicescape Vice President of Engineering, says that the company chose Amazon Web Services for hosting some the data, and Amazon Kinesis for real-time processing of it. The company also uses the Altiscale Data Cloud for getting useful information out of all data. The Altiscale solution is based on the Hadoop open source distributed computing and processing platform.
But when it comes to recommendations, a lot of his advice isn’t necessarily technical. Instead, it’s about how companies need to organize themselves, and how to make sure they can get staff with the right big-data skills.
“It’s important to have people with big-data experience, but also who are flexible enough to learn new technologies, because there’s always new big-data technologies coming out – the rate of growth is phenomenal,” Milazzo says. “Adaptability and a willingness to experiment is key.”
Kyle Patton, Altiscale Director of DevOps, adds that “It’s important that the people who are involved in the application scenarios are also involved in the recording and analysis of the data. Having that end-to-end perspective is important. With this kind of data analytics, you often don’t know what data you’re looking for, and it’s important that people who record and analyze the data have direct exposure to the ultimate use to which the data will be put.”
Milazzo concurs, and adds, “We go even one step beyond that. Originally we were organized along traditional lines where our development team and operations team were separate. But the more we got into it, the more we realized that you need a single DevOps team where the same people are doing the development and the operations piece of it. Combining the two functions was a very good move for us.”




