by Jeff Vance

5 Big Data projects that could change your life

News
Aug 4, 201412 mins

Real-world Big Data projects that are already paying rewards.

Most over-hyped technology trends wear out their welcome pretty quickly, which should make skeptics among us wary about Big Data. However, while Big Data is being touted as the latest trend that will change the world, the skeptics aren’t as, well, skeptical as they were about cloud and social.

That’s probably because Big Data is generating real-world wins for the companies embracing it. Already, Big Data analytics is starting to fundamentally change such disparate disciplines as pharmaceutical research, sales and marketing, and product development.

Many use cases, such as smart cities and driverless cars, even get us excited about a Jetsons-like existence where the world around us seems to anticipate our needs. Those scenarios may be the future of Big Data, but they’re not the “now” of Big Data.

“There’s a big difference between what is technologically feasible and what is practical,” says Don DeLoach, CEO of Infobright, a data analytics company. “Look at two trends driving Big Data: Internet of Things and Machine to Machine communications. Both have been around for a long time, but the increasing sophistication of sensors and a corresponding drop in price, along with the proliferation of various wireless communications options, means that what was once technologically possible in theory is now becoming practical.

Some of our most ambitious Big Data dreams just haven’t entered into the realm of the practical yet. The technology is there for a self-driving car, but the infrastructure is not. Yet, the now is still pretty compelling.

“If you want to keep track of what’s really happening with Big Data, just follow the money,” DeLoach said. “Where ROI is the most obvious, that’s where people will invest.” And invest they have.

The ROI for Big Data in health care, vehicle telematics, and online marketing is already clear. That doesn’t mean we won’t eventually see driverless cars and super-smart cities. It just means they’re not yet practical enough to attract large-scale investments.

Here are five Big Data projects that are straddling the line between the practical and the pie-in-the-sky possible, and these projects, or ones like them, could very well change your life:

The Human Genome Project revolutionizes medicine

When the Human Genome Project got off the ground in 1990, we didn’t think of it as a Big Data project, but that’s what it was. By the time a complete human genome was mapped in 2003, some of the precursors to the Big Data movement had already started to percolate in the tech world.

So, it’s no surprise that the health care and pharmaceutical sectors are two of the most aggressive early adopters of Big Data tools, since there’s already a successful track record in place.

The Human Genome Project has also illustrated a sort of Moore’s Law of Big Data. You can already get an incomplete, but useful, snapshot of your genome from sites like 23andMe for $100 or less, and the push to drive down the cost of mapping your entire personal genome for that same price is already well underway. Prices have fallen each and every year. You can map your complete genome now for somewhere between $1,000 and $5,000. Back in 2007, this would have cost you $1 million, minimum.

Startups such as Life Technologies (recently acquired by Thermo Fisher Scientific) and InVitae are doing their best to make genome mapping something everyone can afford, which will lead to personalized treatments for everything from cancer to rheumatoid arthritis.

Emory University Hospital and IBM develop ICU of the future

Emory University Hospital is using software from IBM and Excel Medical Electronics (EME) for a research project that has a goal of creating advanced, predictive medical care for critical patients through real-time streaming analytics.

Emory is testing a new system that can identify patterns in physiological data in order to instantly alert clinicians to danger signs in patients. In a typical ICU, a dozen different streams of medical data light up the monitors at a patient’s bedside – including heart physiology, respiration, brain waves, and blood pressure. This constant feed of vital signs is transmitted as waves and numbers and routinely displayed on computer screens at every bedside. Currently, it’s up to doctors and nurses to rapidly process and analyze all this information in order to make medical decisions.

Today, any small deviation from the norm, which could be an early warning sign, usually goes unnoticed.

The system being piloted at Emory uses EME’s BedMasterEX, IBM InfoSphere Streams, and Emory’s analytics engine to collect and analyze physiological patient data in real time. The new system will enable clinicians to acquire, analyze, and correlate medical data at a volume more quickly than they could even dream of a few years back.

“Accessing and drawing insights from real-time data can mean life and death for a patient,” says Tim Buchman, MD, PhD, and director of critical care at Emory University Hospital. “Through this new system we will be able to analyze thousands of streaming data points and act on those insights to make better decisions about which patient needs our immediate attention and how to treat that patient. It’s making us much smarter in our approach to critical care.”

The software identifies patterns that could indicate serious complications like sepsis, heart failure, or pneumonia, aiming to provide real-time medical insights that clinicians can act on immediately.  

Penn State’s Salis Lab helps researchers engineer synthetic organisms

Howard M. Salis, an assistant professor in Penn State University’s chemical engineering department, taught himself how to code and built a high-performance computing web portal, the Salis Lab, that enables researchers in the synthetic biology and metabolic engineering fields to use computational methods to design synthetic organisms.

“Microorganisms are the best chemists on Earth,” Salis says. “If we learn to take advantage of them, we can manufacture a whole diversity of products. In the past, genetic engineering was more like tinkering and trial and error.”

In other words, genetic engineering was more like natural selection itself, random and slow, but with a much smaller pool of subjects.

“Synthetic biology, on the other hand, is more of an engineering discipline. We want to quantify everything. We develop bio-physical models that we can use to make quantitative predictions about what will happen when DNA mutates in various ways,” Salis explains.

Synthetic biology involves extremely complex algorithms, so the project is hosted on AWS Elastic Compute Cloud, which can scale up or down as needed. The number of possible mutations in a short DNA sequence is greater than the number of atoms in the universe. Salis Lab has become wildly popular with more than 2,000 biotechnology researchers designing over 30,000 synthetic DNA sequences through the web portal in the last two years.

The applications for this are as varied as the researchers’ imaginations. One goals is to figure out a way engineer microorganisms that will provide a fuel source economically competitive to the use of fossil fuels. A more mundane use case is developing pigments for blue jeans.

Even more amazing is the predictive power researchers can tap into. “Using our models, we can actually predict evolution,” Salis says. “We can simulate the effect of DNA mutations to predict the most probably course evolution will follow.”

Eventually, this could allow researchers to develop microorganisms that are resistant to evolution.

The possible use cases of this are staggering. There are billions of microorganisms in the world, and each has parts of its genome that we could potentially put to work for us. It’s an enormous Big Data challenge to sequence those genomes, quantify and catalog them, and, finally, predict how to combine them in useful ways. But it’s a challenge that researchers like Salis are eager to tackle. 

Georgetown’s Global Insight Initiative tackles “Big Problems”

Georgetown University’s Global Insight Initiative pulls data from around the world to gain insights into societal trends. The Global Insight Initiative analyzes data, but first needs to pull it, organize it, then package it to answer very complex questions.

“The world is a really complex system; there are 7 billion people interacting and competing for resources,” says J.C. Smart, director of the Global Insight Initiative at Georgetown University. The world has 40,000 cities, 12 million miles of roads, 800 million automobiles, and on and on. “Understanding how those all interact, and how they’re all dependent on one another is a very complex system. It’s a system of systems. That is Big Data, but more to the point, when you’re looking at the planet, it is Big Knowledge.”

The Global Insight Initiative needed data integration tools to manage the data volume in order to improve their knowledge base. “The knowledge base, just to give you an estimate of the number of things we’re talking about, we’re talking about a trillion objects and a quadrillion relationships,” Smart explains.

Kapow Software worked with Georgetown University’s Global Insight Initiative to automate high-volume data integration in order to expand the Initiative’s knowledge base. This involves accessing 20,000+ web sources from 162 countries representing 42 native languages to look at the planet and derive that “Big Knowledge.” Before automation, this process was so manually intensive that it would take dozens of people to find, pull, and organize documents and other web artifacts. After that, where do you find the time or resources to analyze the collection of information?

The Global Insight Initiative used Kapow’s software to create automated data integration flows (which you can think of as info-gathering robots). Once deployed, these infobots let a single user (who need not have any coding skills) run and manage thousands of automated data integration applications at any time in order to explore an integrated view of what could be wildly disparate data.

Now, the Global Insight Initiative will try to find answers to really hard “Big Problems,” such as: How do we best deploy water resources? How do we minimize the spread of diseases? How do we manage power distribution? How do we manage locations of health-care clinics to offer access to as many people as possible, and how do we position physician resources when disasters and catastrophes hit?

LA ExpressPark seeks to ease congestion and reduce pollution

Los Angeles’ downtown district has experienced significant growth over the last decade, transforming itself from a part of the city best known for skid row to an entertainment and business hot spot. Along with the growth, however, came tremendous traffic clutter. As drivers search for open spaces, they often circle the block for 30 minutes or more.

“Parking in downtown Los Angeles had become an expensive game of chance,” says David Cummins, senior vice president and managing director, parking and justice solutions, Xerox.

Making matters worse, on-street parking prices at meters rarely matched demand. Prices were uniform within a given area and were often the same, or cheaper, than garages a few blocks away. According to research from UCLA professor Donald Shoup, as much as 74 percent of congestion in downtown areas is due to drivers hunting for street parking. In a city where people already drive too much, there was really no incentive for drivers to park farther away.

To better match supply and demand, and hopefully ease congestion along the way, the city brought in Xerox to develop the LA ExpressPark parking system. Xerox installed hockey-puck-sized sensors in parking spaces to detect vacancies. Then, to better align supply with demand, Xerox developed an algorithm-based dynamic pricing engine to raise fares in high-occupancy blocks (to encourage turnover) and lower fares on fairly empty blocks (to encourage people to go out of their way a little bit).

As a transplant to L.A., I’ve often puzzled over the fact that Angelenos would prefer to circle the block endlessly than simply park two blocks away and endure a five-minute walk. What didn’t occur to me is that part of the problem, apparently, was a lack of awareness. If people knew parking was convenient (and cheaper) two blocks away, more would take advantage of it.

What happened when supply and demand were brought in line was that rates were actually lowered for 60 percent of meters, while they went up on only 20 percent of them. (Others stayed the same.)

To direct drivers to those empty spaces, new variable message signs have been deployed, signs which can be updated automatically as conditions change. The information is also shared with smartphone apps, such as Parker and Park Me, as well as the L.A. City website. Soon, Xerox intends to push the data directly to the navigation systems of the vehicles, which would automatically direct drivers to the nearest open spot to their destinations and perhaps even automatically pay for the parking, as well.

The early results have been promising. The city is already benefiting from increased overall usage in the less busy areas of town, and, even though rates are lower overall, revenue is already up 2 percent.

Better yet, congestion has started to ease and should improve even more as drivers learn about LA ExpressPark. “Parking managers now have complete and immediate visibility into what is happening on streets in their city and can make data-driven decisions on everything from rate structures to meter collections. Merge technology combines multiple vendors – from violation tickets processing to maintenance crews to collection crews – so that everything is readily available for the parking authority. Using the data in this way improves performance and creates additional revenues,” Cummins explains.

Cummins notes that the early results from this program prove that data-driven decisions can help change wasteful driver behaviors in order to reduce congestion and pollution.

Jeff Vance is a Santa Monica-based writer. He’s the founder of Startup50, a site devoted to emerging tech startups. Follow him on Twitter @JWVance, or reach him by email at jeff@sandstormmedia.net.