Here's the scoop on one of the fastest growing and most talked-about roles in IT.
endif; ?>After months of high unemployment and a still-wobbly economy, any good news from the jobs market is going to get some traction. But even that doesn’t seem to fully explain the attention surrounding a suddenly very “in” job title: data scientist.
According to CNN, data scientist is one of the best new jobs of 2012, and an article in the Harvard Business Review called it the “sexiest” job of the 21st century.
The allure surrounding this role is in direct correlation with the general market’s interest in big data and analytics — tools of the trade for data scientists, who are tasked with unearthing meaningful correlations within ever-mounting data volumes and turning them into profitable business insights.
What’s more, there is often a unicorn status imbued on people who fit this multifaceted position, which blends computer science, advanced quantitative concepts, business domain knowledge and communication skills. With demand for data scientists exceeding supply, salaries for these workers range in the six figures, according to Matthew Ripaldi, senior vice president at Modis, a staffing firm.
[ALSO: Could data scientist be your next job?]
Recruiters also agree that the data scientist position is fast-growing, even if the number of job postings is not yet staggering. “When we started looking at this position two years ago, there were just eight postings, and now there are 42,” says Tom Silver, senior vice president for North America at job search site Dice.com. “Forty-two out of 83,000 jobs is not huge, but I would suspect postings to grow even more in the future.”
With all the attention, it’s only natural that people with any background in data and computing might wonder, who are these people, and could I become one? We’ve tried to answer some of the most basic questions here.
What is a data scientist?
The answer to this deceptively simple question depends on who you ask. A widely accepted definition comes from Hilary Mason, chief scientist at Bit.ly: someone who can obtain, scrub, explore, model and interpret data.
Neil Raden, CEO at Hired Brains, goes a bit deeper, categorizing data scientists into two groups.
Type I – are true scientists who research and create algorithms and methods, publish papers and actively participate in their discipline’s communications. These individuals are found most often in research, academia and organizations, where new methods and algorithms are the core of the enterprise (think Google, Amazon, Wall Street), Raden says.
Type 2 – the group more often referred to in today’s hiring market — are not scientists but practitioners, Raden continues. These are experts in statistical and mathematical modeling and development, who understand and employ quantitative methods, as well as design, test and deploy models.
Jacob Spoelstra, global head of R&D at Opera Solutions, also distinguishes between what many people broadly categorize as data scientists and the work he and others do at Opera, a provider of predictive analytics as a service.
Opera’s “data scientists” — who would best fit in Raden’s Type 1 category — work at the machine learning level, developing statistical models and pattern recognition algorithms that find and extract predictive intelligence from massive data flows. They turn these findings into directed actions that help improve business by, for instance, reducing financial fraud or detecting risky mortgages. Companies like Google employ hundreds of this type of data scientist, Spoelstra estimates, and of Opera’s roughly 700 employees, about one-third are machine-learning specialists.
Meanwhile, Greta Roberts, CEO at Talent Analytics Corp., believes the current understanding of the data scientist job actually encompasses four functional roles. After conducting a survey that asked data scientists to detail how much time they devoted to 11 analytics functions, four clusters emerged: data preparation professionals (who spend most of their time on data acquisition, preparation and analytics); programmers (who program and do some analytics), managers (who concentrate on data management, administration, presentation, interpretation and design) and generalists (who do a little bit of everything).
[ALSO: What it takes to become a data scientist]
“When I started hearing all the hubub, I thought, ‘Nobody fits that definition — how could they possibly?'” Roberts says. “Because it’s a newer role, I think people are throwing everything in there. And when you over-specify, you come up with a null set.” What many businesses see as a data scientist, Roberts says, is actually a variety of functions performed by a group. And while there still may be a shortage of people to fill these roles, she says, the situation is a far cry from hunting unicorns, as plenty of people possess the natural aptitude to grow into one or more of the needed roles.
What are the necessary skills and credentials?
As Roberts indicates, lists detailing data science skills have proliferated on the Web, and they can be daunting. Most specify experience with advanced math, statistical analysis (including tools such as R, SAS and Stata), programming (including languages such as C, C++, Python and Java), SQL databases, platforms like Hadoop and MapReduce, data mining and modeling, data visualization, creativity, communication skills and business understanding.
And it’s true that data scientists do require skills and capabilities that are distinct from previous generations of data analysts, according to Raden. For instance, they need to be able to handle the wide variety of data available today and the resulting array of analysis that can be employed, he says.
They need programming skills, as well as a background in quantitative methods and an investigative and modeling orientation. And they must be able to discern what is and isn’t meaningful when it comes to the data, Raden continues. Effective data scientists also need sufficient business domain knowledge and the ability to communicate complex subjects to others who lack the background in the tools and methods employed, he says.
What pushes data scientists ahead of other analytics professionals, Ripaldi says, is the ability to communicate – often to the C-suite — what the data tells them, as well as how to act on the findings. “You can analyze all the data you want, but if you can’t articulate what it’s telling you, you’re not a data scientist,” he says. After all, the goal is to advance business strategy, such as reducing customer churn, targeting offers across channels and mitigating financial risks.
Then again, Roberts sees inherent conflict in these requirements, she says. “They have to be able to sit and look at data for days at a time and then flip the switch and be an engaging presenter? That’s two different people.”
Opera – which, again, hires data scientists of the machine-learning variety – looks for people with a background in a quantitative field, an aptitude for mathematics and statistical concepts, the ability to instantiate these concepts in computer programs, comfort with large volumes of data and an interest in solving real business problems.
“We’re comfortable with someone who needs to learn the machine learning algorithms if they demonstrate an affinity for math and problem-solving skills,” says Joseph Milana, Global Head of Analytics. “They may not be an applied mathematician or have built a neural network, but they should demonstrate energy and interest for us to bring them in.”
What background lends itself to becoming a data scientist?
At Opera, most successful applicants have higher levels of academic training and even a Ph.D. “Given the advances in machine learning science and the new techniques that are emerging, scientists do need to have advanced training and be steeped in the latest thinking,” Milana says. Even on Dice, half the data scientist postings specify a Ph.D., according to Silver. “It’s not absolutely necessary, but it’s a major bonus,” he says.
Opera hires across a variety of data-driven disciplines, including computer science, electrical engineering, statistics, mechanical engineering and physics. Such cross-disciplinary knowledge can be useful, Milana says; for instance, he has seen formulas originating in hydrology applied to stock market trading signals.
For the greater pool of data scientists, Raden believes a Ph.D. is unnecessary and that people who now work in business intelligence and quantitative analysis, with some previous exposure to advanced math and statistical modeling, could grow into these roles at a company that offers mentoring and training in key areas like predictive modeling and big data.
Roberts agrees that the focus on specific skills and academic credentials can sometimes be a proxy for how potential candidates think. “What they’re trying to measure is, ‘Do you like to learn?’ but there are a bunch of ways to get at that,” she says. In the Talent Analytics survey, the innate characteristics of data scientists included curiosity, creativity, objectivity, structured thinking and attention to detail, she says. Milana and Spoelstra agree that the most important natural traits they look for in candidates include curiosity, logical thinking, common sense, perseverance, practicality and good judgment.
There is no question that the need for data scientists is only going to increase. But because the role is a relatively new one, it will only see more change over time, both in terms of what these professionals do and how companies organize, attain and develop the needed talent.
“This is a huge opportunity for people in IT, project management and product management that aren’t afraid to learn on their own and burn the midnight oil to figure things out,” Roberts says.
Brandel is a freelance writer. She can be reached at marybrandel@verizon.net.




