Tag

data

Browsing

By Mark Eggleton

Data is referred to as the oil or even the soil of the 21st century. Either analogy is apt as they both illustrate the importance of data and how it will power or feed the world economy in the digital age. Data is the fuel of the future as well as the rich soil which everything will grow.

Its ubiquity already spreads far and wide ranging from the consumer data story such as knowledge of our online search history, financial transactions and social media interactions. More widely, data drives the analysis of whole industries, small businesses or traffic patterns.

It has moved far beyond crunching the numbers to find out your age, gender and income to gaining real time insights into how the world works and (hopefully) how it can be improved.

A good example would be the humble automobile. There are millions of cars on the road yet most of them sit idle over 95 per cent of the time. Better use of data should reduce the number of vehicles on the roads and increase the utilisation rate of cars actually on the roads. Drive share programs and driverless vehicles will mean fewer cars are sold but the data generated by each one as well as each user and their travel patterns will be invaluable.

Yet while every technological soothsayer suggests all of this is just around the corner, there’s still an extraordinary amount we don’t know as the new economy takes shape.

NSW Chief Data Scientist and CEO of the NSW Data Analytics Centre Dr Ian Oppermann is a little more sanguine about where we are now.

“The way companies have been taking advantage of a new digital future is better targeted services, more individual services. Quite often we talk about know your customer or a market size of one.

“For government, we’re helping agencies re-think the delivery of services. It’s doing old things in new ways where you can actually look across barriers, across agency boundaries and silos.”

Oppermann warns organisations should not fall into the trap of making decisions based solely on data as it’s a very simplistic observation of the world.

“What we do with data analytics or with artificial intelligence is we try to recapture the information in that data and then make informed decisions based on the little pieces of information scattered throughout many different sources.” (Dr Ian Oppermann, NSW Chief Data Scientist and CEO of the NSW Data Analytics Centre)

He cautions against blind faith in algorithms or in data “as a dangerous place to be”.

“It’s like following the GPS down a goat track even though when you look outside into the real world you realise you really should not be driving down that road. As long as we’re aware of the risk and as long as we question the results then I think we stand ourselves in good stead to make better decisions.

“What we ultimately want, is to help people use more data from a variety of sources to assist them make better decisions but we also have to build trust.”

Oppermann says trust comes with so many different aspects and it’s an evolving journey. He cites what we already do with banks as an example of how far along the journey we already are when it comes to digital trust.

“The data which a bank holds, is our salaries, it’s our pay – it’s something we can translate to cash but realistically banks are data centres or trust centres. We are quite comfortable having our salary paid into a bank where it goes in as data, sits there as data and we draw it out as data until we use an ATM. All the while it’s just data until we get it to manifest as a polymer bank note.

“We trust it implicitly and explicitly because we have been trusting banks for hundreds of years. Most of the time we’re pretty comfortable dealing with a bank but if you take that same data and say, now this data isn’t money, it is information about me or information about my preferences or other people we don’t actually have that same level of trust.

“Even if the governance processes, the security processes, the decision-making algorithms, are exactly the same we don’t have that same level of trust because we are not used to the idea of a government or a Google or a web services company delivering services to us in a way that we’ve interacted with for hundreds of years.”

For Oppermann, our data journey is just similar to the journey from gold to paper, to bank notes and now to data. He says ultimately trust will build slowly because of reliable and expected performance.

Part of the trust problem exists around the fear of too much information being held by too few. The big data refineries such as Amazon, Alphabet, Facebook and Apple already have a monstrous first mover advantage as do many financial institutions. This has bred the fear they are too large and are in danger of becoming monopolistic in the same way Standard Oil was in the United States in the early 20th century. Many have asked the question as to whether they need to be broken up or heavily regulated.

Partner in Charge, KPMG’s Data & Analytics in the Netherlands Professor Sander Klous, says governments are trying to play catch-up but it’s difficult because the old rules around ethics have been thrown on their head.

“We know how to apply them to human activity but how do you apply ethics to an algorithm? It’s something completely new,” he says.

As to whether the big data refineries should be broken up, he indicates that data is the new element in antitrust considerations. It’s a winner takes all ecosystem, where it becomes impossible to outperform the largest players because of all the data they possess.

“It’s a bit like the big banks where governments wanted to exercise some control over them because they became too big to fail. It’s the same with Google or a Facebook, if either of them broke down tomorrow, you could claim they are probably too big to fail as well,” Klous says.

Klous does suggest the sheer size of the large data refineries will see them eventually broken up because they’re unsustainable.

“The data refineries are basically just really big pipes where raw material is processed and something smart comes out. So, the whole idea that one party is controlling that pipeline is too rigid to be a sustainable model.

“I think what you need is multiple parties that are able to work together in a platform like structure and the data refinery has to become more complex because there is more than one party in control.”

He draws an analogy with traffic lights where one party is in control as opposed to a roundabout which is a simpler concept but more parties in control as long as everyone abides by a simple rule.

“In a roundabout, there is a simple standard where right goes first and then you make your decision to enter and everything flows. The same applies in a data environment where there there can be a simple set of standards that need to be complied with and then as you add intelligence (or information) to the data refinery it informs decisions. You’re not relying on one single party to make the right decisions.”

He says data refineries will eventually turn into this platform model where multiple parties will collaborate to create value and eventually domain specific refineries will develop in areas around eg. health or logistics where the dominant players will work together.

As for the future, Klous says we are inevitably moving towards a smart society but there are things we need to get under control. Ideally, he would like to see some sort of control framework in place that allows individuals to be able to trust what the large data refineries are actually doing.

“We are sorting out how to deal properly with privacy and other ethical issues without losing benefits like more convenience or increased efficiency.”

 

By Mark Eggleton

Sourced from Reports.afr

Researchers urged to hone methods for mining social-media data, or investment in marketing will be wasted.

By MediaStreet Staff Writers

A growing number of people, from marketers to academic researchers, are mining social media data to learn about both online and offline human behaviour. In recent years, studies have claimed the ability to predict everything from summer blockbusters to fluctuations in the stock market.

But mounting evidence of flaws in many of these studies points to a need for researchers to be wary of serious pitfalls that arise when working with huge social media data sets. This is according to computer scientists at McGill University and Carnegie Mellon University.

Such erroneous results can have huge implications on data gleaned from social media. A lot of marketing investment could be placed in the wrong areas.

The challenges involved in using data mined from social media include:

  • Different social media platforms attract different users – Pinterest, for example, is dominated by females aged 25-34 – yet researchers rarely correct for the distorted picture these populations can produce.
  • Publicly available data feeds used in social media research don’t always provide an accurate representation of the platform’s overall data – and researchers are generally in the dark about when and how social media providers filter their data streams.
  • The design of social media platforms can dictate how users behave and, therefore, what behaviour can be measured. For instance, on Facebook the absence of a “dislike” button makes negative responses to content harder to detect than positive “likes.”
  • Large numbers of spammers and bots, which masquerade as normal users on social media, get mistakenly incorporated into many measurements and predictions of human behaviour.
  • Researchers often report results for groups of easy-to-classify users, topics, and events, making new methods seem more accurate than they actually are. For instance, efforts to infer political orientation of Twitter users achieve barely 65% accuracy for typical users – even though studies (focusing on politically active users) have claimed 90% accuracy.

Many of these problems have well-known solutions from other fields such as epidemiology, statistics, and machine learning. The common thread in all these issues is the need for researchers to be more acutely aware of what they’re actually analysing when working with social media data.

Social scientists have honed their techniques and standards to deal with this sort of challenge before. Says Derek Ruths, an assistant professor in McGill’s School of Computer Science, “The infamous ‘Dewey Defeats Truman’ headline of 1948 stemmed from telephone surveys that under-sampled Truman supporters in the general population. Rather than permanently discrediting the practice of polling, that glaring error led to today’s more sophisticated techniques, higher standards, and more accurate polls. Now, we’re poised at a similar technological inflection point. By tackling the issues we face, we’ll be able to realise the tremendous potential for good promised by social media-based research.”

 

Advertising is strangling the web, and that may have to do with the declining value and lack of transparency associated with the players that dominate it. (Image: Anders Emil Møller / Trouble).

By for Big on Data.

The problem with advertising data and what to do about it. Plus, the future of big data architecture, and other stories from the Ad Tech trenches.

The greatest minds of this generation are wasted on advertisement. Or at least, that’s what someone who has been there and done that thinks. Like most successful aphorisms, this raises eyebrows, drives heated discussions, and strikes a point or two.

Marketing and advertising have enormous influence on society at large — business, technology, media, culture, and data. So, discussing with people working on the intersection of those can offer some insights on the state of the union of Big Data and Ad Tech.

Advertising is big, and so is its data

Advertising is a multi-billion dollar business that has been going through the process of digital transformation for a couple of decades already. Some of today’s most advanced, powerful, and influential companies have advertising embedded in their core.

A good part of the innovation that has been driving big data has come about as a response to the needs of advertising at scale before getting a life of its own. MapReduce, for example, the blueprint for Hadoop’s first incarnation, was originally developed and deployed at scale at Google.

But although Facebook and Google, the ‘Big Two,’ are by far the biggest players in the digital advertising space, they are not the only ones. The Ad Tech, or Marketing Tech, scene is booming, and programmatic marketing is taking over quickly.

Mike Driscoll, CEO of Metamarkets, points out that marketing is being digitally transformed and marketers are following suite. Metamarkets is part of the Ad Tech wave, and its core business is to provide marketers with insights on their digital presence.

“The future will be digital,” Driscoll says. “CMOs (Chief Marketing Officers) are turning to CMTOs (Chief Marketing Technical Officers). But as marketers are going digital, they are also starting to have less trust in some of the channels they’re buying from. Investing in technology means they are now able to hold their partners accountable.”

The problem with advertising data

Driscoll has more than anecdotal evidence and opinions here. Metamarkets just published a survey called the Transparency Opportunity, in which it attempts, as it says, to quantify the benefits of trust.

The findings of the survey show that almost half of the brands using programmatic media buying believe lack of transparency is inhibiting its future growth and scale. But what do we talk about when we talk about transparency here?

“All marketers work with the established duopoly — Google and Facebook,” says Driscoll. Metamarkets also works with other platforms, such as Twitter and AOL, but it’s the Big Two that dominate the advertising market. While the market is growing, nearly all of that growth is driven by Google and Facebook.

c07mroxcaaqf0m.jpg
The advertising pie may be growing, but the growth is driven by a duopoly. (Image: Jason Kint)

It’s not hard to see where this is going: Google and Facebook dominating the market and dictating their terms. This presents a problem for all parties involved. For media, depen

dence on advertising translates as dependence on the Big Two. Media are trying to find ways to cope and come up with new models of doing business while maintaining their editorial independence.

For consumers, the ever-increasing volume of advertising means they are constantly bombarded by a barrage of ads. They are told that this is the price they have to pay for having access to free content, and to a certain extent, it is true. The problem is that more and more advertising is strangling content, consumer fatigue is taking over, and the value and effectiveness of advertising is dropping.

Obviously, this presents a problem for advertisers and their clients. “Marketers want more transparency. They would like to get a receipt for what they buy, instead of a powerpoint and a good story,” Driscoll says. “Brands are asking from their partners to provide better analytics. Historically, channels have been providing results, but not analytics on the results.

Digital transformation means that we don’t just buy goods anymore, we also buy data about the goods. Take AWS, for example: When they started, they just provided the service, but by now, you also get analytics to go with it. Major channels need to invest more not just in internal technology, but also in providing better data access to their partners.”

The big guys will just not share

But, seriously, this is Google and Facebook we’re talking about. Are we to believe that the most iconic data-driven organizations in the world can’t make the right data available to marketers? “Ask any marketer and they’ll tell you — the big guys will just not share. It’s not in their interest to be transparent, but rather to be as less transparent as possible,” Driscoll says.

transparencyinhibiting.jpg
Almost half of brands being advertised see lack of transparency as a problem. Image: MetaMarkets

So, what can be done to deal with this? The real power of marketers is the power to check them, according to Driscoll. “If you look at the leaders emerging in Fortune 500, the next generation of marketers are technologists, and they are demanding independent audits and data. Consider this:

For a long time, advertisers wanted to know if their ads were viewed or not. Facebook says, ‘OK, we’ll measure it ourselves.’ And they got away with it for a while. They reported their own view-ability stats, just like NBC used to report how many people viewed their own shows.

In the last year, that has changed. Marketers says, ‘We will not send advertisements to Facebook, unless we have an independent source of truth.’ So, Facebook responded by providing access to their data to a company called Moat, which built a business model around auditing Facebook data.

That does not mean Facebook and Co., will give away all of their data — you also have to consider privacy issues here. But when you talk to brands, even though they will not say that in public, they are actually doing that. When you have budgets in the tends of millions, you can do that — pull data out and do what every marketer would like to do: Build a unified view over their channels.”

Analyst super powers

But what about the rest of the world, the ones that don’t have the budget to cope with this? Perhaps regulation would be needed, so if the big guys won’t share, someone should make them?

“We’ve been hearing rumours about the Congress getting involved, but for most businesses that would be the last resort. It’s not the ideal solution for marketers or media companies, especially considering the all-time low approval ratings in the US right now,” Driscoll says. For him, the answer is in marketers investing more in analytics.

On the one end of the continuum, organizations can do it all themselves, using infrastructure like Hadoop and analytics tools that sit on top and can help them collect and analyze the data they need. On the other end, Metamarkets touts itself as the right solution for marketers.

Metamarkets is a domain-specific solution that builds on four pillars: Fast data exploration, intuitive visualization, collaboration, and intelligence. Driscoll elaborates: “Scale is a requirement, and we are quickly moving towards streaming events and data.

Interactive visualization helps you understand what’s going on. You need more than dashboards. Dashboards may update, but the questions they answer stay the same. You need collaboration — like Slack for data, that helps teams communicate and share methods and insights.

And you need intelligence. In analytics, you spend 80 percent of your time preparing data and 20 percent actually doing analysis. We have ETL connectors for a multitude of platforms that help get the data where you need them. Plus, it’s one thing to show data, and another thing to search for insights.”

Metamarkets tries to look at what analysts do and automate that to suggest root causes. For example, a campaign running behind targets is something that can be monitored using metrics. But to get to the reason why this is happening, an analyst would slice and dice data per region or demographics.

Metamarkets says they can automate this process and suggest root causes, evolving from tracking statistical significant signals to deriving business-focused insights. “We let analysts specify metrics they are interested in, and then perform root cause analysis for them. We believe in machine and human working side by side, not in replacing analysts, but in giving them super-powers,” Driscoll says.

Data at advertising scale and the future of pipelines

As Metamarkets has been on the forefront of data at advertising scale, and Driscoll himself has served as its CTO, he shared some insights on the evolution of big data architecture: “We have been pushing the limits of scale, so we encounter problems before others do,” he says.

This has resulted in MetaMarkets developing and releasing Druid, an open-source distributed column store. “We created Druid because we needed it and it did not exist, so we had to build it. And then we open sourced it, because if we had not, something else would have come along and replaced it.”

image04-768x352.png
MetaMarkets has been evolving its data pipeline, but still not turning to Kappa architecture on the grounds that its clients are not ready for it. (Image: MetaMarkets)

Druid is seeing some traction in the industry. Case in point, when Hortonworks’ engineers recently presented their work on the combined used of Hive & Druid at the DataWorks EMEA Summit, they attracted widespread interest. Could this mean there may be a valid case for building a business around Druid?

“We have the largest deployment in production, and we love being part of the community. Druid is used by the likes of Airbnb and Ali Baba. But we have no plans of building a business around it. We don’t believe the future is around data infrastructure, which is becoming a commodity, and we don’t want to be competing against the Googles of the world there.

Sure, this may be working for companies built around Hadoop, but commercialization of open source needs widespread adoption to succeed. But I can tell you that Cloudera and Hortonworks are looking to add Druid to their stack and to the range of services they offer.”

Read also: Has the Hadoop market turned a corner? | Open source big data and DevOps tools: A fast path to analytics applications | Finding the anomalies in big data with machine learning | Cloudera’s new data science tool aims to boost big data and machine learning for businesses (TechRepublic)

Driscoll does not believe in horizontally expanding Metamarkets, even though its experience in building data pipelines at scale could in theory be applied to other domains beyond advertising. Its own pipeline has been evolving, going from Hadoop to Spark and from Storm to Samza.

“Spark is more mature and it meets our needs at this point, and we also feel about the same way about Samza,” he says. “But we see streaming as the future of our pipeline. When you work with streaming, there’s a sort of CAP theorem equivalent that applies there.

In distributed data stores, you have consistency, availability and partition tolerance, and you can pick two of those that your system supports simultaneously. In streaming data, you have accuracy, velocity, and volume, and your system can only support two of those simultaneously.

This is why we think the model supported by Apache Beam, Google Data Flow, and Apache Flink will be key going forward. When streaming at scale, there’s no such thing as objective truth, so you have to rely on statistical approximation and on using watermarks.

Do we see our current Lambda architecture giving way to a flattened, Kappa architecture? When you work on the bleeding edge of real-time architecture, the ability of organizations like Metamarkets that are in the business of integrating data from other sources is important.

But when it comes to other companies, not many are yet at the point where they can stream data out. Only the most sophisticated, agile companies out there are able to do this. At this point, only about 50 percent of our clients are there.”

By for Big on Data

Sourced from ZDNet

Sourced from CRAIN’S New York Business

Time Inc. is planning to sell some magazines or other properties as the struggling publisher tries to push ahead with a digital strategy and move past months of talks with potential acquirers.

The owner of Sports Illustrated and People will look to offload “relatively smaller” titles in its portfolio and other “non-core” assets, chief executive Rich Battista said Wednesday on a conference call. He didn’t name the assets.

Battista added that Time is open to joint ventures with other companies and interested in an outside investor who could provide capital “for a particular opportunity.”

Last month, Time announced that it was sticking with its online strategy rather than sell itself after months of negotiations with potential suitors, including Meredith Corp. and a group including Pamplona Capital Management and Jahm Najafi. New York-based Time was said to be holding out for more than $20 a share.

The shares slumped as much as 19% to $12.20 on Wednesday. The magazine publisher reported first-quarter revenue of $636 million, missing the $642 million average of analysts’ estimates. Its net loss widened as print advertising sales declined 21%. It also cut its dividend. Like other magazine publishers, Time is struggling to transform itself as print advertising dries up and the lion’s share of digital advertising dollars goes to Facebook Inc. and Google.

The magazine owner has spent months restructuring its business and replacing senior management, hoping to persuade advertisers to pour money into its magazine titles. This fall, Time plans to introduce a Sports Illustrated online video service with documentaries and insights from the magazine’s reporters, part of its growing push into video. Some of Time’s smaller titles include Sunset magazine and What’s On TV, which is based in the U.K.

Investor challenge

On an earnings conference, one of its investors demanded more detail about Time’s strategic plan.

“You constantly refer to this strategic plan, but you provide no numbers for the shareholders to basically grasp what this company will look like in two or three years,” said Leon Cooperman, of Omega Advisors Inc., which owns 3.9% of the magazine publisher, according to data compiled by Bloomberg.

“I think it’s incumbent upon the company to share with the shareholders, the people that have the money invested, what the strategic plan would yield,” Cooperman said. “Because I’m pretty confident that this company can be sold today at at least $18 a share.”

Cooperman urged the company to hold an analyst day and reveal its strategic plan in more detail. “Then we can make an intelligent decision whether we should agitate for a sale or be patient and give you guys a chance to do your magic,” he said.

Battista replied that Time hired an adviser to cut costs and believes it can reach $1 billion in digital revenue, but did not provide a timeline. In an interview, Battista said the company would provide some profit guidance going forward and “other insights when appropriate.”

“We feel really excited and confident in our plan,” Battista said.

Sourced from CRAIN’S New York Business

By Tobi Elkin.

A new report by ad-tech provider Blue Venn finds that 72% of marketers consider data analysis more important than social media skills.

The report, “Customer Data: The Monster Under the Bed?,” incorporates research from 200 U.S. and U.K. marketers, with the goal of identifying the attributes most needed to compete in the data-centric marketing landscape.

Key findings include:

–Data management is now considered more vital than social media (65%), Web development (31%), graphic design (23%), and search engine optimization (13%).

–However, 27% of marketers are still handing over the process of data analysis to IT departments.

–The focus on understanding and synthesizing customer data is especially strong at large enterprises, where four out of five marketers consider data analysis to be a “vital” skill.

–Data segmentation and modeling are also considered highly sought-after marketing skills, ranking higher than both Web development and graphic design within the enterprise space.

“In the age of big data, marketers have a better opportunity than ever before to truly understand their customers’ decision-making processes. Unfortunately, as it stands, most marketers simply don’t have the time, the knowledge or the tools necessary to undertake this task in a practical and effective way,” stated Anthony Botibol, marketing director at BlueVenn.

By

Sourced from MediaPost