Tag

data

Browsing

Experts say the privacy promise—ubiquitous in online services and apps—obscures the truth about how companies use personal data

You’ve likely run into this claim from tech giants before: “We do not sell your personal data.”

Companies from Facebook to Google to Twitter repeat versions of this statement in their privacy policies, public statements, and congressional testimony. And when taken very literally, the promise is true: Despite gathering masses of personal data on their users and converting that data into billions of dollars in profits, these tech giants do not directly sell their users’ information the same way data brokers directly sell data in bulk to advertisers.

But the disclaimers are also a distraction from all the other ways tech giants use personal data for profit and, in the process, put users’ privacy at risk, experts say.

[Companies] saying they don’t sell data to third parties is like a yogurt company saying they’re gluten-free…. It’s a misdirection.

Ari Ezra Waldman, Northeastern University School of Law

Lawmakers, watchdog organizations, and privacy advocates have all pointed out ways that advertisers can still pay for access to data from companies like Facebook, Google, and Twitter without directly purchasing it. (Facebook spokesperson Emil Vazquez declined to comment and Twitter spokesperson Laura Pacas referred us to Twitter’s privacy policy. Google did not respond to requests for comment.)

And focusing on the term “sell” is essentially a sleight of hand by tech giants, said Ari Ezra Waldman, a professor of law and computer science at Northeastern University.

“[Their] saying that they don’t sell data to third parties is like a yogurt company saying they’re gluten-free. Yogurt is naturally gluten-free,” Waldman said. “It’s a misdirection from all the other ways that may be more subtle but still are deep and profound invasions of privacy.”

Those other ways include everything from data collected from real-time bidding streams (more on that later), to targeted ads directing traffic to websites that collect data, to companies using the data internally.

How Is My Data at Risk if It’s Not Being Sold? 

Even though companies like Facebook and Google aren’t directly selling your data, they are using it for targeted advertising, which creates plenty of opportunities for advertisers to pay and get your personal information in return.

The simplest way is through an ad that links to a website with its own trackers embedded, which can gather information on visitors including their IP address and their device IDs.

Advertising companies are quick to point out that they sell ads, not data, but don’t disclose that clicking on these ads often results in a website collecting personal data. In other words, you can easily give away your information to companies that have paid to get an ad in front of you.

If the ad is targeted toward a certain demographic, then advertisers would also be able to infer personal information about visitors who came from that ad, Bennett Cyphers, a staff technologist at the Electronic Frontier Foundation, said.

For example, if there’s an ad targeted at expectant mothers on Facebook, the advertiser can infer that everyone who came from that link is someone Facebook believes is expecting a child. Once a person clicks on that link, the website could collect device IDs and an IP address, which can be used to identify a person. Personal information like “expecting parent” could become associated with that IP address.

“You can say, ‘Hey, Google, I want a list of people ages 18–35 who watched the Super Bowl last year.’ They won’t give you that list, but they will let you serve ads to all those people,” Cyphers said. “Some of those people will click on those ads, and you can pretty easily figure out who those people are. You can buy data, in a sense, that way.”

Then there’s the complicated but much more common way that advertisers can pay for data without it being considered a sale, through a process known as “real-time bidding.”

Often, when an ad appears on your screen, it wasn’t already there waiting for you to show up. Digital auctions are happening in milliseconds before the ads load, where websites are selling screen real estate to the highest bidder in an automated process.

Visiting a page kicks off a bidding process where hundreds of advertisers are simultaneously sent data like an IP address, a device ID, the visitor’s interests, demographics, and location. The advertisers use this data to determine how much they’d like to pay to show an ad to that visitor, but even if they don’t make the winning bid, they have already captured what may be a lot of personal information.

With Google ads, for instance, the Google Ad Exchange sends data associated with your Google account during this ad auction process, which can include information like your age, location, and interests.

The advertisers aren’t paying for that data, per se; they’re paying for the right to show an advertisement on a page you visited. But they still get the data as part of the bidding process, and some advertisers compile that information and sell it, privacy advocates said.

In May, a group of Google users filed a federal class action lawsuit against Google in the U.S. District Court for the Northern District of California alleging the company is violating its claims to not sell personal information by operating its real-time bidding service.

The lawsuit argues that even though Google wasn’t directly handing over your personal data in exchange for money, its advertising services allowed hundreds of third parties to essentially pay and get access to information on millions of people. The case is ongoing.

“We never sell people’s personal information and we have strict policies specifically prohibiting personalized ads based on sensitive categories,” Google spokesperson José Castañeda told the San Francisco Chronicle in May.

Real-time bidding has also drawn scrutiny from lawmakers and watchdog organizations for its privacy implications.

In January, Simon McDougall, deputy commissioner of the United Kingdom’s Information Commissioner’s Office, announced in a statement that the agency was continuing its investigation of real-time bidding (RTB), which if not properly disclosed, may violate the European Union’s General Data Protection Regulation.

“The complex system of RTB can use people’s sensitive personal data to serve adverts and requires people’s explicit consent, which is not happening right now,” McDougall said. “Sharing people’s data with potentially hundreds of companies, without properly assessing and addressing the risk of these counterparties, also raises questions around the security and retention of this data.”

Few Americans realize that some auction participants are siphoning off and storing ‘bidstream’ data to compile exhaustive dossiers about them.

Letter to ad tech companies from six U.S. senators

And in April, a bipartisan group of U.S. senators sent a letter to ad tech companies involved in real-time bidding, including Google. Their main concern: foreign companies and governments potentially capturing massive amounts of personal data about Americans.

“Few Americans realize that some auction participants are siphoning off and storing ‘bidstream’ data to compile exhaustive dossiers about them,” the letter said. “In turn, these dossiers are being openly sold to anyone with a credit card, including to hedge funds, political campaigns, and even to governments.”

On May 4, Google responded to the letter, telling lawmakers that it doesn’t share personally identifiable information in bid requests and doesn’t share demographic information during the process.

“We never sell people’s personal information and all ad buyers using our systems are subject to stringent policies and standards, including restrictions on the use and retention of information they receive,” Mark Isakowitz, Google’s vice president of government affairs and public policy, said in the letter.

What Does It Mean to “Sell” Data?

Advocates have been trying to expand the definition of “sell” beyond a straightforward transaction.

The California Consumer Privacy Act, which went into effect in January 2020, attempted to cast a wide net when defining “sale,” beyond just exchanging data for money. The law considers it a sale if personal information is sold, rented, released, shared, transferred, or communicated (either orally or in writing) from one business to another for “monetary or other valuable consideration.”

If you are a social media company and you’re providing advertising and people pay you a lot of money, you are selling access to them.

Mary Stone Ross, a co-author of the California Consumer Privacy Act

And companies that sell such data are required to disclose that they’re doing so and allow consumers to opt out.

“We wrote the law trying to reflect how the data economy actually works, where most of the time, unless you’re a data broker, you’re not actually selling a person’s personal information,” said Mary Stone Ross, chief privacy officer at OSOM Products and a co-author of the law. “But you essentially are. If you are a social media company and you’re providing advertising and people pay you a lot of money, you are selling access to them.”

But that doesn’t mean it’s always obvious what sorts of personal data a company collects and sells.

In T-Mobile’s privacy policy, for instance, the company says it sells compiled data in bulk, which it calls “audience segments.” The policy states that audience segment data for sale doesn’t contain identifiers like your name and address but does include your mobile advertising ID.

Mobile advertising IDs can easily be connected to individuals through third-party companies.

Nevertheless, T-Mobile’s privacy policy says the company does “not sell information that directly identifies customers.”

T-Mobile spokesperson Taylor Prewitt didn’t provide an answer to why the company doesn’t consider advertising IDs to be personal information but said customers have the right to opt out of that data being sold.

So What Should I Be Looking for in a Privacy Policy? 

The next time you look at a privacy policy, which few people ever really do, don’t just focus on whether or not the company says it sells your data. That’s not necessarily the best way to assess how your information is traveling and being used.

And even if a privacy policy says that it doesn’t share private information beyond company walls, the data collected can still be used for purposes you might feel uncomfortable with, like training internal algorithms and machine learning models. (See Facebook’s use of one billion pictures from Instagram, which it owns, to improve its image recognition capability.)

Consumers should look for deletion and retention policies instead, said Lindsey Barrett, a privacy expert and until recently a fellow at Georgetown Law. These are policies that spell out how long companies keep data, and how to get it removed.

She noted that these statements hold a lot more weight than companies promising not to sell your data.

“People don’t have any meaningful transparency into what companies are doing with their data, and too often, there are too few limits on what they can do with it,” Barrett said. “The whole ‘We don’t sell your data’ doesn’t say anything about what the company is doing behind closed doors.”

Feature Image Credit: Gabriel Hongsdusit

Sourced from The Markup

By

How to become a data-driven organization and why is this key for future business success.

A global pandemic has played a big part in elevating the data literacy of the ‘everyday man’, essentially demystifying data science for many of us. There now seems to be a much clearer connection between ‘what the data says’ and ‘what action we’ll take’. It’s a good example of data-led decision-making to drive the best possible outcomes.

In the same way, businesses can also drive better outcomes if they understand this connection and use it to transform their business model to one that is driven by data. Building a data-driven future is what will give businesses their competitive edge at a time when ’survival of the fittest’ counts the most.

With this in mind, many organizations are currently investing in data projects of one kind or another. Whether it’s data analytics, big data, AI, machine learning, data science or any other area of focus, the interest and investment in efforts to become ‘data-driven’ has been given significant extra impetus by the experiences of the last 12 months.

Whether the objective is to drive efficiencies, create competitive advantage or improve decision-making processes, it’s important to remember that isolated or independent data projects do not make you data-driven. Instead, the race to create a data-driven business infrastructure should be seen as a strategic journey, where organizations position data so that it empowers and delivers on business objectives. Success depends on transforming business models so that the whole is greater than the sum of its parts.

So, what does it take to become a data-driven organization? Here are a series of core principles that together, can help build a solid foundation, focus and measurable progress for using data as a strategic asset:

Leadership

The entire process rests on effective leadership, where a top-down perspective aligns the business with data strategy. Without this approach, it can prove impossible to instigate the culture shift required to truly become data-driven and ensure that initiatives are given the right emphasis, support and representation, as well as driving the education required at a leadership level.

Skills

Next, it’s important to evaluate the relevant skills – and the gaps – within existing teams. For instance, it’s not unusual for analytics skills to be spread across departments, but in creating the right focus, business leaders need to transition to a core, centralized practice to ensure consistency. This does not necessarily mean that teams have to be changed, but organizations must create best practice processes to focus their efforts. Ultimately, building a community of data professionals who share knowledge and work together can be hugely beneficial, even if they don’t work in the same teams on a daily basis.

Best practice

Looking more closely at best practice, the objective should be to move from sporadic and isolated data driven initiatives siloed in each department to an approach which ensures consistency of approach across the organization. This should always be based on a common understanding of how to deliver value from data effectively.

Governance

As skills and best practice processes become integrated into a data driven culture, it becomes more important to ensure governance increases. Indeed, establishing best-in-class governance and frameworks is essential to ongoing data-driven transformation, because it enables leadership to track progress against goals. In practical terms, leaders need to work with data practitioners to ensure that initiatives meet business objectives, that there is consistency in delivery and prioritization, as well as in the platforms and technologies used. At the same time, every organization must meet their data compliance obligations, especially relating to sensitive or personal information.

Education

Increasing the impact of a data-driven strategy is not just a matter of bringing the specialists together. Educating the business at large about the possibilities of analytics is an important part of the process so the whole business can share a common language around analytics and dispel preconceptions of what analytics can and can’t achieve.

Prioritization

As the impact of education efforts take effect, and business interest and knowledge of the potential of data driven decisions grows, many organizations find they are presented with a wide range of potential initiatives. Clearly, prioritization then becomes important, and key questions about each idea and option should include: will an initiative add significant, measurable value? Is the organization ready to implement data driven initiatives that may deliver meaningful results? Is the right data and platform available to make it work, and is the organization in a position to adopt the new practices each initiative will require?

Measurement

With priorities determined and actively being implemented, the process requires a structure to measure success in a consistent way so that all stakeholders can see the data driven program at work, rather than isolated instances of innovation. This is often pivotal for organizations in their efforts to move away from a series of data science projects to being a truly data-driven company.

There’s no doubt that investing time and resources in developing a data-driven culture can radically improve insight and decision making. In today’s rapidly changing business environment, spotting new opportunities and challenges, improving processes and working with greater insight into the variables that affect business success is vital. By adopting a rounded process that addresses these critical areas, businesses have the best chance of succeeding in their mission not just to become data-driven, but in their wider digital transformation strategy.

Feature Image Credit: (Image credit: Shutterstock / carlos castilla)

By

Sourced from ITProPortal

By

The pandemic forced D&A leaders to step up research and analysis to respond effectively to change and uncertainty, the firm says.

While much of the loudest buzz surrounding the impact of COVID-19 was focused on the dramatic shift from on premises to remote work, the pandemic further affected every aspect of the enterprise, which includes data and analytics technology. The uncertainty of what the tech industry would face forced D&A leadership to quickly find tools and processes — and put them in place — so they could identify key trends and prioritize to the company’s best advantage, said Rita Sallam, research vice president at Gartner, in the company’s recently released information.

Gartner has now identified 10 trends as “mission-critical investments that accelerate capabilities to anticipate, shift and respond.” It recommended that D&A leaders review these trends and consider and apply as necessary. Following is a summary from Gartner of the trends:

Trend 1: Smarter, responsible, scalable AI

Artificial intelligence and machine learning are key factors. Businesses must apply new techniques for smarter, less data-hungry, ethically responsible and more resilient AI solutions.  When smarter, more responsible, scalable AI is applied, organizations will be able to “leverage learning algorithms and interpretable systems into shorter time to value and higher business impact,” Gartner’s report said.

Trend 2: Composable data and analytics

Composable data and analytics leverages components from multiple data, analytics and AI solutions to quickly build flexible and user-friendly intelligent applications to help D&A leaders make the correlation between the discovered insights to actions they must execute. Open, containerized analytics architectures make analytics capabilities more composable.

Public or private, data is unquestionably moving to the cloud and composable data, rendering analytics “a more agile way to build analytics applications enabled by cloud marketplaces and low-code and no-code solutions.”

Trend 3: Data fabric is the foundation

D&A leaders use data fabric to help address “higher levels of diversity, distribution, scale and complexity in their organizations’ data assets,” as a result of increased digitization and  “more emancipated” consumers.

Data fabric applies analytics in order to constantly monitor data pipelines; data fabric “uses continuous analytics of data assets to support the design, deployment and utilization of diverse data to reduce time for integration by 30%, deployment by 30% and maintenance by 70%.”

Trend 4: From big to small and wide data

Using historical data for ML and AI models was rendered irrelevant, once changes based on the pandemic had an extreme effect on business. D&A leaders need a greater variety of data for better situational awareness because human and AI decision making grows more complex and demanding.

Therefore, D&A leaders need to choose analytical techniques that can use available data more effectively and they can with more insight that now requires less data.

“Small and wide data approaches provide robust analytics and AI, while reducing organizations’ large data set dependency,” Sallam said in a press release. “Using wide data, organizations attain a richer, more complete situational awareness or 360-degree view, enabling them to apply analytics for better decision making.”

Trend 5: XOps

DataOps, MLOps, ModelOps and PlatformOps, which comprise XOps, are necessary to achieve efficiencies and economies of scale through DevOps and using best practices of reliability, reusability and repeatability. This also reduces duplication of technology and processes and enabling automation.

Operationalization must be addressed initially and not as an afterthought because the latter is why most analytics and AI projects fail.  The report said,  “If D&A leaders operationalize at scale using XOps, they will enable the reproducibility, traceability, integrity and integrability of analytics and AI assets.”

Trend 6: Engineering decision intelligence

D&A leaders can make engineering decisions more accurate, repeatable, transparent and traceable, as decisions grow more automated and augmented. Gartner refers to “engineering decision intelligence,” which applies to a series of decisions  of business processes as well as grouped emergent decisions and consequences.

Trend 7: Data and analytics as a core business function

D&A is now making the shift into a core business function, rather than a secondary activity. D&A now is a shared business asset aligned to business results. Gartner noted that D&A silos break down because of better collaboration between central and federated D&A teams.

Trend 8: Graph relates everything

Graphs form the foundation of most modern data and analytics capabilities and are reliant on the foundation to find relationships between people, places, things, events and locations across a wide variety of data assets. D&A leaders rely on graphs as quick answers to complex business questions, which require contextual awareness and an understanding of the nature of connections and strengths across multiple entities.

Gartner predicts that by 2025, graph technologies will be used in 80% of data and analytics innovations, up from 10% in 2021, facilitating rapid decision making across the organization.

Trend 9: The rise of the augmented consumer

Today, most business users use predefined dashboards and manual data exploration, but this can lead to incorrect conclusions and flawed decisions and actions. Time spent in predefined dashboards will progressively be replaced when users’ needs can be delivered  with automated, conversational, mobile and dynamically generated insights customized through a predefined dashboard.

“This will shift the analytical power to the information consumer, the augmented consumer, giving them capabilities previously only available to analysts and citizen data scientists,” Sallam said.

Trend 10: Data and analytics at the edge

Support for data, analytics and other technologies are found in edge computing environments, closer to assets in the physical world and outside IT’s purview. Gartner predicts that by 2023, over 50% of the primary responsibility of data and analytics leaders will comprise data created, managed and analyzed in edge environments.

Gartner concluded: “D&A leaders can use this trend to enable greater data management flexibility, speed, governance, and resilience. A diversity of use cases is driving the interest in edge capabilities for D&A, ranging from supporting real-time event analytics to enabling autonomous behavior of things.”

Gartner Data and analytics summit

Gartner analysts offer more analysis on data and analytics trends at the Gartner Data & Analytics Summits 2021, taking place virtually May 4-6 in the Americas, May 18-20 in EMEA, June 8-9 in APAC, June 23-24 in India, and July 12-13 in Japan. Follow news and updates from the conferences on Twitter using #GartnerDA.

This article was originally published on TechRepublic

By

Sourced from ZDNet

Sourced from Forbes

Some agency clients aren’t able to address important questions their marketing partners need answered in order to devise the best strategy to meet their needs. Luckily, analytics tools can help agencies uncover illuminating data points that clients can’t provide up front.

The key to informing a strategy that will achieve a client’s marketing goals is to identify which specific types of data you’re looking for before diving into the analysis. Below, experts from Forbes Agency Council share 11 of the most valuable pieces of information you can glean by analysing your clients’ Google Analytics.

1. What Attracts Versus Repels

As communications experts, we love reviewing Google Analytics to better understand how customers are engaging with a brand and what’s attracting them versus repelling them. This establishes information that allows us to develop more compelling content strategies. You’re able to see the level of leads coming from media relations and placed articles, which is a strong indicator of campaign success. – Kathleen Lucente, Red Fan Communications

2. The Client’s Audience

At the end of the day, the most valuable element of successful marketing is understanding the consumer. Google Analytics can provide some insight into a client’s audience. Combining this with other data sets and marrying the research with strategic analysis can inform an insight-driven marketing strategy. This can inspire consumer targeting, creative, media and more. – Marc Becker, The Tangent Agency

3. ROI On Marketing Investments

No matter what, you want to make sure that you are getting ROI on any marketing investment. Even if your Google Analytics are telling a positive story, if you aren’t getting actual ROI, there is data that either is not accurate or needs to be looked at holistically. There should always be a system of checks and balances, and all touch points should be telling the same story. – Jessica Hawthorne-Castro, Hawthorne LLC

4. Return On Ad Spend Performance 

The most important piece of data you can glean from Google Analytics is the ROAS performance of your clients’ media buying across the various websites they are advertising on. By tracking where the users are coming from and tracking their activity on your clients’ sites, you can determine their ROAS. You can then shift media investment to the top-performing websites. – Dennis Cook, Gamut. Smart Media from Cox.

5. The Source Of Relevant Traffic

Analysing their clients’ Google Analytics allows agencies to see where relevant traffic is coming from, identify trends and target opportunities. Additionally, optimizing your campaigns based on the data feedback will lead to higher conversion rates. – Jordan Edelson, Appetizer Mobile LLC

6. Time On Page

Time on page is the most important Google Analytics statistic. Once you get traffic to your site, do they stay? What content do they consume? How much mindshare do they give you? What pages are sticky and not transactional? Time on page tells you what prospects value and where they give your ideas credence. Know this, and you’ll know your audience. – Randy Shattuck, The Shattuck Group

7. Where Viewers Leave The Website

The pages where viewers are leaving the client’s website at abnormally high rates is where to focus. By finding out what pages are causing website viewers to drop off the most, clients can analyze these pages and make necessary adjustments to better grab the attention of future visitors. – Stefan Pollack, The Pollack Group

8. Behaviour Flow

Behaviour Flow is still my favourite feature offered by Google Analytics. Studying the flow of the visitors and the path they take while interacting with a website helps business owners understand what a page means to the customer. This information helps business owners understand how to prioritize and optimize pages to offer visitors a better user experience. – Ahmad Kareh, Twistlab Marketing

9. Goal Conversion Data

Google Analytics can be overwhelming, so a great place to start is by looking at a client’s goal conversions (the number of visitors that took the action your client intended for them to take). This one area can give quick insight into how and why a website was built, as well as whether or not the site is performing the way it’s meant to. If goals have not yet been set up, this is a great opportunity to start a conversation with your client about short- and long-term objectives. – Carey Kirkpatrick, CKP

10. The Most Popular Content

Simply looking at your website’s most popular content can tell you if that website really serves your target customer. All too often, content serves another purpose or user. My agency’s example is that the bio I wrote for our vice president was the most popular piece of content, which proved that web visitors came to copy that bio rather than to hire our agency. – Jim Caruso, M1PR, Inc. d/b/a MediaFirst PR – Atlanta

11. Device Usage

One often overlooked piece of data in Google Analytics is device usage. All clients basically have two websites: a desktop site and a mobile site. Understanding what visitors are doing on both sites is critical, especially when it comes to advertising and landing pages. – T. Maxwell, eMaximize

Sourced from Forbes

By

The most interesting part of a study from Sidecar not shared on Monday in Search & Performance Marketing Daily points to the percentage that marketers rely on data vs. instinct to make marketing and advertising decisions.

Sidecar surveyed 146 marketing professionals in the retail industry. The majority of respondents were based in the U.S., with the remainder in Canada. All reported that they contribute to ecommerce marketing efforts at their company. The study was fielded between September and October of 2020.

When marketers were asked whether their team makes decisions based on data versus experience and instinct, the balanced response was 50% data and 50% instinct, with 24% of respondents reporting this way.

From here the findings become quite unbalanced. Only 1% of participants in the survey said they base their decisions on 100% instinct and zero percent data, and 1% base their decisions on 90% instinct and 10% data. Some 7% base their decisions on 80% instinct and 20% data, and 18% base their decisions on 70% instinct and 30% data.

When flipping the percentages, the findings are a bit surprising. It turns out that none base their decisions on 0% instinct and 100% data. It does get better, however. Only 4% base their decisions on 10% instinct and 90% data, while 10% base their decisions on 10% instinct and 80% data, and 16% base their decisions on 30% instinct and 70% data.

Some 62% of ecommerce marketing teams are making half or more of their decisions based on instinct rather than data, indicating significant headroom to become more data-driven.

In this new year marketers need to think differently to drive growth and connect with consumers. Thinking differently has important implications for marketers in terms of hiring.

Automation will find a home in more companies this year, from ad testing to keyword analysis, and audience segments and performance trend analysis. Among C-level executives, 82% want to automate bid adjustments, while 59% want to automate ad testing, 53% want to automate retargeting, 47% want to automate bid analysis, and 41% want to automate performance trend analysis.

What will marketing teams look like in 2021 as they reach consumers? Ecommerce marketers, for example, plan to grow their internal and extended teams. Some 66% plan to hire vendors and 67% plan to hire in-house talent.

Enterprise and small businesses plan to hire marketers with affiliate marketing and SEO experience.

Enterprise companies plan to hire content marketing to round out the top three, whereas small companies plan to hire those with video production experience.

Midsize companies are looking for specialists with experience in social media, video production and data analytics.

By

Sourced from MediaPost

This is probably the most common question I got asked beside “How did you land your job in Data Science/ Data Analytics?” I will write another blog on my job hunting journey, so this will focus on how to get the industry exposure without that gig yet.

I gave a talk on this topic before at DIPD @ UCLAthe student organization dedicated to increasing diversity and inclusion in the fields of Product and Data that I co-founded. However, I aim to expand this topic and make it accessible to a broader audience.

And there it goes, I hope this post will potentially inspire more and more data enthusiasts to start their own blogs.

This may be a tough time for many of us, but it’s also a prime time to turbocharge and level up your skill sets in data science and analytics. If your employment got impacted at this time, treat the unfortunate as a great opportunity to take a break, reflect and kickstart your personal project — things that are luxurious when time does not allow.

“When one door closes, another opens” — Alexander Graham Bell

Hardship does not determine who you are, it’s your attitude and perseverance that define your values. Let’s get right into it!

Where to start?

Photo by Carl Heyerdahl via Unsplash

Start small and scale up

Before we start any project, first narrowing down your interests. This is your personal project so you will have full autonomy over it. Find something that makes you tick and gets you motivated to devote your time!

There will be a lot of challenges along the way that may discourage or sidetrack you from accomplishing the project, the thing that keeps you going should be the analysis topic that strongly aligns with your interest. It does not have to be something out of the world. Ask yourself what is important to you and why should we care about it.

When I first started, I knew that I wholeheartedly care about mental health and the ways to gain more mindfulness. So I dug deeper into analyzing the top 6 guided meditation apps to understand which one will be most suitable for my preferences.

Getting inspirations

Photo by Road Trip with Raj via Unsplash

Read, read, and read!

One of the most important key factors that I learned through my research assistant position at CRESST UCLA is to balance the workload between analysis and literature review. What this means is that we need to find what has been done in the past and figure out which additions or unique aspects you can contribute on top of the findings. My reading sources vary from Medium, Analytic Vidhya, statistics books to any relevant sources that I can find on the internet.

Take my Subtle Couple Traits analysis for example. There has been some work done in the space of music taste analysis via Spotify API, but no one has really delved into movies yet. So I took this chance and discovered the intersection of our couple’s cult favorites for music and movies.

Finding the right toolbox

Photo by Giang Nguyen via MinfulR on Medium

Now you get to this step where you need to figure out which data to collect and find the right tools for the job. This part has always resonated intrinsically with my industry experience as a data analyst. It’s the most challenging and time-consuming part indeed.

My best tip for this stage of analysis is to ask a lot of practical questions and come up with some hypotheses that you need to answer or justify through data. We have to also be mindful of the feasibility of the project, otherwise, you can be more flexible in terms of tweaking your approach towards a more doable one.

Note that you can use the programming language that you are most comfortable with 🙂 I believe that either Python or R has its own advantages and great supporting data packages.

An example from my past project can crystalize this strategy. I was curious about the non-pharmaceutical factors that correlate to the suppression of COVID-19 so I listed out all of the variables I can think of such as weather, PPEs, ICU beds, quarantines, etc. then I began massive research on the open-source data sets.

“All models are wrong, but some are useful” — George Box

Since I did not have a background in public health, building predictive models for this type of pandemic data was a huge challenge. I first started with some models I’m familiar with such as random forest or Bayesian ridge regression. However, I discovered that pandemic typically follows the trend of a logistic curve in which the cases grow exponentially over a period of time until it hits the inflection point and levels out. This refers to the compartmental models in epidemiology. It took me almost 2 weeks to learn and apply this model to my analysis but the result was extremely mesmerizing. And I eventually wrote a blog about it.

The process

If you are working in the Data Science/Analytics field, this is not new to you — “80% of a data scientist’s time consists of preparing (simply finding, cleansing, and organizing data), leaving only 20% to build models and perform analysis.”

Photo by Impulse Creative

The process of cleaning data may be cumbersome, but when you get it right, your analysis will be more valuable and significant. Here’s the typical process I take for my analysis workflow:

1) Collecting Data

2) Cleaning Data

Many more…

3) Project-based techniques

  • (NLP) Sentimental analysis, POS, topic modeling, BERT, etc.
  • (Predictions) Classification/Regression model
  • (Recommendation System) Collaborative Filtering, etc.

Many more…

4) Write up insights and recommendations

Connecting the dots

This is the most important part of the analysis. How do we connect the analysis insights into a real-life context and making actionable recommendations? Regardless of your project’s focus, whether it’s about machine learning, deep learning or analytics, what problem is your analysis/model trying to solve?

Photo by Quickmeme

Imagining that we build a highly complex model to predict how many Medium readers will clap for your blog. Okay, so how’s this important?

Link it to potential impacts! If your post receives more endorsement from claps, it may get curated and featured more often on Medium platform. And if more paying Medium readers find your blog, you can probably earn more money through the Medium Partner Program. Now that’s an impact!

However, it’s not always about profit-driven impact, it could be social, health, or even environmental impact. This is just one example of how you can make the connections between technical concepts with real-world implementation.

Roadblocks

You may hit a wall at some points during the journey. My best piece of advice is to proactively seek help!

Besides from reaching out to friends, colleagues, or mentors to ask for advice, I often found it helpful to search or post questions on online Q&A platforms like Stack Overflow, StackExchange, Github, Quora, Medium, you name it! While seeking for solutions, be patient and creative. If the online solutions have not yet solved your problems, try to think of another way to customize the solution for the characteristics of your data or the version of the code.

The art of writing is rewriting.

When I first published my first data blog to Medium, I found myself re-visiting my post and fixing some sentences or wording here and there. Don’t be discouraged if you notice some typos or grammar mistakes after releasing it, you can always go back and edit!

Since it is our personal project, there’s no obligation on whether you must finish it. Hence, prioritization and disciplines play a crucial role throughout the journey. Set a clear goal for your project and lay out a timeline to achieve it. At the same time, don’t spread yourself too thin since it may cause you to lose interest.

Understand your timeline and capacity! I often push my personal project in a sprint of 2 to 4 weeks to finish during break or the weekends. In order to organize your sprint and track your progress, you can refer to some Agile framework that can be found through collaboration software like Trello or Asana. As long as you make progress even the smallest one, your success shall flourish some day. So keep going and don’t give up!

Closing Remarks

The first step is always the hardest. If you don’t think that the project is ready yet, give yourself some time to fine-tune and share it!

Nothing will be perfect at first. But by shipping it to the audiences, you would know what to improve for later projects — I adopted this principle wholeheartedly from product management perspectives.

I used to be not good at communicating my thoughts structurally and clearly (which I’m still trying to improve), but by pushing myself out of the comfort zone, I have gone extra miles from where I was. I hope this will, to some degree, inspire you to start your first data blog. Believe in yourself, be brave and reach out to me or anyone in your network if you need help along the way!

“Faith is taking the first step even when you don’t see the whole staircase.” — Martin Luther King

Photo by Glen McCallum via Unsplash

By Giang Nguyen

Sourced from towards data science

By

The video-call provider has apologised for sending data to Facebook without users’ permission, showing that we must be vigilant about the tech we use.

A couple of months ago, Zoom was a dull, if successful, videoconferencing app that not many people knew about. Now, it is a household name and an integral part of many of our quarantined lives. We conduct business meetings on it; we chat to our mates on it; some people even have sex parties on it.

Yet there are growing concerns over what it does with users’ data. You may think you are working from the privacy of your own home, but the software is probably sharing a lot more information about you than you realise. Zoom has an attention-tracking feature, for example, which notifies the host of some video calls if participants click away to look at something else. The company has actively promoted this feature to educators, explaining it’s a good way to monitor which of your students is slacking off.

In any article about privacy violations, it is pretty much a given that Facebook will be mentioned. This is no exception. Recent analysis by Vice found that Zoom’s iOS app was sending analytics data to Facebook, even when the user did not have a Facebook account and even though this was not addressed in Zoom’s privacy policy. This data included things such as the user’s location and the device’s advertiser identifier information, a unique ID that lets companies send you targeted ads. On Friday, Zoom issued a statement saying “whoops!’” and announcing it had updated its software to stop sending iOS data to Facebook.

I am not saying that you should boycott Zoom and communicate via carrier pigeon. However, as we are forced to live even more of our lives online, let’s not stop holding tech companies to account. Let’s not stop trying to safeguard our right to privacy. Our civil liberties are most fragile during times of crisis. Governments around the world are already using this pandemic to bolster the surveillance state. If we don’t stay vigilant, our privacy will be lost before you can say “Zoom”.

Since you’re here…

… we’re asking readers like you to make a contribution in support of our open, independent journalism. In these frightening and uncertain times, the expertise, scientific knowledge and careful judgment in our reporting has never been so vital. No matter how unpredictable the future feels, we will remain with you, delivering high quality news so we can all make critical decisions about our lives, health and security. Together we can find a way through this.

We believe every one of us deserves equal access to accurate news and calm explanation. So, unlike many others, we made a different choice: to keep Guardian journalism open for all, regardless of where they live or what they can afford to pay. This would not be possible without the generosity of readers, who now support our work from 180 countries around the world.

We have upheld our editorial independence in the face of the disintegration of traditional media – with social platforms giving rise to misinformation, the seemingly unstoppable rise of big tech and independent voices being squashed by commercial ownership. The Guardian’s independence means we can set our own agenda and voice our own opinions. Our journalism is free from commercial and political bias – never influenced by billionaire owners or shareholders. This makes us different. It means we can challenge the powerful without fear and give a voice to those less heard.

Your financial support has meant we can keep investigating, disentangling and interrogating. It has protected our independence, which has never been so critical. We are so grateful.

We need your support so we can keep delivering quality journalism that’s open and independent. And that is here for the long term. Every reader contribution, however big or small, is so valuable. Support the Guardian from as little as €1 – and it only takes a minute. Thank you.

Feature Image Credit: ‘Let’s not stop holding tech companies such as Zoom to account.’ Photograph: Christian Sinibaldi/The Guardian

By

Sourced from The Guardian

By

Developing a holistic data strategy

Enterprises of all sizes, all over the world, have now recognized that data is an integral part of their business that cannot be ignored. While each enterprise may be at a different stage of their personal data journey – be it reducing operational costs or pursuing more sophisticated end goals, such as enhancing the customer experience – there is simply no turning back from this path.

In fact, businesses are at the stage where data has the power to define and drive their organisations overall strategy. The findings from a recent study by Infosys revealed that more than eighty-five percent of organisations globally have an enterprise-wide data analytics strategy already in place.

This high percentage is not surprising. However, the story does not end with just having a strategy. There are numerous other angles that enterprises must consider and act on before we can deem a data journey as successful.

Developing a data strategy

First, enterprises need a calculated strategy which covers multiple facets. Second, the real life implementation of the strategy must be seamlessly carried out – and this is where the challenge lies for all enterprises.

Consider having to create a comprehensive and effective strategy for your company. Data strategy is no longer about simply identifying key metrics and KPIs, developing management roles or creating operational reports, or working on technology upgrades. Rather, its reach extends to pretty much all corners of the business.

In short, data strategy is so tightly integrated with business today, that it is in the driver’s seat, which is a momentous shift from more traditional approaches of the past.

What are the characteristics of a good, strong data strategy?

Creating a good, strong data strategy begins with ensuring complete alignment with the organisation strategy. The data strategy must be closely aligned to the organisational goal, be it around driving growth or increasing profitability or managing risk or transforming business models.

Not only that, but the data strategy must be nimble and flexible, allowing periodic reviews and updates to keep pace with wider changes in the business and market. The data strategy should be able to drive innovation, creating a faster, better and more scalable approach.

A strong data strategy must be built in a bi-directional manner so that it can enable tracking of current performance using business intelligence to provide helpful pointers for the future. This approach is only possible if organisations choose to adopt a multi-pronged data strategy that encompasses people, technology, governance, security and compliance. Importantly, organisations must also choose to adopt an appropriate operating model.

Taking a holistic approach to data

A holistic approach includes developing a defined vision, having a clear structure around the team and factoring in the current skill set of the team. This is in addition to considering what the enterprise can reasonably anticipate in the future and identifying mechanisms to successfully drive the change across the organisation.

The technology component involves having a distinct vision, assessing the existing solution landscape, all the while being cognizant of the latest technological trends and arriving at a path that fits well with overall organisational goals and the technology vision.

Governance, security, and compliance are other critical aspects of a good data strategy. Integrity, hygiene and ownership of data, plus relevant analytics on the data to determine the Return On Investment on data strategy, are all essential steps which cannot be forgotten. We cannot overstate the importance of security.

Adherence to compliance has assumed significance with various regulations in play all over the world, such as GDPR in Europe and new data privacy laws in California and Brazil for example.

In essence, the data strategy must define a value framework and have a reliable mechanism to track the returns to justify the investments made. About fifty percent of respondents to our survey agreed that having a clear strategy chalked out in advance is essential to ensuring an execution that is effective in practice and goes off without any hiccups.

Identifying the best strategy is essentially pointless if the execution falters

Many obstacles have the power to prevent the flawless execution of a data strategy. Copious challenges in the technology arena can arise in various forms, for example: having the knowledge to choose the right analytics tools, lack of availability of people with the right skill set, upskilling, reskilling and training the workforce with the necessary skills for the world of tomorrow and so on. Most of the challenges articulated by respondents to the Infosys survey arose in the execution phase of a data strategy.

While these challenges may appear daunting in the first instance, they can be addressed with careful planning and preparation. Being prepared and equipped for multiple geographies, multiple locations, multiple vendors, talent acquisition and good quality training are just some of the numerous possible ways companies can begin working towards smooth execution of their digital strategy.

Feature Image Credit: Image credit: Pixabay

By

Gaurav Bhandari, AVP and Head of Data & Analytics Consulting at Infosys.

Sourced from techradar.pro

Sourced from DIGIDAY

While you can’t plan for uncertainty, you can prepare for it. The Advertising Association is encouraging the industry to plan for Brexit as the risks of the UK leaving the EU without a deal on 31 October 2019 are high.

In its remit of representing the interests of the UK advertising industry, the Advertising Association has brought together key pieces of information to ensure businesses have contingencies in place to continue receiving personal data lawfully in the event of a no-deal Brexit. This is intended to provide guidance, and does not replace legal advice.

The UK’s data protection regime is currently governed by the EU’s General Data Protection Regulations (GDPR) and the UK’s Data Protection Act 2018 (DPA 2018). If your organisation receives personal data from the EEA you will still need to abide by both GDPR and the DPA 2018 even after Brexit.

Assessing data adequacy
As the UK is currently a member of the EU, there are no restrictions on the flow of personal data and other EEA Member States. Article 45 of the GDPR states that the European Commission needs to assess the relevant country’s laws to determine whether they are essentially equivalent or “adequate” to that of EU ones.

The UK has announced that it will allow the flow of personal data to the EEA regardless of a deal being in place and will recognise existing European Commission data adequacy decisions. However, the EU has not yet made a similar commitment towards the UK. This is because on leaving the EU, the UK will become a ‘third country’. And while the UK remains an EU member, the European Commission will not conduct this assessment. Unfortunately, this means if we leave the EU without a deal we will not have a data adequacy decision in place to facilitate the free flow of personal data from the EEA.

Standard Contractual Clauses
In the absence of an adequacy decision, GDPR states that personal data can be transferred to a third country or an international organisation if there are appropriate safeguards. There are a number of recognized safeguards, but most appropriate to businesses are the implementation of Standard Contractual Clauses (SCCs).

SCCs are a standard set of contractual terms and conditions for the transfer of personal data which both the data exporter and the data importer enter into. They include contractual obligations that help to protect personal data when it leaves the EEA and ensure compliance with GDPR. SCCs only relate to the transfer of personal data, so they can be incorporated into a wider contract that covers other business terms. One of the key benefits of using these SCCs is that they are approved by the European Commission.

Binding corporate rules
If you are a multinational operating in the UK and in one or more EEA country, then Binding Corporate Rules are required to transfer personal data between the different parts of the Group located in the UK and the EEA.

US Privacy Shield
If you send data to a US Privacy Shield organisation, the Privacy Shield participant will need to update their public commitment to specifically reference the UK, in addition to the EU. There is further information on the US government’s Privacy Shield website. In addition, the ICO has published guidance for organisations about international data transfers.

Data Protection Lead Authority
If the ICO is your lead Data Protection Authority, you may need to review your operations to assess whether you can still have a lead authority and benefit from the one-stop-shop following Brexit.

Appointing a data representative.
If you are a data controller or processor that is subject to GDPR but not established in the EEA – as will be the case when the UK leaves the EU – you have an obligation to designate a data representative based in the EEA. This representative will be the go-to person to deal with individuals and DPAs in the EEA. The UK plans to oblige non-UK controllers who are subject to the UK data protection framework to appoint representatives in the UK if they are processing UK data on a large scale.

It’s important to regularly check the GOV.UK website for updates. The ICO has a page dedicated to Brexit that covers the implications for data protection and data transfers in more detail and its SCC tool provides template contracts. If you need more information about your obligations and what you need to do to comply, we recommend seeking legal advice.

For more information on matters relating to Brexit, visit the Advertising Association website: https://www.adassoc.org.uk/policy-areas-category/Brexit/

Sourced from DIGIDAY

By Paul Matthews

In 2018, the world has been shaken by the usage of big data: the Cambridge Analytica scandal, which was related to the allegedly illegal buying and selling process of data points and data-related pieces from the British company, has put data science into the spotlight of the “mainstream business” world. After this scandal, in fact, data has surpassed oil as the most valuable asset on Earth. Let’s analyse why and, most importantly, how this has happened.

Data Points: A Commercially Powerful Numerical Value

For “data point”, we intend a numerical value which, when associated with a specific entity (i.e. a person, a company), combines preferences, comments and tastes (from a numerical perspective) in order for a software to automatically elaborate them. The power of data points stands in the fact that, when properly analysed, they could give thorough insights on a particular user’s preference on a specific topic. The “exploitation” of Facebook searches on the Brexit topic, for example, was elaborated using data points to provide highly tailored ads to the people who were either searching for “leave the UK” and related keywords. Although this may sound slightly political, it was actually confirmed by Cambridge Analytica itself last year after they (and Facebook) were fined for over $2 billion for buying and selling private pieces of information (data points).

Data Science: An Enterprise Niche Sector Going Mainstream

The possibility of creating tailored ads based on numerical values has intrigued business owners worldwide to the point in which they decided to open data science-related divisions in companies which weren’t exactly at an “enterprise” level. Data elaboration, acquisition, science and Python development professional figures have been recruited in small and medium companies worldwide massively, in the past 7 months. Despite a specific GDPR section strictly regulating data acquisition and processing, data scientists have definitely “gone mainstream” in the recent past.

From fintech to eCommerce, to pure lead generation, the usage of data science has become a constant in 2019.

Some Business Sectors Have Been Getting More Results Than Others…

As mentioned above, data processing and science have been used by a variety of businesses in the past months. Fintech and real estate have been the most successful ones, in terms of lead generation tailored onto data. Sectors like bridging loans, development finance and similar have seen a net 35% increase in organic investment in terms of hiring Python developers who were able to process such delicate data to prepare targeted, tailored and highly convertible ads for social media channels. Lead generation has become very dependant on data in the recent past.

To Conclude

The usage of data in 2019 has definitely become a mainstream procedure. In the nearest future, we can safely say that GDPR rules will become even more strict: with more specific regulations on the acquisition and storing, data is still far away from being fully regulated.

By Paul Matthews

Paul Matthews is a Manchester-based business and tech writer who writes in order to better inform business owners on how to run a successful business. You can usually find him at the local library or browsing Forbes’ latest pieces. Paul is currently consulting a bridging loans company in Manchester.