Conversational AI drives better customer experiences - Build What's Next
Blog

Conversational AI drives better customer experiences

3364

Of your peers have already read this article.

3:30 Minutes

The most insightful time you'll spend today!

Find out what's new in Google Contact Center AI solution, including Agent Assist, Custom Voice, and the Speech-to-Text On-Prem, the first of Google Cloud's hybrid AI offerings.

Conversational AI is opening up a new world of possibilities in areas like customer experience, user engagement, and access to content. 

In Cloud AI, we’ve taken Google’s groundbreaking machine learning models in speech and natural language processing and applied them to the contact center space, radically improving the customer experience while also driving down operational costs. It’s shifting the focus of contact centers from the backroom, out of sight (and mind) to the boardroom and the strategic heart of the business.

Our solution, called Contact Center AI (CCAI), is an accelerator of digital transformation as organizations all over the world figure out how to support their customers during these challenging times. One customer that’s redefined the possibilities of AI-powered conversation using CCAI, is Verizon.

Whether through voice calls or chat, Verizon customers will no longer need to go through menu prompts or option trees; they simply say or type their question, and CCAI’s natural-language recognition feature finds the best way to assist them. No stilted speech or robot-like commands.

“These customer service enhancements, powered by the Verizon collaboration with Google Cloud, offer a faster and more personalized digital experience for our customers while empowering our customer support agents to provide a higher level of service,” said Shankar Arumugavelu, SVP & CIO at Verizon. 

CCAI is also driving cost savings without cutting corners on customer service. In the past, to improve customer satisfaction (CSAT), you had to hire more agents, increasing operating costs. Conversely, if you reduced the number of agents, your CSAT scores went down. Contact Center AI (CCAI) lets you have both.  

CCAI does this by expediting customer requests using virtual agents, assisting live agents, and providing insights on customer interactions to drive improvements back into the system. The result is a better experience for your customers—they get answers faster—and you get lower operating costs. It allows human agents to focus on higher value tasks in their customer interactions.

New products and features add richer customer experience 

We are continuing to invest in CCAI and today we are excited to announce the following new capabilities:

Dialogflow CX is the latest version of Dialogflow, a development suite used by over a million developers for building natural, rich conversational experiences into mobile and web applications, smart devices, bots, interactive voice response systems, popular messaging platforms and more. This new version of Dialogflow is optimized for large contact centers that deal with complex (multi-turn) conversations and it is truly omnichannel – you build it once and deploy it everywhere – in your contact centers and digital channels. Dialogflow CX features a new visual builder to create, build and manage virtual agents. It’s available now, in beta.

Dialogflow CX brings conversation state management to a whole new level

Lukasz Rewerenda
Principal Solutions Architect, Randstad (Netherlands)

Agent Assist for Chat is a new module for Agent Assist that provides agents with continuous support over “chat” in addition to voice calls, by identifying intent and providing real-time, step-by-step assistance. Agent Assist enables agents to be more agile and efficient and spend more time on difficult conversations, giving both the customer and the agent a better experience. It transcribes calls in real time, identifies customer intent, provides real-time, step by step assistance (recommended articles, workflows, etc.), and automates call dispositions.

For customers in regulated industries, Agent Assist can remove the risk of agents providing inaccurate information (which can happen due to high agent turnover and limited training). Agent Assist can also surface the latest discount information, deals and special offers, which can be hard for agents to keep track of as this information changes frequently. 

Custom Voice, available in beta, is a new capability for CCAI and our Text-to-Speech API that lets you create a unique voice to represent your brand across all your customer touchpoints, instead of using a common voice shared with other organizations. By taking advantage of the custom Text-to-Speech model created with Custom Voice, you can define and choose the voice profile that suits your business and adjust to changes without scheduling studio time with voice actors to record new phrases.

At Google Cloud, we are committed to ensuring that our products and features are in alignment with our AI Principles. If you are interested in Custom Voice, there is a review process to ensure that your use case is aligned with our AI principles. Learn more about Custom Voice and how to sign up here

Bring Google’s groundbreaking speech models on-prem

For many customers, an all-in approach to public cloud is not an option, which is why we’re extending our AI capabilities to run on-prem. Last week we announced Speech-to-Text On-Prem, the first of our hybrid AI offerings, now generally available. Speech-to-Text On-Prem gives you full control over speech data and since it runs in your own data center, it’s easy to comply with your data residency and compliance requirements. At the same time, Speech-to-Text On-Prem uses state-of-the-art speech models from Google researchers that are more accurate, smaller, and require less computing resources to run than existing solutions.

Take a look at how businesses saved costs and improved customer experience with CCAI in the new commissioned study conducted by Forrester Consulting: “New Technology: The Projected Total Economic Impact™ Of Google Cloud Contact Center AI.” To learn more about CCAI and conversational AI, see how our customers are using the solution and check out these sessions at Next OnAir:

How-to

How TeamSnap Improved Return on Ad Spend Significantly

5605

Of your peers have already read this article.

3:30 Minutes

The most insightful time you'll spend today!

Using a combination of Google Analytics 360, Google BigQuery, and Tableau, TeamSnap’s marketing team reallocated $300,000 in underperforming ad spend, achieving a 200% ROI in just two days.

Anyone who has ever coached or played on a sports team, or had a child involved in sports, knows how difficult scheduling and logistics can be. From game and practice schedules, to uniforms and who’s bringing the snacks, it can be a lot for coaches, administrators, parents, and players to manage.

It’s no wonder that TeamSnap, a sports team, club, and tournament management app, has exploded in popularity worldwide. By syncing events to everyone’s personal calendars and providing messaging and payment tracking, TeamSnap makes communication and organization easy.

Achieves 200% ROI in 2 days by reallocating $300,000 in underperforming ad spend. Improves customer engagement, generating $4 million in additional customer value each year.

TeamSnap markets its app to coaches, players, and clubs via targeted YouTube ads. It also uses Google AdWords and DoubleClick to advertise on search results and run programmatic campaigns. These methods have been highly effective, helping TeamSnap grow to millions of users worldwide and become one of the most popular apps in the iOS app store.

As its business and data grew, TeamSnap was challenged to track ROI and measure the customer journey across channels and devices over time. The company’s marketing budget grew quickly, making it even more important to spend wisely. With data in Google Analytics 360DoubleClick Campaign Manager, Google AdWords, and Salesforce, TeamSnap needed a way to link and correlate those data sources in a scalable, timely, and cost-effective way to understand the true impact of its digital marketing across websites and mobile apps.

To avoid the painstaking manual process of pulling data from multiple sources, TeamSnap began using Google Analytics 360, which integrates with Google BigQuery, to provide a fully managed big data analysis service. TeamSnap analyzes the data using Tableau, which connects directly to Google BigQuery for fast analytics and helps the company share and collaborate on that information with self-service ease.

“Before Google Analytics 360, Google BigQuery, and Tableau, tracking our return on ad spend was difficult because we had so much data. We don’t have that problem anymore because we’ve moved to real-time reporting. We find additional revenue growth opportunities almost daily.”
-Ken McDonald, Chief Growth Officer, TeamSnap

The combination allows TeamSnap to easily track the activity of millions of users with self-service ease, without worrying about the scalability or availability of the big data platform.

“Before Google Analytics 360, Google BigQuery, and Tableau, tracking our return on ad spend was difficult because we had so much data,” says Ken McDonald, Chief Growth Officer at TeamSnap. “We didn’t always have insights to make the best choices. We don’t have that problem anymore because we’ve moved to real-time reporting. We find additional revenue growth opportunities almost daily.”

Making Ad Dollars Work Harder

TeamSnap now automatically imports unsampled Google Analytics 360 logs into the Google BigQuery data warehouse. To import data from other sources such as Google AdWords, DoubleClick, and YouTube, TeamSnap uses Google BigQuery Data Transfer Service. With all relevant data consolidated in Google BigQuery, TeamSnap can use Tableau to perform advanced analytics on its digital marketing, executing ad-hoc analyses in seconds, while eliminating data sampling issues, to improve accuracy. These analyses can also be reused and shared with internal and external stakeholders via Tableau Online, promoting governed reuse and consistency.

“Integration between Google Analytics 360 and Google BigQuery is seamless, giving us much more confidence in our A/B testing. We’re constantly finding new and interesting ways to use our digital marketing data. Often, making a simple change can increase revenue by hundreds of thousands of dollars a year.”
-Ken McDonald, Chief Growth Officer, TeamSnap

“Using Google Analytics 360 and Google BigQuery with Tableau to track our return on ad spend is ideal,” says Ken. “It’s easy to use SQL to query the data or explore it with drag-and-drop ease.”

With Google BigQuery, Ken and his team can bring all the data from the TeamSnap billing systems, internal CRM, and other Google services into one straightforward dataset that everyone uses. With Tableau, users are able to perform self-service analytics on this data and provision analyses via shared dashboards that communicate the same consistent truth across the company. These dashboards provide a single view of the business to discover new patterns and questions worth analyzing.

All of this results in enormous time savings because no one is re-inventing the wheel. “Using these tools, we immediately reallocated $300,000 of ad spend that was performing poorly, generating 200% ROI in the first two days,” says Ken.

Ken now spends his time analyzing data instead of trying to pull it all together, identifying pockets of inefficient spend in real time and reallocating those marketing dollars toward better performing campaigns.

“Before, we could only focus on the largest campaign-level datasets because it was so time consuming to pull the data,” he says. “With Google BigQuery and Tableau, we can examine our advertising ROI much more granularly and reallocate more than $10 million in ad spend annually to grow the company faster and more efficiently.”

More Effective A/B Testing

To make sure it is delivering the best customer experiences, TeamSnap uses Google Optimize to run A/B tests on its website. It uses Google BigQuery and Tableau to verify and supplement these findings by measuring longer-term customer behavior across devices, spanning both web and mobile apps.

By pulling in data from Google Optimize, Google Analytics 360, Salesforce, and in-house billing and CRM systems, and understanding it with Tableau, TeamSnap has increased the accuracy and effectiveness of its A/B testing, gaining a more complete picture of customer onboarding and activity. In some cases, it found that short-term indicators it previously trusted were actually poor predictors of long-term behavior.

“Integration between Google Analytics 360 and Google BigQuery is seamless, giving us much more confidence in our A/B testing,” says Ken. “We’re constantly finding new and interesting ways to use our digital marketing data. Often, making a simple change can increase revenue by hundreds of thousands of dollars a year.”

Improving Product Quality

TeamSnap also uses Google BigQuery and Tableau to improve its own product, tracking customer activity at such a granular level that usability and functionality issues can be exposed and addressed faster. It’s also increasing customer engagement by verifying that potential customers are coming in through the right onboarding path—for example, a coach versus a player, or a consumer versus a club or other sports business. Using A/B testing to make sure customers are routed to the appropriate flow, TeamSnap drove $4 million in additional customer value each year.

“We initially chose Google BigQuery and Tableau to help with marketing, but we realized quickly that they could help us on the product side as well,” says Ken. “Most of the testing we do is about making things better and easier for our customers, and we’re accelerating that process with Google BigQuery and Tableau.”

Whitepaper

ESG Report: Economic Advantages of Google BigQuery OnDemand Serverless Analytics

DOWNLOAD WHITEPAPER

5890

Of your peers have already downloaded this article

10:30 Minutes

The most insightful time you'll spend today!

Traditional big data solutions require a significant upfront investment, ongoing maintenance of hardware and software, and manual provisioning to match compute and storage resources with demand. Google’s serverless analytics warehouse does away with all that extra work, helping businesses focus on what matters most: getting value from their data.

Enterprise Strategy Group (ESG), which examined the economic value propositions of Google BigQuery and alternative big data solutions, reports that BigQuery enables organizations to:

  • Save up to 88 percent on data warehousing over a three-year period
  • Achieve a faster time to value by getting up and running quickly
  • Empower more employees to become citizen data scientists

Eliminate maintenance tasks so teams can spend more time gaining insights

Download the complete report to learn more.

5135

Of your peers have already watched this video.

44:51 Minutes

The most insightful time you'll spend today!

Case Study

The Strange Phenomenon AI Revealed at Ride-Hailing Company Go-Jek

Go-Jek, Indonesia’s first billion-dollar startup, has seen an incredible amount of growth in both users and data over the past two years. Many of the ride-hailing company’s services are backed by machine learning models hosted on Google Cloud Platform. Models range from driver allocation, to dynamic surge pricing, to food recommendation, and process millions of bookings every day, leading to substantial increases in revenue and customer retention.

By embracing Google Cloud, Go-Jek has overcome many of the technical challenges brought on by its rapid growth. BigQuery has become the cornerstone of their data foundation, scaling seamlessly to meet their immense data storage and processing needs.

Using Pub/Sub as an event stream and Dataflow for unified batch and stream processing has prevented inconsistencies in production data, while simultaneously reducing costs through intelligent resource allocation.

Together, these technologies allow Go-Jek to react immediately to real world events, whether by retraining models with ML Engine, or refreshing data in a low latency data store like BigTable.

Find out how Go-Jek leverages Google Cloud and other lessons they have learned scaling machine learning.

Case Study

Home Depot’s Interconnected Retail Experience by Virtue of Google Cloud Migration for SAP Applications

8247

Of your peers have already read this article.

1:30 Minutes

The most insightful time you'll spend today!

In 2017, home improvement brand Home Depot began its Google Cloud migration for SAP applications. The migration strategy had a multifold impact on the brand's CX and retail shopping experiences with better flexibility, speed and scale.

With nearly 2,300 stores, The Home Depot is the world’s largest home-improvement chain — a brand that professional contractors and DIYers alike have come to depend on. The home improvement industry continues to experience unprecedented demand and dramatic increases in online ordering accompanied by expanding consumer expectations for things like curbside pickup and same day delivery. The Home Depot’s decision to migrate to cloud-based infrastructure, including the migration of the company’s SAP applications on Google Cloud which began in 2017, has set it up for success in an increasingly digital world, and helped the company adapt to changing market conditions quickly. 

Interconnected retail at scale

Building on a strong customer-first philosophy, The Home Depot aims to create what it calls interconnected retail—allowing customers to shop however, whenever, and wherever they want. “So many companies are focused on omni-channel retail,” explains Sam Moses, Vice President of Corporate Systems. “At The Home Depot, we wanted to take it to the next level. Interconnected retail puts the customer at the center of everything and enables them to shop in store, online, or both. Customers can begin a transaction online and continue in-store, or vice-versa.”

To support this strategy, the company’s SAP environment needed to be more agile. Running  everything on-premises, from central finance to POS systems, meant that The Home Depot’s IT teams experienced redundancy and repetitive, manual processes. Their data warehouse needed an upgrade to process and analyze growing and increasingly diverse data sets. The Home Depot chose to migrate its SAP environment to Google Cloud to support both the velocity and scale needed for the business as well as critical analytics capabilities needed for its bold digital initiatives. “We chose Google Cloud to support our SAP implementation. Our decision had a lot to do with the relationship between Google Cloud and SAP and also for the applications and services that are offered by Google Cloud, like BigQuery, which are helping to enable data and analytics within our organization,” Moses explains.

After migrating its SAP applications—including S/4HANA, its customer activity repository (CAR), general ledger, e-commerce system, enterprise data warehouse and more to Google Cloud, the company now has the speed, scale and flexibility to tackle enormous spikes in the business, all while staying fully available for their customers. Additionally, The Home Depot was able to transform its financial systems and make them more agile to deliver critical information across multiple business functions in real time.

Maximizing data insights to support customer experiences

By migrating to Google Cloud, The Home Depot is leveraging Google Cloud analytics to build   the industry’s most efficient supply chain including more robust demand forecasting, supplier lead times, estimated delivery times and more, all while maintaining better security than before. “We experienced unprecedented change in our customers’ behavior and their buying patterns, which puts a lot of pressure on our supply chain,” explains Moses. “So having the ability to leverage data and analytics gives us insights to know exactly what it is that our customers need.” 

The company’s analysts now use BigQuery ML for machine learning directly against the company’s BigQuery data and use AutoML to determine the best model for predictions. The Home Depot’s engineers have also adapted BigQuery to monitor, analyze, and act on application performance data across all its stores and warehouses in real time—capabilities that were not as seamless in the on-premises environment.

With hundreds of projects on Google Cloud, The Home Depot’s cloud journey is well on track, but the company is always looking to the future. “As our customers’ needs have continued to evolve, and as technology has continued to evolve, our relationship with Google will continue to advance — to be able to innovate together, to be able to find new solutions together, to better serve our customers.”

Learn more about how The Home Depot is renovating its retail operation with SAP on Google Cloud.

10140

Of your peers have already watched this video.

1:00 Minutes

The most insightful time you'll spend today!

How-to

Predict User Churn on Gaming Apps with Google Analytics Data using BigQuery ML

User retention can be a major challenge for mobile game developers. According to the Mobile Gaming Industry Analysis in 2019, most mobile games only see a 25% retention rate for users after the first day. To retain a larger percentage of users after their first use of an app, developers can take steps to motivate and incentivize certain users to return. But to do so, developers need to identify the propensity of any specific user returning after the first 24 hours. 

In this blog post, we will discuss how you can use BigQuery ML to run propensity models on Google Analytics 4 data from your gaming app to determine the likelihood of specific users returning to your app.

You can also use the same end-to-end solution approach in other types of apps using Google Analytics for Firebase as well as apps and websites using Google Analytics 4. To try out the steps in this blogpost or to implement the solution for your own data, you can use this Jupyter Notebook

Using this blog post and the accompanying Jupyter Notebook, you’ll learn how to:

  • Explore the BigQuery export dataset for Google Analytics 4
  • Prepare the training data using demographic and behavioural attributes
  • Train propensity models using BigQuery ML
  • Evaluate BigQuery ML models
  • Make predictions using the BigQuery ML models
  • Implement model insights in practical implementations

Google Analytics 4 (GA4) properties unify app and website measurement on a single platform and are now default in Google Analytics. Any business that wants to measure their website, app, or both, can use GA4 for a more complete view of how customers engage with their business. With the launch of Google Analytics 4, BigQuery export of Google Analytics data is now available to all users. If you are already using a Google Analytics 4 property, you can follow this guide to set up exporting your GA data to BigQuery.

Once you have set up the BigQuery export, you can explore the data in BigQuery. Google Analytics 4 uses an event-based measurement model. Each row in the data is an event with additional parameters and properties. The Schema for BigQuery Export can help you to understand the structure of the data.  

In this blogpost, we use the public sample export data from an actual mobile game app called “Flood It!” (AndroidiOS) to build a churn prediction model. But you can use data from your own app or website. 

Here’s what the data looks like. Each row in the dataset is a unique event, which can contain nested fields for event parameters.

  SELECT *
FROM `firebase-public-project.analytics_153293282.events_*`
TABLESAMPLE SYSTEM (1 PERCENT)
table

This dataset contains 5.7M events from over 15k users.

  SELECT 
    COUNT(DISTINCT user_pseudo_id) as count_distinct_users,
    COUNT(event_timestamp) as count_events
FROM
  `firebase-public-project.analytics_153293282.events_*
count

Our goal is to use BigQuery ML on the sample app dataset to predict propensity to user churn or not churn based on users’ demographics and activities within the first 24 hours of app installation.

data

In the following sections, we’ll cover how to:

  1. Pre-process the raw event data from GA4
    1. Identify users & the label feature
    2. Process demographic features
    3. Process behavioral features
  2. Train classification model using BigQuery ML
  3. Evaluate the model using BigQueryML
  4. Make predictions using BigQuery ML
  5. Utilize predictions for activation

Pre-process the raw event data

You cannot simply use raw event data to train a machine learning model as it would not be in the right shape and format to use as training data. So in this section, we’ll go through how to pre-process the raw data into an appropriate format to use as training data for classification models.

This is what the training data should look like for our use case at the end of this section:

user id

Notice that in this training data, each row represents a unique user with a distinct user ID (user_pseudo_id). 

Identify users & the label feature

We first filtered the dataset to remove users who were unlikely to return the app anyway. We defined these ‘bounced’ users as ones who spent less than 10 mins with the app. Then we labeled all remaining users:

  • churned: No event data for the user after 24 hours of first engaging with the app.
  • returned: The user has at least one event record after 24 hours of first engaging with the app.

For your use case, you can have a different definition of bounce and churning. Also you can even try to predict something else other than churning, e.g.:

  • whether a user is likely to spend money on in-game currency 
  • likelihood of completing n-number of game levels
  • likelihood of spending n amount of time in-game etc.

In such cases, label each record accordingly so that whatever you are trying to predict can be identified from the label column.

From our dataset, we found that ~41% users (5,557) bounced. However, from the remaining users (8,031),  ~23% (1,883) churned after 24 hours:

  SELECT
    bounced,
    churned, 
    COUNT(churned) as count_users
FROM
    bqmlga4.returningusers
GROUP BY 1,2
ORDER BY bounced
boucned

To create these bounced and churned columns, we used the following snippet of SQL code. 

  ...
#churned = 1 if last_touch within 24 hr of app installation, else 0
IF (user_last_engagement < TIMESTAMP_ADD(user_first_engagement, 
      INTERVAL 24 HOUR),
    1,
    0 ) AS churned,
#bounced = 1 if last_touch within 10 min, else 0
IF (user_last_engagement <= TIMESTAMP_ADD(user_first_engagement, 
      INTERVAL 10 MINUTE),
    1,
    0 ) AS bounced,
...

You can view the Jupyter Notebook for the full query used for materializing the bounced and churned labels. 

Process demographic features

Next, we added features both for demographic data and for behavioral data spanning across multiple columns. Having a combination of both demographic data and behavioral data helps to create a more predictive model. 

We used the following fields for each user as demographic features:

  • geo.country
  • device.operating_system
  • device.language

A user might have multiple unique values in these fields — for example if a user uses the app from two different devices. To simplify, we used the values from the very first user engagement event.

  CREATE OR REPLACE VIEW bqmlga4.user_demographics AS (
  WITH first_values AS (
      SELECT
          user_pseudo_id,
          geo.country as country,
          device.operating_system as operating_system,
          device.language as language,
          ROW_NUMBER() OVER (PARTITION BY user_pseudo_id ORDER BY event_timestamp DESC) AS row_num
      FROM `firebase-public-project.analytics_153293282.events_*`
      WHERE event_name="user_engagement"
      )
  SELECT * EXCEPT (row_num)
  FROM first_values
  WHERE row_num = 1 #first engagement
);

Process behavioral features

There is additional demographic information present in the GA4 export dataset, e.g. app_info, device, event_params, geo etc. You may also send demographic information to Google Analytics through each hit via user_properties. Furthermore, if you have first-party data on your own system, you can join that with the GA4 export data based on user_ids. 

To extract user behavior from the data, we looked into the user’s activities within the first 24 hours of first user engagement. In addition to the events automatically collected by Google Analytics, there are also the recommended events for games that can be explored to analyze user behavior. For our use case, to predict user churn, we counted the number of times the follow events were collected for a user within 24 hours of first user engagement: 

  • user_engagement
  • level_start_quickplay
  • level_end_quickplay
  • level_complete_quickplay
  • level_reset_quickplay
  • post_score
  • spend_virtual_currency
  • ad_reward
  • challenge_a_friend
  • completed_5_levels
  • use_extra_steps

The following query shows how these features were calculated:

  WITH
  events_first24hr AS (
    SELECT
      e.*
    FROM
      `firebase-public-project.analytics_153293282.events_*` e
    JOIN
      bqmlga4.returningusers r
      ON
        e.user_pseudo_id = r.user_pseudo_id
    WHERE
      TIMESTAMP_MICROS(e.event_timestamp) <= r.ts_24hr_after_first_engagement
  )
SELECT
  user_pseudo_id,
  SUM(IF(event_name = 'user_engagement', 1, 0)) AS cnt_user_engagement,
  # ... repeated for all behavior data ... 
  SUM(IF(event_name = 'use_extra_steps', 1, 0)) AS cnt_use_extra_steps,
FROM
  events_first24hr
GROUP BY
  1

View the notebook for the query used to aggregate and extract the behavioral data. You can use different sets of events for your use case. To view the complete list of events, use the following query:

  SELECT
    event_name,
    COUNT(event_name) as event_count
FROM
    `firebase-public-project.analytics_153293282.events_*`
GROUP BY 1
ORDER BY
   event_count DESC

After this we combined the features to ensure our training dataset reflects the intended structure. We had the following columns in our table:

  • User ID:
    • user_pseudo_id
  • Label:
    • churned
  • Demographic features
    • country
    • device_os
    • device_language
  • Behavioral features
    • cnt_user_engagement
    • cnt_level_start_quickplay
    • cnt_level_end_quickplay
    • cnt_level_complete_quickplay
    • cnt_level_reset_quickplay
    • cnt_post_score
    • cnt_spend_virtual_currency
    • cnt_ad_reward
    • cnt_challenge_a_friend
    • cnt_completed_5_levels
    • cnt_use_extra_steps
    • user_first_engagement

At this point, the dataset was ready to train the classification machine learning model in BigQuery ML. Once trained, the model will output a propensity score between churn (churned=1) or return (churned=0) indicating the probability of a user churning based on the training data.

Train classification model 

When using the CREATE MODEL statement, BigQuery ML automatically splits the data between training and test. Thus the model can be evaluated immediately after training (see the documentation for more information).

For the ML model, we can choose among the following classification algorithms where each type has its own pros and cons:

model

Often logistic regression is used as a starting point because it is the fastest to train. The query below shows how we trained the logistic regression classification models in BigQuery ML.

  CREATE OR REPLACE MODEL bqmlga4.churn_logreg
TRANSFORM(
  EXTRACT(MONTH from user_first_engagement) as month,
  EXTRACT(DAYOFYEAR from user_first_engagement) as julianday,
  EXTRACT(DAYOFWEEK from user_first_engagement) as dayofweek,
  EXTRACT(HOUR from user_first_engagement) as hour,
  * EXCEPT(user_first_engagement, user_pseudo_id)
)
OPTIONS(
  MODEL_TYPE="LOGISTIC_REG",
  INPUT_LABEL_COLS=["churned"]
) AS
SELECT
  *
FROM
  bqmlga4.train

We extracted monthjulianday, and dayofweek  from datetimes/timestamps as one simple example of additional feature preprocessing before training. Using TRANSFORM() in your CREATE MODEL query allows the model to remember the extracted values. Thus, when making predictions using the model later on, these values won’t have to be extracted again. View the notebook for the example queries to train other types of models (XGBoost, deep neural network, AutoML Tables).

Evaluate model

Once the model finished training, we ran ML.EVALUATE to generate precisionrecallaccuracy and f1_score for the model:

  SELECT
  *
FROM
  ML.EVALUATE(MODEL bqmlga4.churn_logreg)
row

The optional THRESHOLD parameter can be used to modify the default classification threshold of 0.5. For more information on these metrics, you can read through the definitions on precision and recallaccuracyf1-scorelog_loss and roc_auc. Comparing the resulting evaluation metrics can help to decide among multiple models.Furthermore, we used a confusion matrix to inspect how well the model predicted the labels, compared to the actual labels. The confusion matrix is created using the default threshold of 0.5, which you may want to adjust to optimize for recall, precision, or a balance (more information here).

  SELECT
  expected_label,
  _0 AS predicted_0,
  _1 AS predicted_1
FROM
  ML.CONFUSION_MATRIX(MODEL bqmlga4.churn_logreg)
expected

This table can be interpreted in the following way:

actual

Make predictions using BigQuery ML

Once the ideal model was available, we ran ML.PREDICT to make predictions. For propensity modeling, the most important output is the probability of a behavior occurring. The following query returns the probability that the user will return after 24 hrs. The higher the probability and closer it is to 1, the more likely the user is predicted to return, and the closer it is to 0, the more likely the user is predicted to churn.

  SELECT
  user_pseudo_id,
  returned,
  predicted_returned,
  predicted_returned_probs[OFFSET(0)].prob as probability_returned
FROM
  ML.PREDICT(MODEL bqmlga4.churn_logreg,
  (SELECT * FROM bqmlga4.train)) #can be replaced with a proper test dataset

Utilize predictions for activation

Once the model predictions are available for your users, you can activate this insight in different ways. In our analysis, we used user_pseudo_id as the user identifier. However, ideally, your app should send back the user_id from your app to Google Analytics. In addition to using first-party data for model predictions, this will also let you join back the predictions from the model into your own data.

  • You can import the model predictions back into Google Analytics as a user attribute. This can be done using the Data Import feature for Google Analytics 4. Based on the prediction values you can Create and edit audiences and also do Audience targeting. For example, an audience can be users with prediction probability between 0.4 and 0.7, to represent users who are predicted to be “on the fence” between churning and returning.
  • For Firebase Apps, you can use the Import segments feature. You can tailor user experience by targeting your identified users through Firebase services such as Remote Config, Cloud Messaging, and In-App Messaging. This will involve importing the segment data from BigQuery into Firebase. After that you can send notifications to the users, configure the app for them, or follow the user journeys across devices.
  • Run targeted marketing campaigns via CRMs like Salesforce, e.g. send out reminder emails.

You can find all of the code used in this blogpost in the Github repository:

https://github.com/GoogleCloudPlatform/analytics-componentized-patterns/tree/master/gaming/propensity-model/bqml

What’s next? 

Continuous model evaluation and re-training

As you collect more data from your users, you may want to regularly evaluate your model on fresh data and re-train the model if you notice that the model quality is decaying.

Continuous evaluation—the process of ensuring a production machine learning model is still performing well on new data—is an essential part in any ML workflow. Performing continuous evaluation can help you catch model drift, a phenomenon that occurs when the data used to train your model no longer reflects the current environment. 

To learn more about how to do continuous model evaluation and re-train models, you can read the blogpost: Continuous model evaluation with BigQuery ML, Stored Procedures, and Cloud Scheduler

More resources

If you’d like to learn more about any of the topics covered in this post, check out these resources:

Or learn more about how you can use BigQuery ML to easily build other machine learning solutions:

Let us know what you thought of this post, and if you have topics you’d like to see covered in the future! You can find us on Twitter at @polonglin and @_mkazi_.Thanks to reviewers: Abhishek Kashyap, Breen Baker, David Sabater Dinter.

More Relevant Stories for Your Company

Case Study

Serverless and BigQuery Together on Google Cloud: Behind the L’Oreal Beauty Tech Data Platform

Editor's note: In Today's guest post we hear from beauty leader L'Oréal about their approach to building a modern data platform on fully managed services: managing the ingest of diverse datasets into BigQuery with Cloud Run, and orchestrating transformations into relevant business domain representations for stakeholders across the organization. Learn

How-to

Why and How to Migrate to Google BigQuery

Over the past few decades, organizations have mastered the science of data warehousing. They have increasingly applied descriptive analytics to large quantities of stored data, gaining insight into their core business operations. Conventional Business Intelligence (BI), which focuses on querying, reporting, and Online Analytical Processing, might have been a differentiating factor

Case Study

Google had Enough of Flooding in India. Here’s What It Did About It

Over the last century, floods have become the most common and deadly natural disaster on the planet. Many countries currently lack effective early warning systems and alerts. Today, 20 percent of flood fatalities occur in India. To help, Google sent a team to study the Ghaghara River, near Patna. “In

Trend Analysis

Four Key Takeaways from Google and Quantum Metric’s Back-to-School Retail Benchmark Study

Is it September yet? Hardly! School is barely out for the summer. But according to Google and Quantum Metric research, the back-to-school and off-to-college shopping season – which in the U.S. is second only to the holidays in terms of purchasing volume1 – has already begun. For retailers, that means

SHOW MORE STORIES