5528
Of your peers have already watched this video.
1:30 Minutes
The most insightful time you'll spend today!
Pega Systems Migrates SAP Servers to Google Cloud in Just 9 Weeks!
Pega Systems’ financial data on SAP environs were on a hosting platform that lacked agility. By moving nearly 30 SAP servers to Google Cloud in just 9 weeks, Pega Systems was able to unlock data and integrate BigQuery into SAP HANA to deliver personalization for clients and embark on an exciting journey with Google Cloud! Watch now.

3597
Of your peers have already downloaded this article
1:30 Minutes
The most insightful time you'll spend today!
Through customer interviews, a survey, and data aggregation, Forrester concluded that migrating and running SAP on Google Cloud has a number of benefits, including financial benefits.
Download this Forrester infographic to understand the 3-year financial impact it can have on your organization.
Being Cloud-native Means Sustainability and Growth-native for Nuuly!

3411
Of your peers have already read this article.
3:00 Minutes
The most insightful time you'll spend today!
They say black never goes out of style. It’s something the team at Nuuly, URBN’s digital rental and resale business, know well. And it’s not just true of the company’s garments but their gadgets, too.
“I was having an offhand conversation with a UX designer recently,” Rebecca Sandercock, Nuuly’s strategy and insights manager, recalled in a recent interview from the company’s sunny, South Philly headquarters. “The designer was talking about how we had chosen dark mode for a number of interfaces internally because it actually saves so much on electrical output. They had the data to back that decision up, but more importantly, it’s just the kind of thing everyone is thinking about all the time here.”
Of course most every business is thinking about sustainability in some way these days. What makes Nuuly stand out is how quickly it can act, thanks in large part to the technology platform that’s made the entire enterprise possible.
“We’re kind of sustainable by our very nature,” Kim Gallagher, Nuuly’s director of marketing and customer success, said.
While she meant the rental and resale business, which helps customers buy fewer clothes and sees Nuuly items worn many times by many people—Gallagher could just as well have been referring to the sustainability inherent in, and enabled by, cloud computing.

There are the obvious, and oft-cited, advantages, such as how centralized data centers can operate more efficiently (some have been carbon neutral since the beginning). Yet there are even more subtle yet substantial benefits. In a marketplace and climate that are both changing faster and faster, sustainability requires a certain amount of agility. Such adaptability and scalability are intrinsic to the cloud technology that threads its way throughout Nuuly.
It turns out that being cloud native also means being sustainability native—as well as growth native. Since its 2019 launch, Nuuly’s net sales have risen roughly 6x over the first three fiscal years.
Cloud fits any situation
When URBN was developing Nuuly—launching in just 10 months—it chose to create everything from scratch on Google Cloud. Despite being part of a larger organization with decades of history and expertise, the company recognized the limitations presented by legacy systems and, more importantly, the necessity of building a wholly new platform that could be fully responsive.
The company has to react not only to new fashion trends but, crucially, the changing behaviors of customers. And not just their evolving tastes but also shopping habits, delivery preferences, unexpected customer service requests—is this a pattern, or a stain?—and social media chatter.
The pressure for a successful launch was high. The URBN portfolio, which also includes Urban Outfitters, Anthropologie, and Free People, had to keep evolving to satisfy a new generation of shopper who exists in an increasingly crowded and demanding digital marketplace.
The cloud’s responsiveness has thus proven its worth in creating a financially sustainable business as well as an environmentally sustainable one. Those even go hand in hand, as Gallagher points out: “Every customer who keeps renting is one who isn’t buying more occasion-specific clothes that go unworn most of the time.”

Dr. Alan Rosenwinkel, director of data science at URBN, had heard from his team about one garment that has become an emblem for the power of the platform. “It had been rented 25 different times before someone loved it enough to buy it and keep it forever,” he explained. (It’s also an emblem of the power of the cloud, that they would have the data awareness to track a single item so closely.)
It’s a new way of shopping made possible by a new way of computing. As one writer for Business Insider cheered, Nuuly “completely cured my addiction to fast-fashion.”
It makes for a healthy business, too. Sales for fiscal year 2019 exceeded $8 million and surpassed $24 million in 2020—one of the few URBN segments to grow during a tough year for fashion—and reached $47 million in 2021. The subscriber base had grown to 51,000 at the end of January.
Agility drives sustainability drives agility
For digital retailers to achieve such customer enthusiasm often relies as much on how the clothes get there as how they look. And those deliveries turn out to be a prime example of where Nuuly’s sustainability and technology meet.
For now, all shipping is handled through a state-of-the-art distribution center in Bucks County on the Philadelphia outskirts. Garments are shipped six-at-a-time, in fully reusable packaging, using ground transportation to keep the carbon footprint to a minimum. In certain limited geographic areas, shipments were sometimes taking longer than the two to five days most members would find acceptable.

“If we were an older company or weren’t set up from a technology perspective to be agile, we might have just said, ‘All right, we’re going to just go to three-day shipping for everyone,’” Rosenwinkel said. “That would drive up the cost, and the environmental impact. But we were able to be more strategic and more targeted about it.”
By regularly analyzing customer sentiment and retention data through BigQuery and Cloud Composer, and tying those to historical shipping times, Nuuly has been able to understand how shipping speed impacts its customers. Using a custom-built order management system, deliveries can be automatically adjusted to arrive more quickly, particularly when certain regions or days of the week are proving difficult to reach customers in time. These accelerated deliveries, like all Nuuly shipments, are made via UPS’s Carbon Offset program.
“Because we have the data, and the platforms to analyze it all,” Rosenwinkel said, “we’re delivering faster with the bare minimum impact on cost and emissions.”
It’s just one example of how modern retailers must juggle so many demands from consumers, workers, suppliers, and even regulators. Adding sustainability to that mix could be seen as a burden, but Nuuly shows how the right technology can lead to a holistic approach that makes all those interests work together even better.
And it allows for more opportunities and more kinds of sustainability, reaching from the designer’s atelier to the customer’s doorstep.
Sustainable details at every level
Back at the distribution center, workers can experience sustainability in a different way, as data is leveraged to enhance their well-being.
All workers are equipped with customized Android devices that help guide order tracking and fulfillment, plus cleaning, repairs, and reselling of garments as needs change throughout their lifecycle. Yet the insights go even deeper. The data science team has closely analyzed routes and repetitive motions for workers to keep their strenuous jobs as low-impact as the company’s broader environmental footprint.
“Through machine learning, we estimate we’ll be able to save our workers over 300,000 miles of steps over a five year period,” Rosenwinkel said. That’s enough walking to circumnavigate the globe 12.5 times.

The company has also applied ingenuity to one of the most notorious aspects of digital retail: packaging. Made from 100% post-consumer recycled materials like plastic bottles, Nuuly’s reusable carriers require no disposable bags or hangers to send goods back and forth. And once the packages have reached their end of life, the team is working with designers on ways to repurpose them into items that can then be offered for rental or sale on Nuuly.
“We’re really looking at the circular economy from every angle,” Gallagher said. “It’s built into the business.”
That includes not just the research the team did on-site at the Dry Cleaning & Laundry Institute—”We went to laundry school,” Gallagher jokes—but the digital tools that take those lessons even further. The team is always looking for ways to optimize fabric care, both to cut down on chemicals and water usage and extend the life of a garment. By analyzing the lifespan of every item, Nuuly not only makes them last longer but can identify problems faster. Employees even use custom apps to mark stains and damage so repair teams can more easily identify and fix the issues.
With its meticulous inventory tracking, Nuuly can even take marginal items, like a white gown or jeans with a small stain, and turn them into a custom dye job or an upcycling opportunity with a partner. These reworked items are then inserted back into the Nuuly Rent inventory as part of a growing collection of one-of-a-kind pieces, called Re_Nuuly.
Other services have launched quickly and easily thanks to the company’s cloud-enabled backend. Wanting to encourage more community and more reuse, the company created Nuuly Thrift. Debuting last fall after just a year in development, the team built everything from new interfaces to an evolved point of sale system.

It’s enabled Nuuly to offer many garments for rent or sale simultaneously, with “truly real-time inventories,” Rosenwinkel said. “So when it’s gone on one site, it shows up as gone on every site—no more surprises.”
Except for the good kind.
“I like to think we’re helping our customers think about ownership in a completely new way,” Sandercock said. “Once they run out of a use for a garment, they can offer it back, and sell it, and someone else will get to enjoy it and give it a new life. And Nuuly, we get to keep it in the community and keep it in the ecosystem, which is really cool—we’re taking extended responsibility over what happens to the clothing we create.”
4819
Of your peers have already watched this video.
2:30 Minutes
The most insightful time you'll spend today!
Haaretz on Google Cloud Guarantees Faster, Reliable & Responsive Services to its Audience
Israeli centenarian newspaper, Haaretz relied on on-prem infrastructure to serve readers digitally. As the need for scalability, security and business intelligence grew alongside their readership, Haaretz was looking for more than just a cloud-based solution to replace their infrastructure. Inon Gershovitz, CTO, Haaretz takes us through the journey of recreating digital news experience, serving growing reader traffic and keeping up the editorial standards with Google Cloud.
Watch the video to hear from Haaretz’ leaders on delivering personalized content to readers with Google Cloud solutions!
10156
Of your peers have already watched this video.
1:00 Minutes
The most insightful time you'll spend today!
Predict User Churn on Gaming Apps with Google Analytics Data using BigQuery ML
User retention can be a major challenge for mobile game developers. According to the Mobile Gaming Industry Analysis in 2019, most mobile games only see a 25% retention rate for users after the first day. To retain a larger percentage of users after their first use of an app, developers can take steps to motivate and incentivize certain users to return. But to do so, developers need to identify the propensity of any specific user returning after the first 24 hours.
In this blog post, we will discuss how you can use BigQuery ML to run propensity models on Google Analytics 4 data from your gaming app to determine the likelihood of specific users returning to your app.
You can also use the same end-to-end solution approach in other types of apps using Google Analytics for Firebase as well as apps and websites using Google Analytics 4. To try out the steps in this blogpost or to implement the solution for your own data, you can use this Jupyter Notebook.
Using this blog post and the accompanying Jupyter Notebook, you’ll learn how to:
- Explore the BigQuery export dataset for Google Analytics 4
- Prepare the training data using demographic and behavioural attributes
- Train propensity models using BigQuery ML
- Evaluate BigQuery ML models
- Make predictions using the BigQuery ML models
- Implement model insights in practical implementations
Google Analytics 4 (GA4) properties unify app and website measurement on a single platform and are now default in Google Analytics. Any business that wants to measure their website, app, or both, can use GA4 for a more complete view of how customers engage with their business. With the launch of Google Analytics 4, BigQuery export of Google Analytics data is now available to all users. If you are already using a Google Analytics 4 property, you can follow this guide to set up exporting your GA data to BigQuery.
Once you have set up the BigQuery export, you can explore the data in BigQuery. Google Analytics 4 uses an event-based measurement model. Each row in the data is an event with additional parameters and properties. The Schema for BigQuery Export can help you to understand the structure of the data.
In this blogpost, we use the public sample export data from an actual mobile game app called “Flood It!” (Android, iOS) to build a churn prediction model. But you can use data from your own app or website.
Here’s what the data looks like. Each row in the dataset is a unique event, which can contain nested fields for event parameters.
SELECT *FROM `firebase-public-project.analytics_153293282.events_*`TABLESAMPLE SYSTEM (1 PERCENT)

This dataset contains 5.7M events from over 15k users.
SELECTCOUNT(DISTINCT user_pseudo_id) as count_distinct_users,COUNT(event_timestamp) as count_eventsFROM`firebase-public-project.analytics_153293282.events_*

Our goal is to use BigQuery ML on the sample app dataset to predict propensity to user churn or not churn based on users’ demographics and activities within the first 24 hours of app installation.

In the following sections, we’ll cover how to:
- Pre-process the raw event data from GA4
- Identify users & the label feature
- Process demographic features
- Process behavioral features
- Train classification model using BigQuery ML
- Evaluate the model using BigQueryML
- Make predictions using BigQuery ML
- Utilize predictions for activation
Pre-process the raw event data
You cannot simply use raw event data to train a machine learning model as it would not be in the right shape and format to use as training data. So in this section, we’ll go through how to pre-process the raw data into an appropriate format to use as training data for classification models.
This is what the training data should look like for our use case at the end of this section:

Notice that in this training data, each row represents a unique user with a distinct user ID (user_pseudo_id).
Identify users & the label feature
We first filtered the dataset to remove users who were unlikely to return the app anyway. We defined these ‘bounced’ users as ones who spent less than 10 mins with the app. Then we labeled all remaining users:
- churned: No event data for the user after 24 hours of first engaging with the app.
- returned: The user has at least one event record after 24 hours of first engaging with the app.
For your use case, you can have a different definition of bounce and churning. Also you can even try to predict something else other than churning, e.g.:
- whether a user is likely to spend money on in-game currency
- likelihood of completing n-number of game levels
- likelihood of spending n amount of time in-game etc.
In such cases, label each record accordingly so that whatever you are trying to predict can be identified from the label column.
From our dataset, we found that ~41% users (5,557) bounced. However, from the remaining users (8,031), ~23% (1,883) churned after 24 hours:
SELECTbounced,churned,COUNT(churned) as count_usersFROMbqmlga4.returningusersGROUP BY 1,2ORDER BY bounced

To create these bounced and churned columns, we used the following snippet of SQL code.
...#churned = 1 if last_touch within 24 hr of app installation, else 0IF (user_last_engagement < TIMESTAMP_ADD(user_first_engagement,INTERVAL 24 HOUR),1,0 ) AS churned,#bounced = 1 if last_touch within 10 min, else 0IF (user_last_engagement <= TIMESTAMP_ADD(user_first_engagement,INTERVAL 10 MINUTE),1,0 ) AS bounced,...
You can view the Jupyter Notebook for the full query used for materializing the bounced and churned labels.
Process demographic features
Next, we added features both for demographic data and for behavioral data spanning across multiple columns. Having a combination of both demographic data and behavioral data helps to create a more predictive model.
We used the following fields for each user as demographic features:
geo.countrydevice.operating_systemdevice.language
A user might have multiple unique values in these fields — for example if a user uses the app from two different devices. To simplify, we used the values from the very first user engagement event.
CREATE OR REPLACE VIEW bqmlga4.user_demographics AS (WITH first_values AS (SELECTuser_pseudo_id,geo.country as country,device.operating_system as operating_system,device.language as language,ROW_NUMBER() OVER (PARTITION BY user_pseudo_id ORDER BY event_timestamp DESC) AS row_numFROM `firebase-public-project.analytics_153293282.events_*`WHERE event_name="user_engagement")SELECT * EXCEPT (row_num)FROM first_valuesWHERE row_num = 1 #first engagement);
Process behavioral features
There is additional demographic information present in the GA4 export dataset, e.g. app_info, device, event_params, geo etc. You may also send demographic information to Google Analytics through each hit via user_properties. Furthermore, if you have first-party data on your own system, you can join that with the GA4 export data based on user_ids.
To extract user behavior from the data, we looked into the user’s activities within the first 24 hours of first user engagement. In addition to the events automatically collected by Google Analytics, there are also the recommended events for games that can be explored to analyze user behavior. For our use case, to predict user churn, we counted the number of times the follow events were collected for a user within 24 hours of first user engagement:
user_engagementlevel_start_quickplaylevel_end_quickplaylevel_complete_quickplaylevel_reset_quickplaypost_scorespend_virtual_currencyad_rewardchallenge_a_friendcompleted_5_levelsuse_extra_steps
The following query shows how these features were calculated:
WITHevents_first24hr AS (SELECTe.*FROM`firebase-public-project.analytics_153293282.events_*` eJOINbqmlga4.returningusers rONe.user_pseudo_id = r.user_pseudo_idWHERETIMESTAMP_MICROS(e.event_timestamp) <= r.ts_24hr_after_first_engagement)SELECTuser_pseudo_id,SUM(IF(event_name = 'user_engagement', 1, 0)) AS cnt_user_engagement,# ... repeated for all behavior data ...SUM(IF(event_name = 'use_extra_steps', 1, 0)) AS cnt_use_extra_steps,FROMevents_first24hrGROUP BY1
View the notebook for the query used to aggregate and extract the behavioral data. You can use different sets of events for your use case. To view the complete list of events, use the following query:
SELECTevent_name,COUNT(event_name) as event_countFROM`firebase-public-project.analytics_153293282.events_*`GROUP BY 1ORDER BYevent_count DESC
After this we combined the features to ensure our training dataset reflects the intended structure. We had the following columns in our table:
- User ID:
user_pseudo_id
- Label:
churned
- Demographic features
countrydevice_osdevice_language
- Behavioral features
cnt_user_engagementcnt_level_start_quickplaycnt_level_end_quickplaycnt_level_complete_quickplaycnt_level_reset_quickplaycnt_post_scorecnt_spend_virtual_currencycnt_ad_rewardcnt_challenge_a_friendcnt_completed_5_levelscnt_use_extra_stepsuser_first_engagement
At this point, the dataset was ready to train the classification machine learning model in BigQuery ML. Once trained, the model will output a propensity score between churn (churned=1) or return (churned=0) indicating the probability of a user churning based on the training data.
Train classification model
When using the CREATE MODEL statement, BigQuery ML automatically splits the data between training and test. Thus the model can be evaluated immediately after training (see the documentation for more information).
For the ML model, we can choose among the following classification algorithms where each type has its own pros and cons:

Often logistic regression is used as a starting point because it is the fastest to train. The query below shows how we trained the logistic regression classification models in BigQuery ML.
CREATE OR REPLACE MODEL bqmlga4.churn_logregTRANSFORM(EXTRACT(MONTH from user_first_engagement) as month,EXTRACT(DAYOFYEAR from user_first_engagement) as julianday,EXTRACT(DAYOFWEEK from user_first_engagement) as dayofweek,EXTRACT(HOUR from user_first_engagement) as hour,* EXCEPT(user_first_engagement, user_pseudo_id))OPTIONS(MODEL_TYPE="LOGISTIC_REG",INPUT_LABEL_COLS=["churned"]) ASSELECT*FROMbqmlga4.train
We extracted month, julianday, and dayofweek from datetimes/timestamps as one simple example of additional feature preprocessing before training. Using TRANSFORM() in your CREATE MODEL query allows the model to remember the extracted values. Thus, when making predictions using the model later on, these values won’t have to be extracted again. View the notebook for the example queries to train other types of models (XGBoost, deep neural network, AutoML Tables).
Evaluate model
Once the model finished training, we ran ML.EVALUATE to generate precision, recall, accuracy and f1_score for the model:
SELECT*FROMML.EVALUATE(MODEL bqmlga4.churn_logreg)

The optional THRESHOLD parameter can be used to modify the default classification threshold of 0.5. For more information on these metrics, you can read through the definitions on precision and recall, accuracy, f1-score, log_loss and roc_auc. Comparing the resulting evaluation metrics can help to decide among multiple models.Furthermore, we used a confusion matrix to inspect how well the model predicted the labels, compared to the actual labels. The confusion matrix is created using the default threshold of 0.5, which you may want to adjust to optimize for recall, precision, or a balance (more information here).
SELECTexpected_label,_0 AS predicted_0,_1 AS predicted_1FROMML.CONFUSION_MATRIX(MODEL bqmlga4.churn_logreg)

This table can be interpreted in the following way:

Make predictions using BigQuery ML
Once the ideal model was available, we ran ML.PREDICT to make predictions. For propensity modeling, the most important output is the probability of a behavior occurring. The following query returns the probability that the user will return after 24 hrs. The higher the probability and closer it is to 1, the more likely the user is predicted to return, and the closer it is to 0, the more likely the user is predicted to churn.
SELECTuser_pseudo_id,returned,predicted_returned,predicted_returned_probs[OFFSET(0)].prob as probability_returnedFROMML.PREDICT(MODEL bqmlga4.churn_logreg,(SELECT * FROM bqmlga4.train)) #can be replaced with a proper test dataset
Utilize predictions for activation
Once the model predictions are available for your users, you can activate this insight in different ways. In our analysis, we used user_pseudo_id as the user identifier. However, ideally, your app should send back the user_id from your app to Google Analytics. In addition to using first-party data for model predictions, this will also let you join back the predictions from the model into your own data.
- You can import the model predictions back into Google Analytics as a user attribute. This can be done using the Data Import feature for Google Analytics 4. Based on the prediction values you can Create and edit audiences and also do Audience targeting. For example, an audience can be users with prediction probability between 0.4 and 0.7, to represent users who are predicted to be “on the fence” between churning and returning.
- For Firebase Apps, you can use the Import segments feature. You can tailor user experience by targeting your identified users through Firebase services such as Remote Config, Cloud Messaging, and In-App Messaging. This will involve importing the segment data from BigQuery into Firebase. After that you can send notifications to the users, configure the app for them, or follow the user journeys across devices.
- Run targeted marketing campaigns via CRMs like Salesforce, e.g. send out reminder emails.
You can find all of the code used in this blogpost in the Github repository:
What’s next?
Continuous model evaluation and re-training
As you collect more data from your users, you may want to regularly evaluate your model on fresh data and re-train the model if you notice that the model quality is decaying.
Continuous evaluation—the process of ensuring a production machine learning model is still performing well on new data—is an essential part in any ML workflow. Performing continuous evaluation can help you catch model drift, a phenomenon that occurs when the data used to train your model no longer reflects the current environment.
To learn more about how to do continuous model evaluation and re-train models, you can read the blogpost: Continuous model evaluation with BigQuery ML, Stored Procedures, and Cloud Scheduler
More resources
If you’d like to learn more about any of the topics covered in this post, check out these resources:
- BigQuery export of Google Analytics data
- BigQuery ML quickstart
- Events automatically collected by Google Analytics 4
- Qwiklabs: Create ML models with BigQuery ML
Or learn more about how you can use BigQuery ML to easily build other machine learning solutions:
- How to build demand forecasting models with BigQuery ML
- How to build a recommendation system on e-commerce data using BigQuery ML
Let us know what you thought of this post, and if you have topics you’d like to see covered in the future! You can find us on Twitter at @polonglin and @_mkazi_.Thanks to reviewers: Abhishek Kashyap, Breen Baker, David Sabater Dinter.
Google Cloud Next ’22 to Commence in October: Block Your Calendar!

7910
Of your peers have already read this article.
1:30 Minutes
The most insightful time you'll spend today!
We’re excited to announce that Google Cloud Next returns on October 11–13, 2022.
Join us for keynotes from industry luminaries and engage live with Google developers. Explore dynamic content across various learning levels, and dive deep into technologies and solutions spanning the Google Cloud and Google Workspace portfolios. Participate in breakout sessions, demos, and hands-on training. Hear from the world’s leading companies about their digital transformation journeys. You’ll have opportunities to connect with experts, get inspired, and boost your skills. We can’t wait to see you at Next ’22!
It’s too early to determine how the event experience will span the digital and physical worlds, so please stay tuned for updates as we plan with the health and safety of the attendees in mind. In the meantime, mark October 11–13 in your calendar, and visit our event site for updates. For more inspiration, rediscover Next ’21, now available on demand.
More Relevant Stories for Your Company

Mid-Sized B2B Firm Achieves the Business Trifecta with a Single Strategy
Thirteen years’ experience in e-commerce has given Teddy Chan, Chief Executive Officer and Chief Technology Officer, AfterShip, a deep understanding of the challenges of shipping and tracking packages to customers worldwide. “The key problem many merchants face is customers asking ‘where is my order?’ and ‘when I will get the package?’”

Transform Your Business: Comprehensive Cloud Services and Tailored Pricing Plans
As the saying goes, “it’s hard to make predictions, especially about the future.” Some organizations find it challenging to predict what cloud resources they’ll need in months or years ahead. Every organization is on its own unique cloud journey. To help, we’re developing new ways for customers to consume and

How L&T Financial Services Processes 95% of Motorcycle Loans in Less Than Two Minutes
L&T Financial Services is one of the largest lenders in India. India’s demonetization policy in recent years has led to a shift from cash transactions to digital payments. In 2016, the government withdrew 500 and 1000 rupee notes from circulation and encouraged a heavily cash-based population to deposit their canceled notes

IT Team Figures Out Easiest Way to Build Data Pipelines and Create ML Models
Building a strong brand in today's hyper-competitive business environment takes vision. It also requires a flexible, easily managed approach to digital asset management (DAM), so marketing professionals and other stakeholders can easily share, store, track, and manipulate assets to build the brand. Many of today's leading companies, including JetBlue, Slack,






