Poor Product Discovery Causes Shoppers’ Abandonment, Time to Augment ‘Search Experience’ with Google’s Retail Search!

3102
Of your peers have already read this article.
2:30 Minutes
The most insightful time you'll spend today!
How many times has a shopper searched for a product on a store’s website only to get results that aren’t relevant—or worse, provided no search results at all? While most ecommerce sites have search capabilities, few accomplish their ultimate goal: making it easy for customers to discover the products they want.
Some 94% of U.S. consumers abandoned a shopping session because they received irrelevant search results, according to a 2021 survey conducted by The Harris Poll and Google Cloud. It’s a phenomenon known as “search abandonment.” Indeed, poor product discovery experiences can stop a purchase in its tracks and leave shoppers frustrated. Retailers miss out on a staggering $300 billion each year due to search abandonment in the U.S. alone.
With the announcement today of our latest discovery solution, Retail Search, now generally available, retailers worldwide can supercharge their websites and mobile apps with Google-quality search. Built on Google’s technologies that understand user intent and context, the solution helps businesses improve the search and overall shopping experience across all of their digital touchpoints.
Already, early adopters like Lowes, Fnac Darty, and Casas Pernambucanas have been using Google Cloud’s discovery solutions to increase sales conversions, boost basket sizes, and improve customer engagement.
Understanding user intent in search is a hard problem
While we’ve come a long way from the days when search was largely based on keywords and boolean rules, shoppers still struggle to find what they’re looking for. They often have to come up with a perfectly-worded query that a retailer’s site search engine will understand, sometimes rephrasing several times before getting the results they want—if they find anything relevant at all.
Traditional search technologies don’t work in the modern age of online retail, where tens or even hundreds of thousands of items are available on a single ecommerce site.

Today, people expect search engines to understand their intent more deeply, return relevant results faster, and help them discover new products easily with personalized recommendations.
Fortunately, product discovery experiences on retail sites can now offer superior, contextual experiences for consumers, combining marketing and data science insights along with advanced search technologies based on machine learning and artificial intelligence.
Now, through the power of Retail Search, when a shopper searches for a “long black dress with short sleeves and comfortable fit” on an ecommerce site, they should immediately get results for precisely that—rather than refining their search multiple times, or worse, giving up their shopping journey.
Helping retailers solve their search experience woes
Retail Search transforms the shopping experience and makes it easier for shoppers to find relevant products by surfacing results that are intuitive and contextual.
This fully managed service is easily customizable, enabling organizations to craft shopper-focused search experiences. Our site search solution builds upon decades of Google’s experience and innovation in search indexing, retrieval, and ranking. Retailers can make product discovery even easier for shoppers, while optimizing for their business goals with advanced capabilities, including:
- Advanced query understanding that produces better results from even the broadest queries, including non-product searches.
- Semantic search to effectively match product attributes with website content for fast, relevant product discovery.
- Optimized results that leverage user interaction and ranking models to meet specific business goals.
- State-of-the-art security and privacy practices that ensure retailer data is isolated with strong access controls and is only used to deliver relevant search results on their own properties.
Retailers can now shape product discovery into a shopping experience that is timely, relevant, and personalized—all powered by Google Cloud.
Early adopters seeing immediate impact
Several leading retailers around the world have been able to quickly create better online shopping experiences for their customers and capture more digital and omnichannel growth using Google Cloud’s product discovery solutions.
“With limited customer signals and no historical data, descriptive long-tail searches are some of the most challenging queries to understand,” said Neelima Sharma, senior vice president, technology, e-commerce, marketing and merchandising at Lowe’s. “We have been partnering with Google Cloud to give our customers relevant results for long-tail searches and have seen an increase in click-through and search conversion and a drop in our ‘No Results Found’ rate since we launched.”
“Making constant improvements to our website search engines has always been a priority for us as we aim to give our customers a simpler, more customized and enhanced online shopping experience,” said Olivier Theulle, Fnac Darty’s chief ecommerce and digital officer. “As we implement Google Cloud’s site search solution on our Fnac Darty websites and are the first French retailer to do so, we expect the solution to deliver increased conversion rates while also offering greater customer satisfaction.”
“Google Cloud has helped improve the indexing and quality of search results on Pernambucanas’ digital platforms, providing a better experience to our customers and improving sales conversions,” said Fabiano Rustice, chief information officer at Pernambucanas. “On Black Friday in November 2021—the largest retail date in Brazil—we saw a 20% reduction in search refinements per user, for which Retail Search was instrumental. We are proud to be an early adopter and to have Google Cloud as a strategic innovation partner.”
In addition to Google Cloud’s efforts working directly with customers, our ecosystem of key partners, including GroupBy, Lucidworks, GridDynamics, SpringML, and others are also leveraging Google Cloud discovery solutions to provide value added services to their customers.
“In the fast paced and extremely competitive ecommerce environment, we are seeing that forward-looking and innovative retailers are prioritizing product discovery,” said Roland Gossage, CEO of GroupBy. “GroupBy has proudly partnered with Google Cloud to launch its product discovery solutions across several client environments. Already we have seen a more than 10% increase in online revenues, which equates to millions of dollars, and this is just the beginning.”
To learn more, visit Discovery Solutions for Retail or contact your Google Cloud field sales representative.
Enabling Real-time AI with Streaming Ingestion in Vertex AI

2510
Of your peers have already read this article.
2:30 Minutes
The most insightful time you'll spend today!
Many machine learning (ML) use cases, like fraud detection, ad targeting, and recommendation engines, require near real-time predictions. The performance of these predictions is heavily dependent on access to the most up-to-date data, with delays of even a few seconds making all the difference. But it’s difficult to set up the infrastructure needed to support high-throughput updates and low-latency retrieval of data.
Starting this month, Vertex AI Matching Engine and Feature Store will support real-time Streaming Ingestion as Preview features. With Streaming Ingestion for Matching Engine, a fully managed vector database for vector similarity search, items in an index are updated continuously and reflected in similarity search results immediately. With Streaming Ingestion for Feature Store, you can retrieve the latest feature values with low latency for highly accurate predictions, and extract real-time datasets for training.
For example, Digits is taking advantage of Vertex AI Matching Engine Streaming Ingestion to help power their product, Boost, a tool that saves accountants time by automating manual quality control work.“Vertex AI Matching Engine Streaming Ingestion has been key to Digits Boost being able to deliver features and analysis in real-time. Before Matching Engine, transactions were classified on a 24 hour batch schedule, but now with Matching Engine Streaming Ingestion, we can perform near real time incremental indexing – activities like inserting, updating or deleting embeddings on an existing index, which helped us speed up the process. Now feedback to customers is immediate, and we can handle more transactions, more quickly,” said Hannes Hapke, Machine Learning Engineer at Digits.
This blog post covers how these new features can improve predictions and enable near real-time use cases, such as recommendations, content personalization, and cybersecurity monitoring.

Streaming Ingestion enables real-time AI
As organizations recognize the potential business impact of better predictions based on up-to-date data, more real-time AI use cases are being implemented. Here are some examples:
- Real-time recommendations and a real-time marketplace: By adding Streaming Ingestion to their existing Matching Engine-based product recommendations, Mercari is creating a real-time marketplace where users can browse products based on their specific interests, and where results are updated instantly when sellers add new products. Once it’s fully implemented, the experience will be like visiting an early-morning farmer’s market, with fresh food being brought in as you shop. By combining Streaming Ingestion with Matching Engine’s filtering capability, Mercari can specify whether or not an item should be included in the search results, based on tags such as “online/offline” or “instock/nostock.”

Mercari Shops: Streaming Ingestion enables real-time shopping experiment
- Large-scale personalized content streaming: For any stream of content representable with feature vectors (including text, images, or documents), you can design pub-sub channels to pick up valuable content for each subscriber’s specific interests. Because Matching Engine is scalable (i.e., it can process millions of queries each second), you can support millions of online subscribers for content streaming, serving a wide variety of topics that are changing dynamically. With Matching Engine’s filtering capability, you also have real-time control over what content should be included, by assigning tags such as “explicit” or “spam” to each object. You can use Feature Store as a central repository for storing and serving the feature vectors of the contents in near real time.
- Monitoring: Content streaming can also be used for monitoring events or signals from IT infrastructure, IoT devices, manufacturing production lines, and security systems, among other commercial use cases. For example, you can extract signals from millions of sensors and devices and represent them as feature vectors. Matching Engine can be used to continuously update a list of “the top 100 devices with possible defective signals,” or “top 100 sensor events with outliers,” all in near real time.
- Threat/spam detection: If you are monitoring signals from security threat signatures or spam activity patterns, you can use Matching Engine to instantly identify possible attacks from millions of monitoring points. In contrast, security threat identification based on batch processing often involves potentially significant lag, leaving the company vulnerable. With real-time data, your models are better able to catch threats or spams as they happen in your enterprise network, web services, online games, etc.
Implementing streaming use cases
Let’s take a closer look at how you can implement some of these use cases.
Real-time recommendations for retail
Mercari built a feature extraction pipeline with Streaming Ingestion.

The feature extraction pipeline is defined with Vertex AI Pipelines, and is periodically invoked by Cloud Scheduler and Cloud Functions to initiate the following process:
- Get item data: The pipeline issues a query to fetch the updated item data from BigQuery.
- Extract feature vector: The pipeline runs predictions on the data with the word2vec model to extract feature vectors.
- Update index: The pipeline calls Matching Engine APIs to add the feature vectors to the vector index. The vectors are also saved to Cloud Bigtable (and can be replaced with Feature Store in the future).
“We have been evaluating the Matching Engine Streaming Ingestion and couldn’t believe the super short latency of the index update for the first time. We would like to introduce the functionality to our production service as soon as it becomes GA, ” said Nogami Wakana, Software Engineer at Souzoh (a Mercari group company).
This architecture design can be also applied to any retail businesses that need real-time updates for product recommendations.
Ad targeting
Ad recommender systems benefit significantly from real-time features and item matching with the most up-to-date information. Let’s see how Vertex AI can help build a real-time ad targeting system.

The first step is generating a set of candidates from the ad corpus. This is challenging because you must generate relevant candidates in milliseconds and ensure they are up to date. Here you can use Vertex AI Matching Engine to perform low-latency vector similarity matching, generate suitable candidates, and use Streaming Ingestion to ensure that your index is up-to-date with the latest ads.
Next is reranking the candidate selection using a machine learning model to ensure that you have a relevant order of ad candidates. For the model to use the latest data, you can use Feature Store Streaming Ingestion to import the latest features and use online serving to serve feature values at low latency to improve accuracy.
After reranking the ads candidates, you can apply final optimizations, such as applying the latest business logic. You can implement the optimization step using a Cloud Function or Cloud Run.
What’s Next?
Interested? The documents for Streaming Ingestion are available and you can try it out now. Using the new feature is easy: For example, when you create an index on Matching Engine with the REST API, you can specify the indexUpdateMethod attribute as STREAM_UPDATE.
{
displayName: "'${DISPLAY_NAME}'",
description: "'${DISPLAY_NAME}'",
metadata: {
contentsDeltaUri: "'${INPUT_GCS_DIR}'",
config: {
dimensions: "'${DIMENSIONS}'",
approximateNeighborsCount: 150,
distanceMeasureType: "DOT_PRODUCT_DISTANCE",
algorithmConfig: {treeAhConfig: {leafNodeEmbeddingCount: 10000, leafNodesToSearchPercent: 20}}
},
},
indexUpdateMethod: "STREAM_UPDATE"
}After deploying the index, you can update or rebuild the index (feature vectors) with the following format. If the data point ID exists in the index, the data point is updated, otherwise, a new data point is inserted.
{
datapoints: [
{datapoint_id: "'${DATAPOINT_ID_1}'", feature_vector: [...]},
{datapoint_id: "'${DATAPOINT_ID_2}'", feature_vector: [...]}
]
}It can handle the data point insertion/update at high throughput with low latency. The new data point values will be applied in any new queries within a few seconds or milliseconds (the latency varies depending on the various conditions).
The Streaming Ingestion is a powerful functionality and very easy to use. No need to build and operate your own streaming data pipeline for real-time indexing and storage. Yet, it adds significant value to your business with its real-time responsiveness.
To learn more, take a look at the following blog posts for learning Matching Engine and Feature Store concepts and use cases:
3074
Of your peers have already watched this video.
14:30 Minutes
The most insightful time you'll spend today!
Overview of AI Notebooks on Google Cloud
Artificial intelligence and machine learning are one of the most disruptive technologies that enterprises have encountered in the last four or five years. And AI and ML will continue to disrupt enterprises going forward for the next 10 years.
The key to unlocking value with artificial intelligence starts with a model. And the simplest way to develop a model is to use notebooks, and within notebooks to use open source frameworks to develop AI models.
Notebooks are the go-to tool for data scientists to develop and deploy AI/ML models. In this overview video, Suds Narasimhan, Product Manager, Google Cloud, walks you through an overview of cloud AI notebooks, Google Cloud’s managed Jupyter lab notebook service for enterprises on the Google Cloud.
He explores how cloud AI notebooks can help enterprise data scientists explore data quickly and develop an AI and ML model and deploy it into production.
Analytics Hub for Secure Data Sharing and Analytics Unlocks True Data Value and Insights

6749
Of your peers have already read this article.
2:30 Minutes
The most insightful time you'll spend today!
Customers tell us that sharing and exchanging data with other organizations is a critical element of their analytics strategy, but it’s hamstrung by unreliable data and processes, and only getting harder with security threats and privacy regulations on the rise.
Furthermore, traditional data sharing techniques use batch data pipelines that are expensive to run, create late arriving data, and can break with any changes to the source data. They also create multiple copies of data, which brings unnecessary costs and can bypass data governance processes. These techniques do not offer features for data monetization, such as managing subscriptions and entitlements. Altogether, these challenges mean that organizations are unable to realize the full potential of transforming their business with shared data.
To address these limitations, we are introducing Analytics Hub, a new fully managed service, available in Q3, in preview, that helps you unlock the value of data sharing, leading to new insights and increased business value. With Analytics Hub you get:
- A rich data ecosystem by publishing and subscribing to analytics-ready datasets.
- Control and monitoring over how your data is being used, because data is shared in one place.
- A self-service way to access valuable and trusted data assets, including data provided by Google. For example, a unique dataset from Google Search Trends will be available, that you can query and combine with your own data.
- An easy way to monetize your data assets without the overhead of building and managing the infrastructure.
Built on a decade of cross-organizational sharing
While Analytics Hub is a new service, it builds on BigQuery, Google’s petabyte-scale, serverless cloud data warehouse. BigQuery’s unique architecture provides separation between compute and storage, enabling data publishers to share data with as many subscribers as you want without having to make multiple copies of your data. With BigQuery, there are no servers to deploy or manage, which means that data consumers get immediate value from shared data. Data can be provided and consumed in real-time using the streaming capabilities of BigQuery and you can leverage the built in machine learning, geospatial, and natural language capabilities of BigQuery or take advantage of the native business intelligence support with tools like Looker, Google Sheets, and Data Studio.
BigQuery has had cross-organizational, in-place data sharing capabilities since it was introduced in 2010. We took a look at usage metrics in BigQuery and found that over a 7 day period in April, we had over 3,000 different organizations sharing over 200 petabytes of data. These numbers don’t include data sharing between departments within the same organization.

As you can see, data sharing in BigQuery is already popular. But we want to make it easier and even more scalable.
Raising the bar on data sharing
To make data sharing easier and more scalable in BigQuery, Analytics Hub introduces the concepts of shared datasets and exchanges. As a data publisher, you create shared datasets that contain the views of data that you want to deliver to your subscribers. Next, you create exchanges, which are used to organize and secure shared datasets. By default, exchanges are completely private, which means that only the users and groups that you give access to can view or subscribe to the data. You can also create internal exchanges or leverage public exchanges provided by Google. Finally, you publish shared datasets into an exchange to make them available to subscribers.
Data subscribers search through the datasets that are available across all exchanges for which they have access and subscribe to relevant datasets. This creates a linked dataset in their project that they can query and join with their own data. Subscribers pay for the queries that they run against the data while the publisher pays for the storage of the data. Data providers can add new data, new tables, or new columns to the shared dataset and these will be immediately available to subscribers. In addition, the publisher can track subscribers, disable subscriptions, and see aggregated usage information for the shared data.
Analytics Hub makes it easy for you to publish, discover, and subscribe to valuable datasets that you can combine with your own data to derive unique insights. Here are some types of data that will be available through Analytics Hub:
- Public datasets: Easy access to the existing repository of over 200 public datasets, including data about weather and climate, cryptocurrency, healthcare and life sciences, and transportation.
- Google datasets: Unique, freely-available datasets from Google. One example of this is the COVID-19 community mobility dataset. Another example is the forthcoming Google Trends dataset, which will provide the top 25 search terms and top 25 rising search terms over a 5 year window in 210 distinct locations in the US. Trends data can be used by everyone in the organization to gain insights into what customers care about.
- Commercial (paid for) datasets: We are working with leading commercial data providers to bring their data products to Analytics Hub. If you are interested in delivering your data via Analytics Hub, we’re also introducing Data Gravity, an initiative that provides storage benefits and new distribution paths for data published through Analytics Hub.
- Internal datasets: We know that data sharing can be challenging in larger organizations. Analytics Hub can be used for internal data, for example, to share standardized customer demographics with your sales engineering and data science teams.
Customers and partners using Analytics Hub

“Google Search Trends data has always been an important tool for our WPP agency data teams. At WPP we believe that data variety is a superpower which is why we are excited to use the new Trends dataset availability within BigQuery, plus the launch of Analytics Hub. The best creativity in the world is informed by data insights, and influenced by what people search for, so the operational efficiencies we’ll gain via the Analytics Hub and the insights we can drive with Trends data are just phenomenal.”
—Di Mayze Global Head of Data and AI, WPP

“Equifax Ignite is our shared data analytics environment within our Equifax data fabric. We are excited to partner with Google to leverage Analytics Hub and BigQuery to deliver data to over 400 statisticians and data modelers as well as securely sharing data with our partner financial institutions.”
—Kumar Menon, SVP Data Fabric and Decision Science, Equifax

“The flow of data and insights between our teams at Deloitte and our clients is paramount for building truly transformational data cultures. With its purpose-built architecture for secure data exchanges and sharing analytics resources, Google Cloud’s Analytics Hub can help provide significant operational efficiencies for how Deloitte teams support our clients’ data-driven initiatives within their industry ecosystems. It will also help minimize the worries about scale, privacy and security, or the administrative burden associated with each.”
—Navin Warerkar, Managing Director, Deloitte Consulting LLP, and US Google Cloud Data & Analytics GTM Lead

“Crux Informatics is proud to partner with Google to support the launch of Analytics Hub, removing friction for those who need access to analytics-ready data. With thousands of datasets from over 140 sources, Crux Informatics will accelerate access to data on Analytics Hub and together provide a more efficient and cost effective solution to deliver datasets in Google Cloud’s ecosystem.”
—Will Freiberg, CEO, Crux Informatics
Next steps for Analytics Hub
This is just the beginning for Analytics Hub. As we get to preview and general availability, we will be adding additional capabilities, including workflows for publishing and subscribing, publishing analytics assets (Looker Blocks, Data Studio reports, Connected Google Sheets) along with the shared data, the ability for data publishers to specify query restrictions on the usage of their data, and making it easy for data publishers to create sandbox environments for subscribers to work with their data, even if they are not yet on Google Cloud. We will provide features in Analytics Hub for monetization of data, including managing subscriptions, data entitlements, and billing.
Please sign up for the preview, which is scheduled to be available in the third quarter of 2021. In the meantime, you can learn more about BigQuery and how to leverage its built-in data sharing capabilities. Please go to g.co/cloud/analytics-hub to register your interest in Analytics Hub.
Enhancing SAP Build Process Automation with Google Document AI and Google Workspace

4144
Of your peers have already read this article.
4:00 Minutes
The most insightful time you'll spend today!
SAP Build Process Automation is designed to optimize business processes and boost efficiency. The platform helps both business users and developers alike digitize core workflows and incorporate artificial intelligence (AI) into time consuming and error-prone manual tasks.
All digital paths can benefit from automation. The pandemic, supply chain shortages, and other disruptive events have upped the pressure on businesses and their workers to perform in more efficient and flexible ways.
Google Cloud and SAP have responded by providing an integrated toolbox that can fundamentally change the way businesses operate — all while creating value for their customers.
With a focus on taking process automation to an even more advanced level and removing inefficiencies from workflows, SAP has introduced integrations with Google Cloud Document AI, and Google Workspace for SAP Build Process Automation customers. This integration can reduce or eliminate many repetitive and error-prone tasks, so that companies can help save money, operate more productively, and scale more easily.
AI unleashes innovation
Continuous advances in digital systems and advanced technology introduce new opportunities to rethink workflows. SAP recognizes the role that AI-powered automation can play in transforming workflows, and that the benefits of doing so extend beyond basic time savings and cost cutting. By plugging Google Cloud Document AI into the application, SAP Build Process Automation’s low-code, no-code platform enables SAP to help its customers in lines of business and IT integrate multiple applications while democratizing access to governed machine learning technology.
Machine learning and process automation can drive efficiency with SAP S/4HANA
Customers using SAP S/4HANA can build workflow improvements into all major core processes, such as order entry, invoice creation, asset posting, and many others, helping them to be faster and more efficient in the process. Google Cloud Document AI extracts key elements — including addresses, article numbers, price, quantity, and more — within emails, PDFs, handwritten notes, and other formats.
Integrated automation can drive results for invoice and purchase order processing
An example of the improvements these integrations have made to SAP Build Process Automation is the use of AI, productivity tools, and automation for processing purchase orders. In the past, workers had to manually select relevant orders in their Gmail accounts and extract key data from large PDF files, including the order date, supplier details, article numbers, quantity, unit price, and the total amount. Then, workers would have to enter all of the individual line items one by one into the SAP S/4HANA system.
Today, through SAP’s integrated automation platform, customers can automatically extract and organize data by Google Workspace (Google Sheets, Google Drive, and Gmail) and Document AI. This works by extracting data contained in Gmail attachments, downloading it into Google Drive and extracting fields using Document AI’s pretrained models. The tool then enters the consolidated order information from Google Sheets into SAP S/4HANA. For example, the Canton of Zurich in Switzerland experienced a significant reduction of workload to process compensation forms once the organization implemented this automation. Furthermore, implementing the automation can avoid audit and compliance issues, and improve data quality.
The screenshot below shows how a workflow operates within the SAP Build Process Automation software.

Sales, procurement, finance, and other functions are also able to handle more strategic work that delivers greater value to customers with these SAP and Google Cloud integrations. They’ve also boosted both security and regulatory compliance for customers, including those in the financial services space.
For example, the Google Cloud Document AI technology can process thousands of supply chain invoices, validates and systematically approves them. And, over time, the machine learning and AI components improve the analysis process and ensure that data processed with SAP Build Process Automation is adhering to a company’s best practices and essential regulatory requirements.
SAP and Google Cloud put automation to work
Achieving the most accurate, efficient business processes is possible with an automation framework that embeds collaboration and productivity applications while giving lines of business and IT users access to machine learning technology. SAP Build Process Automation combined with Google Cloud Document AI, and Google Workspace has proven to be a catalyst in driving innovation and business transformation for customers across multiple industries, leading to an average of 22% to 30% faster time to market, and significant financial gains, including up to 20% accounts payable savings potential. An example of these improvements includes those experienced by TasNetworks, which had a 25% reduction in back-office processing efforts after implementing these technologies.
To learn more about how Google Cloud and SAP are building solutions for accelerating business value, visit cloud.google.com/solutions/sap. You can find more information about SAP Build Process Automation at sap.com/build-automation.
To start your transformation journey today, choose the SAP Business Technology Platform region that’s best for you. We’re also happy to announce that Google Cloud offers the first and only option to run BTP in the cloud in India — learn more here: Google Cloud’s newest SAP Business Technology Platform Region.
Lufthansa: Wind Forecasting with Google Cloud ML Helps Increase On-time Flights

2501
Of your peers have already read this article.
3:30 Minutes
The most insightful time you'll spend today!
The magnitude and direction of wind significantly impacts airport operations, and Lufthansa Group Airlines are no exception. A particularly troublesome kind is called BISE: it is a cold, dry wind that blows from the northeast to southwest in Switzerland, through the Swiss Plateau. Its effects on flight schedules can be severe, such as forcing planes to change runways, which can create a chain reaction of flight delays and possible cancellations. In Zurich Airport, in particular, BISE can potentially reduce capacity by up to 30%, leading to further flight delays and cancellations, and to millions in lost revenue for Lufthansa (as well as dissatisfaction among their passengers).
Being able to predict this kind of wind well in advance lets the Network Operations Control team schedule flight operations optimally across runways and timeslots, to minimize disruptions to the schedule. However, predicting speed and magnitude can be incredibly difficult to model and thus to predict— which is why Lufthansa reached out to Google Cloud.
Machine learning (ML) can help airports and airlines to better anticipate and manage these types of disruptive weather events. In this blog post, we’ll explore an experiment Lufthansa did together with Google Cloud and its Vertex AI Forecast service, accurately predicting BISE hours in advance, with more than 40% relative improvement in accuracy over internal heuristics, all within days instead of the months it often takes to do ML projects of this magnitude and performance.

“Being impressed with Google’s technology and prowess in the field of AI and machine learning, we were certain that my working together with their expert, to combine our technology with their domain expertise, we would achieve the best results possible,“ said Christian Most, Senior Director, Digital Operations Optimization at Lufthansa Group.
Collecting and preparing the dataset
The goal of Lufthansa and Google Cloud’s project was to forecast the BISE wind for Zurich’s Kloten Airport using deep learning-based ML approaches, then to see if the prediction surpasses internal heuristics-driven solutions and gauge the ease of use and practicality of the deep learning approach in production.
Since deep learning-based techniques require large datasets, the project relies on Meteoswiss simulation data, a dataset consisting of multiple meteorological sensor measurements collected from several weather stations across Switzerland over the past five years. By using this dataset, we obtained data on factors like wind direction, speed, pressure, temperature, humidity and more, at a 10 min resolution, along with some information about the location of the weather stations, such as altitude. These factors, which we hypothesized to be predictive of the BISE, ended up carrying valuable signals, as we would see later.
This collected data was next subjected to an extensive cleaning and feature engineering process using Vertex AI Workbench, in order to prepare the final dataset for training. The cleaning phase included steps to drop the features, or rows, that contained too many missing values, or failed statistical tests for entropy, etc. Since the direction of wind is a circular feature (between 0 and 360 degrees), this column/feature was replaced with two features: the corresponding sine and cosine embedding. The dataset was then flattened such that the columns contained all the relevant features and sensor measurements from all the weather stations at a particular 10-minute interval.
Since the target variable — i.e,. BISE — was not directly available, we engineered a proxy target variable for BISE called “tailwind speed around runway,” which above a certain threshold indicates the presence of BISE along the runway.
Forecasting wind in the Cloud
Once the dataset was ready, Lufthansa and Google Cloud evaluated several options before deciding to experiment and tune Vertex AI Forecast, Google’s AutoML-powered forecasting service, in order to achieve optimum results. Vertex Forecast is capable of the required feature engineering, neural architecture search, and hyper parameter tuning, and it is managed by Google Cloud to score in the top 2.5% in the M5 Forecasting Competition on Kaggle, in a completely automated fashion. These qualities made it an excellent choice for Lufthansa, to reduce the manual overhead of creating, deploying, and maintaining top performing deep learning models.

The raw data files were loaded from cloud storage, preprocessed on Vertex AI Workbench. Then, a training pipeline was initiated on Vertex AI Pipelines, which performed the following steps in sequence:
The .csv data file was loaded from Cloud Storage into a Vertex AI managed dataset.
A Vertex AI forecasting training job was initiated with the dataset, and it was also registered as a model in the Vertex AI Model Registry.
Upon completion, the model was evaluated on the test set, and the model’s predictions and the input features and ground truth of the test set, were stored in a user-defined table in BigQuery. Several test metrics were also available on the service and model dashboards.
One of the biggest challenges was the severe imbalance in the dataset, as measurements with BISE were very far and few in between. In order to account for this, instances where BISE occurred, as well the occurrences temporally close to them, were upweighted using weights calculated with methods including Inverse of Square Root of Number of Samples (ISNS), Effective Number of Samples (ENS), and Gaussian reweighting. The formulas for the methods are given below. These weights were supplied as separate columns in the dataset, and were iteratively used thereafter by the service as the “weight” column.
ISNS

ENS

Weighted gaussian

Results and next steps


In the above figures, the x-axis represents the forecast horizon and the Y-axis shows the respective metrics (Recall/F1-score). As shown after multiple experiments, we can see Vertex AI Forecast achieved higher recall and precision t (red bar), outperforming Lufthansa’s internal baseline heuristics, with the performance gap widening steadily as the forecast horizon extends further into the future. At the two-hour mark, our custom-configured Vertex AI Forecast model improved by 40% relative to the internal heuristics and 1700% compared to the random guess baseline. As we saw with other experiments, at a six-hour forecast horizon, the performance gap widens even more, with Vertex AI Forecast in the lead. Since forecasting BISE a few hours in advance is very beneficial to prevent flight delays for Lufthansa, this was a great solution for them.
“We are very excited to be able to not only do accurate long term forecasts for the BISE, but also that Vertex AI Forecasting makes training and deploying such models much easier and faster, allowing us to innovate rapidly to serve our customers and stakeholders in the best possible manner,” said Swiss Oliver Rueegg, Product Owner, Swiss International Airlines.
Lufthansa plans to explore productionizing this solution by integrating it into their Operations Decision Support Suite, which is used by the network controllers in the Operations Control Center in Kloten, as well as to work closely with Google’s specialists to integrate both Vertex AI Forecast and other of Google’s AI/ML offerings for their use cases.
More Relevant Stories for Your Company

Streamline Your Business Processes with Google Cloud’s Custom Document Splitter
Businesses rely on processing an inflow of documents to drive processes and make decisions. Many such documents are combined into a single file. For example, a loan application may have a driver’s license, paystub, W2, bank statement, and other document types within a single file. The complexity of handling many

KLM’s Doubles Bookings With the Same Spend With Machine Learning
GOALS Develop smarter, more effective media buying models through dataDrive relevant advertisingScale predictive modelling across all touchpoints in the customer journey APPROACH Combined contextual data to create a predictive model with granular layers Activated data in real-time RESULTS 40% lower cost per bookingMore than twice as many bookings at same

Manhattan Associates and Google Cloud: How the Partnership Accelerates Future of Digital Retail
While the shift to digital business and the cloud has been well under way for some years now, organizations today have a new sense of urgency due to COVID-19. Delivering digital transformation is no longer a ‘nice to have’ option, rather, it is an operational imperative. Taking advantage of the

Google Cloud’s No-Cost Skill Badge: Up Your Generative AI Game
Generative AI is a rapidly expanding technology with a wide range of potential applications. Google Cloud Learning is thrilled to offer a new, no-cost Generative AI Fundamentals skill badge. This skill badge is designed for anyone eager to learn about the power of generative AI. No technical skills or prior knowledge






