Multicloud Mindset: Thinking About Open Source and Security in a Multicloud World - Build What's Next
Blog

Multicloud Mindset: Thinking About Open Source and Security in a Multicloud World

2777

Of your peers have already read this article.

1:30 Minutes

The most insightful time you'll spend today!

Need some helpful best practices for thinking about security in multicloud environments? Here's a blog discussing the impact of open source and novel security challenges in the multicloud world.

There’s never been a better time to talk about multicloud, and the Google Cloud Multicloud Mindset series on Twitter Spaces was created to do just that! This series takes place once every two weeks and features live conversations with top experts about the latest multicloud topics. You can join the 15-minute Q&A to ask your top questions and listen to episodes later offline for up to 30 days after we chat.

If you happened to miss our last few episodes, we recommend checking out our introduction blog to the series for what you missed. Let’s dive into our latest episodes, discussing the impact of open source and novel security challenges in multicloud environments.

Episode #5: ‘The intersection of open source and multicloud’

Open source technology has been an integral part of computing since its earliest era, predating even the birth of technology hubs like Silicon Valley. Open source projects have been responsible for giving us some of the most popular software in the world, such as Mozilla Firefox and the operating system Linux.

In the fifth episode, we sat down with Mike Coleman, Cloud Developer Advocate at Google Cloud, and took a closer look into the history of open source technologies, the role they play in a multicloud world, and the developer perspective on using these technologies to do their work.

The concept of multicloud anchors on the ability to run workloads across clouds and being able to pick the providers that are best suited for specific parts of workloads. Adopting open source technologies and languages empower companies to use the tools they need, regardless of cloud provider, without the fear of getting locked into a specific provider.

“As you think about moving across different environments, whether that be cloud to cloud, or developer desktop to ultimate destination, whether that be your data center or the cloud. Open source software allows you to do that…and multicloud is just an extension of that. This idea that I need to run the same software wherever I go.” — Mike Coleman, Cloud Developer Advocate at Google Cloud

If you’ve ever wanted a developer’s take on the impact of multicloud and the influence of open source in software development and digital transformation trends, you’ll want to tune into this episode.

You can access the full conversation on Twitter Spaces.

Episode #6: ‘Novel challenges in security with multicloud’

In the sixth episode of the series, we chatted with Dr. Anton Chuvakin, Security Advisor at Office of the CISO at Google Cloud, about how security leaders and architects are shifting away from traditional security models, which are increasingly insufficient for multicloud environments.

As more organizations adopt multicloud approaches, the question of how to maintain security in these complex environments and the increasing burden on SecOps teams is top of mind. As Dr. Chuvakin noted, the challenges in the cloud facing more traditional teams range from types of telemetry and logs to volumes and lack of clarity on detection use cases. However, these issues intensify when extended to include multiple clouds, where learning how to do something on one provider may be completely different on another.

“If you end up multicloud, you need to know public cloud and how it works at a better level than you would if you’re going to a single provider. Just like if you’re trying to repair three cars, you need to first learn how to repair cars. You need to have more cloud knowledge to do multicloud, not less. You need to have more powerful superpowers in the public cloud computing area because you can’t just learn one provider and call it a day.” — Dr. Anton Chuvakin, Security Advisor at Office of the CISO at Google Cloud

During the discussion, he offered three tips for tackling multicloud security:

  1. Learn cloud more, not less if you’re going multicloud. Multicloud requires more cloud knowledge because you can’t learn a single provider and call it a day. You’ll need to understand the differences in order to be able to secure multiple cloud environments.
  2. Focus on learning cloud identity management and how it compares to your traditional identity management service functions. Start with identifying the differences and similarities in what you see in one cloud and then continue with other clouds you use.
  3. Explore where your threat areas change in cloud environments when you plan detection and response activities to understand if your detection is covered across clouds.

If your organization is embracing multicloud, this is a great episode to listen and learn more about cloud security, the primary considerations and challenges facing security teams, and some helpful best practices for thinking about security in multicloud environments.

We’ll be sharing the latest topics and episodes with you every month in this blog series. Until next time.

Blog

HarbourBridge Schema Assistant Allows Quick, Bulk Migration to Cloud Spanner

5479

Of your peers have already read this article.

1:30 Minutes

The most insightful time you'll spend today!

Google Cloud announces the open-source HarbourBridge Schema Assistant for a guided schema-design workflow for migrating from MySQL or PostgreSQL to Cloud Spanner. Learn more.

Today we’re announcing the HarbourBridge Schema Assistant, which provides a guided schema-design workflow for migrating from MySQL or PostgreSQL to Spanner. HarbourBridge imports dump files (from mysqldump or pg_dump) or directly connects to your source database, and converts the source database schema to an equivalent Spanner schema. The new Schema Assistant capability displays the source schema and Spanner schema side-by-side, highlights errors and walks you through a series of steps to validate and optimize your Spanner schema. It also produces a browsable assessment report with an overall migration-fitness score for Spanner, a table-by-table detailed analysis of type mappings and a list of features used in the source database that aren’t supported by Spanner. It supports editing of table and column names, column types, primary keys and constraints, as well as dropping of tables, columns, foreign keys and secondary indexes.

The new Schema Assistant complements HarbourBridge’s existing data and schema migration capabilities and is a critical step towards our goal of building a complete open-source migration toolkit. HarbourBridge continues to support command-line schema and data migration and turn-key Spanner evaluation.

Complementing the bulk data migration capabilities of HarbourBridge, we are also announcing the ability to migrate change events from MySQL to Cloud Spanner.

image4.png
HarborBridge takes your MySQL or PostgreSQL schema and translates it to a Spanner schema. It will provide you with a detailed report of all the changes and spanner fit scoes.

Supported Features in Schema Assistant

  1. Global type mapping. Users can customize the global mapping for how types should be mapped to Spanner consistently across the schema. For example, mapping large integers in source schema to Spanner’s NUMERIC.
  2. Local type mapping. Users can override the custom type mapping for a given table/column.
  3. Session management. A session keeps track of all the changes made to the schema mapping.
  4. Customization of secondary indexes. Users can add, edit and delete secondary indexes to optimize their Spanner performance.
  5. Customization of foreign keys and interleaved tables. Table interleaving is an important design consideration when migrating to Cloud Spanner as explained in more detail in this blog post.
image1.png
Global type mapping from MySQL to Spanner
image3.png
tables and columns mapping from source to destination

Features in the pipeline

We are already working to further expand the supported set of schema editing features and welcome your feedback. We are particularly excited to expand the Schema Assistant’s design recommendations for optimizing Spanner schemas e.g. in-depth recommendations for primary key design.

HarbourBridge is open source and we gladly accept contributions from the wider community.

Blog

STAC-M3 Tick History Analytics in Google Cloud Benchmark Results Reveals it is 18X Faster than Previous Version

4764

Of your peers have already read this article.

5:00 Minutes

The most insightful time you'll spend today!

Google Cloud's redesigned STAC -M3™ benchmark suite was recently audited to assess performance of tick analytics stack and other data points that are relevant for financial firms to perform I/O intensive and compute-intensive tasks with market data.

The Securities Technology Analysis Center (STAC®), an organization that improves technology discovery and assessment in the finance industry through dialog and research, recently audited the STAC-M3™ benchmark suite on Google Cloud (SUT ID KDB211210). These enterprise tick-analytics benchmarks assess the ability of a solution stack such as database software, servers, and storage, to perform a variety of I/O-intensive and compute-intensive operations on historical market data.

Following up on our previous STAC-M3 benchmark audit (SUT ID KDB181001), a redesigned Google Cloud architecture leveraged the most recent version of kdb+ 4.0, the time-series database from KX, and achieved significant improvements: 35 out of 41 benchmarks ran faster in the new cluster – by up to 18x faster than Google Cloud’s prior results. Key highlights include the following:

Compared to the previous STAC-M3 Antuco suite results on Google Cloud:

  • Was faster in 13 of 17 mean response-time benchmarks
  • Was 18x faster – a 94% reduction in run time – in the version of Year-High Bid that allows caching (STAC-M3.ß1.1T.YRHIBID-2.TIME), which also set an overall record for all published results
  • Had 9x higher throughput in Year-High Bid (STAC-M3.ß1.1T.YRHIBID.MBPS)

Compared to the previous STAC-M3 Kanaga suite results on Google Cloud:

  • Was faster in 22 of 24 mean response-time benchmarks
  • Was over 10x faster in all four Market Snapshot workloads (STAC-M3.ß1.10T.YR[2,3,4,5]-MKTSNAP.TIME)
  • Had 5x the throughput in Year-High Bid involving 2 years of data (STAC-M3.ß1.1T.2YRHIBID.MBPS)

“The STAC-M3 standard was designed by financial firms to reveal the performance of tick analytics stacks. Generational improvements like those exhibited by Google Cloud’s most recent STAC-M3 audit, are important data points for firms evaluating new architectures for performance and scale,” said Peter Nabicht, President of STAC.

These performance results may translate to real-world advantages that may be difficult for investment firms to achieve in static and costly on-premises environments: immediate answers in high data velocity markets, more thoroughly explored research theories by adding data or new quantitative approaches, and reduced costs by releasing cloud resources more quickly.

STAC-M3: High-speed tick analytics


Designing for record-breaking results

In our STAC-M3 audit, the stack under test (SUT) was designed to take advantage of horizontal scalability in the cloud by sharding data across independent compute nodes. The cluster of 12 Google Compute Engine N2 instances was powered by Intel Cascade Lake, with each node using 32 vCPUs, 160GiB of memory, and 9TiB of local NVMe SSDs. The full STAC-M3 Antuco and Kanaga data set was split across the cluster and kdb+ scripts distributed queries between nodes.

Figure 1.  Sharded kdb+ 4.0 STAC-M3 architecture on Google Cloud.

This configuration was the sweet spot for this particular workload, but this architecture does not need to be limited to 12 nodes for other workloads – the data sharding algorithm could scale to any number of nodes as required by workload demands. Since scaling out the cluster in this manner increases the total pool of available storage, this architecture can continue scaling out to petabytes of storage across hundreds of nodes.

The ability to spawn large clusters with hundreds of thousands of processors on demand at low cost, and to delete the resources when jobs complete, not only changes the economics of running computations on large financial data sets, it also opens up opportunities to explore solutions to new types of problems that were previously overlooked due to the constraints of fixed hardware on-premises. You can check the pricing of this VM configuration using the Google Cloud Pricing Calculator. The costs can be reduced even further by using preemptible VMs.

While the new cluster used a similar number of nodes, cores, and total memory as the previously-audited cluster, the redesigned architecture allowed us to harness the low latency and high throughput of Local NVMe SSDs.

Resources on demand

The cluster was created on demand using Terraform and Ansible during testing and auditing. The use of infrastructure as code (IaC) techniques ensured that the cluster, fully loaded with the STAC-M3 data set, could be created when needed and then removed when benchmarking was complete. It also meant that the cluster configuration was enforced by code on each deployment, eliminating configuration variance and drift. The full IaC definition to create the cluster can be retrieved from the report in the STAC Vault.

Each time the cluster was created, data was streamed to Local SSDs from Google Cloud Storage, our reliable and secure object storage, at up to the line rate of 32Gbps per node. The entire 57TiB STAC-M3 Antuco and Kanaga data was replicated from Cloud Storage to local storage in approximately 20 minutes.

Figure 2. STAC-M3 sharded data set.

Since each node was independent and responsible for its own shard of data, doubling the cluster size would cut the synchronization time in half, or copy twice as much data in the same amount of time. Using higher bandwidth options of up to 100Gbps would triple the possible throughput for a relatively small incremental cost, trading an approximately 11%-23% price increase at current list prices for a 200% data synchronization performance increase. Taking advantage of fast networking to cache sharded data in parallel to a large cluster makes storing bulk data in Cloud Storage viable for even the largest workloads.

For quants working on vast data sets in sprawling compute clusters, the ability to fully describe infrastructure as declarative code, create elastic resources on demand, cache data quickly from cheap bulk storage, and turn resources off when computations complete is a dramatic change compared to waiting months to grow on-premises clusters – and a compelling reason to use cloud infrastructure.

To see how we designed and optimized the cluster for API-driven cloud resources, read our new whitepaper.

STAC-A2™: Calculating derivatives risk


In 2018, we showed that cloud instances can outperform bare metal when analyzing large tick history data sets in the demanding suite of STAC-M3 benchmarks. Last year, Google Cloud’s partner Appsbroker showed that the same was true for calculating derivatives risk in STAC-A2 on Google Cloud. You can read about how Appsbroker built its record-breaking STAC-A2 compute cluster on Google Cloud in its blog post, or access the STAC Report directly. Here are the highlights:

Compared to all other publicly reported solutions, this solution, based on a cluster of 10 virtual machines, had:

  • The highest throughput (STAC-A2.β2.HPORTFOLIO.SPEED)
  • The fastest cold time in the large problem size (STAC-A2.β2.GREEKS.10-100k-1260.TIME.COLD)

Compared to a solution involving an 8-node, on-premises cluster (SUT ID INTC181012), this 10-node, cloud-based solution:

  • Had 5 times the maximum paths (STAC-A2.β2.GREEKS.MAX_PATHS)
  • Had 10% greater throughput (STAC-A2.β2.HPORTFOLIO.SPEED)
  • Was 18% faster in cold runs of the large problem size (STAC-A2.β2.GREEKS.10-100k-1260.TIME)
  • Was 9% faster in cold runs of the baseline problem size (STAC-A2.β2.GREEKS.TIME.COLD)

Finding market advantages with Google Cloud


Across the investment management industry, every firm is seeking many of the same competitive advantages. However, finding unique opportunities and managing larger and larger data sets is becoming a major strain. Cloud is fundamentally changing how quants tackle the problem while empowering them to manage risk and generate higher returns.

Building on-premises computing clusters with tens or hundreds of thousands of cores and petabytes of storage requires huge up-front investments and lead time measured in months or years. Google Cloud makes the same scale available to its customers, provisioned on demand and paid per use. More importantly, the elasticity of cloud resources enables agility that is simply not available in a fixed data center cluster – the agility to explore, experiment, iterate, and respond to markets faster than before.

Scaling out to tens of thousands of cores in minutes and then removing the resources immediately not only changes the speed at which questions can be answered; it encourages different and more frequent questions, asked simultaneously on many independent clusters, free from the constraints of fixed on-premises hardware.

It is this flexibility and power that enables financial services firms to leverage larger data sets and get results, backtest, research, and analyze large amounts of data, faster and whenever they need it.

Download our whitepaper to learn more about our latest STAC-M3 tick history analytics benchmark results and how to optimize cloud infrastructure for high-speed market data analysis.

Case Study

Lending DocAI Shortens Borrowers’ Journey on Roostify

11821

Of your peers have already read this article.

3:00 Minutes

The most insightful time you'll spend today!

Google Cloud's Lending DocAI automated Roostify's document processing for home application with multi-language support, allowing the provider of enterprise cloud apps for mortgage and home lenders manage upto thousands of borrowers on daily basis.

The home lending journey entails processing an immense number of documents daily from hundreds of thousands of borrowers. Currently, home lending document processing relies on some outdated digital models and a high dependency on manual labor, resulting in slow processing times and higher origination costs. Scaling a business that sorts through millions of documents daily, while increasing efficacy and accuracy, is no small feat. When it comes to applying for a mortgage loan, consumers expect a digital experience that’s as good as the in-person one. Roostify simplifies the home lending journey for lenders and their customers.

No time to spare: Overcoming document processing challenges with AI

Roostify provides enterprise cloud applications for mortgage and home lenders. In order to empower its customers to deliver a better, more personalized lending experience, they needed to automate and scale their in-house document parsing functionality.

As a key component of its document intelligence service, Roostify is leveraging Google Cloud’s Lending DocAI machine learning platform to automate processing documents required during a home loan application process, such as tax returns or bank statements with multi-language support. This partnership delivers data capture at scale, enabling Roostify customers to automatically identify document types from the uploaded file and to extract relevant entities such as wages, tax liabilities, names, and ID numbers for further processing, and make things move faster in the cumbersome lending process. 

Roostify’s solutions leverage Google Cloud’s Lending DocAI, which is built on the recently announced Document AI platform, a unified console for document processing. Customers can easily create and customize all the specialized parsers (e.g., mortgage lending documents and tax returns parsers) on the platform without the need to perform additional data mapping or training. All Google Cloud’s specialized parsers are fine-tuned to achieve industry-leading accuracy, helping customers and partners confidently unlock insights from documents with machine learning. Learn more about the solution from the GA launch blog and the overview video.

Integrating Lending DocAI’s intelligent document processing capabilities into the Roostify platform means more innovation for their customers and tangible results: faster loan processing times, fewer document intake errors, and lower origination costs. Additional support in Google Lending DAI for other languages and more documents like global Know Your Customer (KYC) documents or payroll reports is in the near future.

Full integration of AI solutions

Working together with Roostify’s platform team, we were able to help them solve their document processing challenge through integration of various GCP products such as Lending DocAI (LDAI), Data Loss Prevention (DLP) for redacting sensitive data, BigQuery for data warehousing and analytics, and Firestore for API status. To make it very safe and secure, all data was encrypted end-to-end at Rest and in Transit. LDAI won’t require any training data to process. It is an easy plug and play API.

Here is a sneak peek in the high level deployment architecture for LDAI in Roostify environment:

LDAI in Roostify.jpg

Here are the steps for processing data:

  1. Receives document processing request from the client.
  2. API Function directs requests to the pre-processing service. For Async requests a processing ID is generated and returned to the caller.
  3. Pre-processing service sends the request for further processing (Long/short PDF conversion), calling other microservices and receives back the responses. Any error in the response received is then sent to the response processing service. 
  4. If the response is synchronous, the pre-processing service directs it to the LDAI Invoker service. 
    1. If the response is asynchronous, the pre-processing service feeds it into the Cloud Pub/Sub service.
  5. Cloud Pub/Sub service feeds the response back to the LDAI Invoker service.
  6. LDAI Invoker service routes the request to the Google LDAI API for classification if there are multiple pages in the document.
  7. Document will be split based on LDAI response and then saved in a GCS bucket for temporary storage.
  8. LDAI entity interface for single page processing and then LDAI Invoker sends LDAI results to LDAI Response Processing
  9. If a request is a synchronous request the LDAI Response Processor sends results to the API Function so that it can complete the synchronous call and respond to the rConnect caller.
    1. If the request is an asynchronous request the LDAI Response Processor will respond to the caller’s webhook and complete the transaction.
  10. Finally, Data stored in the GCP bucket will be deleted.

All the responses that come from the LDAI API can optionally feed into BigQuery via the Response Processor, after parsing it through Data Loss Prevention (DLP) API to redact the PII/sensitive information.  Throughout the processing of both asynchronous and synchronous requests all transactions are logged using Cloud Logging.  For asynchronous transactions, the state is maintained throughout the process using Cloud Firestore.

Roostify currently uses this technology to power two different solutions: Roostify Document Intelligence and Roostify Beyond™. Roostify Document Intelligence is a real-time document capture, classification, and data extraction solution built for home lenders. It ingests documents uploaded by borrowers and loan officers, identifies the relevant documents, and extracts and classifies key information. Roostify Document Intelligence is available as a standalone API service to any home lender with any digital lending infrastructure already in place. 

Roostify Beyond™ is a robust suite of AI-powered solutions that enables home lenders to create intelligent experiences from start to close. It combines powerful data, insightful analytics, and meaningful visualization to streamline the underwriting process. Roostify Beyond™ is currently available only to Roostify customers as part of an Early Adopter program and will be rolled out to the market later this year.

field confidence level.jpg
Lenders can set the desired field confidence level. An extracted field that does not meet the set field confidence will display a warning indicator to borrowers asking them to validate the uploaded document.
Beyond algorithms.jpg
If the Beyond algorithms aren’t sure about the document (i.e., with lower confidence in the classification result than that set by the admin), the user sees a message asking them to validate the task.

Through this partnership, Roostify has enabled its customers to adopt a data-first approach to their home lending processes, which will lead to improved user experiences and significantly reduced loan processing times.

Fast track end-to-end deployment with Google Cloud AI Services (AIS)

Google AIS (Professional Services Organization), in collaboration with our partner Quantiphi, helped Roostify deploy this system into production and fast-tracked the development multifold to generate the final business value.

The partnership between Google Cloud and Roostify is just one of the latest examples of how we’re providing AI-powered solutions to solve business problems.

Blog

Google’s Research and Data Insights Solutions to Power Drug Development and Clinical Research

3712

Of your peers have already read this article.

1:30 Minutes

The most insightful time you'll spend today!

Explore how Google's latest research and data insights solutions are based on 'three' key functionalities to empower researchers' find relevant answers and work collaboratively, without downtime.

In order to be successful, research needs to be replicable, so scientists can build on past work and insights. However, an article in Nature warned that as much as 50% of published drug development research could not be reproduced in subsequent trials. As a result, promising drug candidates sometimes led to disappointment, as well as wasted time and money, when key findings could not be replicated. 

The shift to cloud computing helps solve this problem because it allows researchers to use open-source tools that work across platforms. As demand for cloud computing rises, our customers have asked us for more ready-made solutions to assure reproducibility of results by their collaborators, regardless of the platform they are using. They asked for secure and effective collaboration tools as well as faster time-to-insight from any type of data.

We listened. Google’s new research and data insights solution includes three sets of functionalities to address these key challenges. Each can be activated on demand and may be eligible for subscription pricing. “HPC in a box” offers abstract complexity to run high performance computing (HPC) workloads by automatically managing your cluster in the most effective manner. It integrates seamlessly with some of the industry’s most-used schedulers like Slurm and PBS. It makes it easier than ever to answer bigger questions faster by accessing Google’s fast, powerful hardware like TPUs and GPUs, all for one predictable flat fee for eligible workloads. Healthcare Innovation Hub provides healthcare-specific functionality to help ingest, aggregate, and de-identify any type of healthcare data in its original format. It unlocks cross-modality analysis and collaboration and empowers researchers with harmonization tools to overcome healthcare interoperability issues. Google Cloud Real-World Insights (formerly FDA MyStudies) accelerates and streamlines drug development and clinical trials to address urgent medical challenges with reproducible results.

The solution enables researchers to ask new questions, get answers more quickly, and work more collaboratively–with no wait times or down times. Institutions can scale to more ambitious projects and generate actionable, real-time insights from any data source–all while staying within budget.

Many top research centers have already found it faster and more cost effective to shift from downloading and storing data on their own servers to storing and analyzing data on Google Cloud. Here are some of the real-world projects already yielding breakthroughs:

Our partners, such as Atos, Burwood, Omnibond, Mavenwave, Quantiphi, and Deloitte, can help you first design and develop, then install and implement your own solution, including training. To assess your institution’s needs and develop a customized plan for your next-generation research solution with research and insights, contact our sales team.

Blog

Rethinking retail with Google Cloud Retail Search

2813

Of your peers have already read this article.

3:30 Minutes

The most insightful time you'll spend today!

With Cloud Retail Search, improving the shopping experience is easier for retailers. Read this blog to know how Cloud Retail Search is a great solution to help reduce churn, and improve conversion and retention.

Cloud Retail Search, part of Discovery Solutions For Retail portfolio, helps retailers significantly improve the shopping experience on their digital platform with ‘Google-quality’ search. Cloud Retail Search offers advanced search capabilities such as better understanding user intent and self-learning ranking models that help retailers unlock the full potential of their online experience.

Google Cloud’s Discovery Solutions For Retail are a set of services that can help retailers improve their digital engagement and are offered as part of our industry solutions.

Executive Summary

Retailers are always working on trying to keep up with the ever changing consumer expectations and trying to forecast the next trend that can impact sales and revenue.

The pandemic brought its own (and largely new) set of challenges which further complicated the issue over the last two years. The retailers were forced to adapt to the new consumer (low physical touch) behavior in which the browsing and product research was largely digital (endless aisle) and accelerated other trends such as buy online and pick up in stores (BOPIS), curbside pick up and pick up lockers. According to a McKinsey Global Survey from early last year, the pandemic has accelerated the pace of digital transformation by several years.

The National Retail Federation (NRF) estimates that retail sales are expected to grow between 6% and 8% in 2022 (slower growth rate than in 2021), as consumers spend more on services instead of goods, deal with inflation and higher food & gas prices due to geopolitical disruptions in the world.

And the competition continues to be fierce as ever. Amazon continues its dominance in the U.S. retail world and new PYMNTS data shows that Amazon’s share of US Ecommerce sales hit an all-time high of 56.7% in 2021.

Customers now have more choices than ever on how they want to engage with the retailers, where they want to spend the money and make their purchase. They also have increased expectations from the retailers around providing a high quality product discovery experience, which is forcing the retailers to invest heavily on improving customer engagement on their digital platforms to boost conversion rate and overall customer loyalty.

This is where Retail Search can help by providing an enhanced search experience that uses Google-quality search models to understand the customer intent and takes into account the retailer’s first party data (such as promotions, available inventory and price) for ranking results.

How is Google Cloud Retail Search Different

The Ecommerce platform on-site search use case is not new and retailers have been trying to solve it effectively for the last two decades. Most retailers recognize that search is a critical service on the platform and have spent countless resources to improve and fine tune it over the years. Yet the challenge remains. According to a Baymard Institute study in late as 2019, 61% of sites still required their users to search by the exact product type jargon the site uses.

However, users now expect the same robust and intuitive search features as is offered by Google.com and other popular web platforms, who seem to have the uncanny ability to intelligently interpret and yield relevant results to complex search queries.

Google’s decades of experience and research in search technology benefits Cloud Retail Search solution and that is what differentiates it from the competition.

  • Advanced Query Understanding: Retail Search can provide more relevant results for the same query due to better query understanding features and knowing when to broaden or narrow the query results. While most search engines still rely largely on keyword based or matching tokens results, Retail Search has the advantage of being able to leverage Google search algorithms to return highly relevant results for product listings and category pages.
  • Semantic Search: Intent recognition is a key requirement for semantic search and identifying what the customers mean when they enter the query is a key strength of Retail Search. This is critical for retailers since this has a direct impact on Clickthrough rate, Conversion rate and the Bounce Rate.
  • Personalized Search Results: Another key differentiator for Retail Search is its ability to leverage user interaction data and ranking models to provide hyper personalized search results. Retailers are able to optimize search performance to deliver desired outcomes: better engagement, revenue, or conversions.
  • Self-Learning and Self-Managed Solution: Retail Search models get better over time because of the self-learning capabilities built into the solution. In addition, the service is fully managed, which saves precious resources needed to keep it running and managing its set up.
  • Strong Security Controls: The service runs on Google Cloud and follows security best practices to keep our customers’ data secure. Google never shares model weights or customer data across customers using the Retail API or other Discovery Solution products. For more details about this data use, see a description of Retail API data use.

High Level Conceptual View

Here is a simplified high-level view of Retail Search API. Retailers can call the API for the given search query and get back the results which can then be displayed on their digital properties.

The returned results contains two types of information:

  • Search results: Query search results including product listings and category pages based on advanced query understanding and semantic search.
  • Dynamic faceted search attributes: Faceted Search is a feature that allows further refinement of the search by providing ways to apply additional filters while returning results.

Retail Search needs the following datasets as input to train its machine learning models for search:

  • Product Catalog: Information about the available products including product categories, product description, in-stock availability, and pricing.
  • User Events: This is the clickstream data that contains user interaction information such as clicks and purchases.
  • Inventory / Pricing Updates: Incremental updates to in-stock availability and pricing as that information is updated.

(Keeping the product catalog up to date and recording user events successfully is crucial for getting high-quality results. Set up Cloud Monitoring alerts to take prompt action in case any issues arise).

Retailers also have the ability to set up business/config rules to customize their search results and optimize for business revenue goals such as Clickthrough rate, Conversion rate, Average size order etc.

How to get started

Retail Search is generally available now and anyone with a Google cloud account can access it. If you don’t already have an account, you can start with a trial account for free here.

Establish a Success Criteria: It’s important to establish a success criteria for measuring the effectiveness of Retail search. Get a consensus on which factor(s) you want to include in scope for measuring the effectiveness of Retail Search. This could include one or two from the following: Search Conversion Rate, Search Average Order Value, Search Revenue Per Visit and Null Search Rate (No Results Found).

  • Measuring Performance: Retail dashboards provide metrics to help you determine how incorporating the Retail API is affecting the results. You can view summary metrics for your project on the Analytics tab of the Monitoring & Analytics page in Cloud Console.
  • Set up A/B Experiments: To measure the performance of Retail Search with another search solution, you can set up A/B tests using a third-party experiment platform such as Google Optimize.

Summary:

As retailers try to navigate through the post-pandemic world where supply chain failures and digital transformation acceleration are major focus areas, they now also have to keep a close eye on the recent geopolitical challenges resulting in rising inflation and costs.

While we can all agree that in-store shopping will continue to be a major source of revenue, it is also important for retailers to tweak the in-store experience for the digital world. Trends such as buy online and pick up in stores (BOPIS), curbside pick up and pick up lockers are here to stay.

Given all the above, consumer engagement and digital experience is more important now than ever before. The cost of search abandonment is way too high and has both short and longer term impact. Retail Search is a great solution to help reduce churn, improve conversion and retention. It provides Google-quality search models to help understand customer intent and the retailers have the ability to set up business/config rules to optimize search results for business revenue goals such as Clickthrough rate, Conversion rate and Average size order.

More Relevant Stories for Your Company

Case Study

Tencent Africa Cuts Cost and Improves Stability with Google Cloud

As one of Africa’s leading technology companies, Tencent Africa is responsible for WeChat operations on the continent. WeChat Africa has successfully launched a number of features including the WeChat Wallet, a mobile payment service for smartphones, which enables seamless and secure transactions for friends to send money to one another,

Blog

Know the Leaders of Google Cloud Public Sector Community

At Google Cloud, being a strategic partner is part of our DNA. Whether it’s listening closely to our customers, helping to build team skills for innovation or simply being there (since we know the cloud is 24/7), we get excited about working hands-on with customers to deliver new solutions.  As

Case Study

FFF Enterprises See 80% Improvements in Speed at Lower Cost by Moving SAP Data to Google Cloud

FFF Enterprises, a pharmaceutical distributor of lifesaving biopharma products, vaccines and plasma products deployed SAP on Google Cloud to leverage its ability to scale server demands, reduce costs and improve speed and performance. Learn how the migration helps FFF empower healthcare to care!

Case Study

IT Team Figures Out Easiest Way to Build Data Pipelines and Create ML Models

Building a strong brand in today's hyper-competitive business environment takes vision. It also requires a flexible, easily managed approach to digital asset management (DAM), so marketing professionals and other stakeholders can easily share, store, track, and manipulate assets to build the brand. Many of today's leading companies, including JetBlue, Slack,

SHOW MORE STORIES