Google Research: Themes from 2021 and Beyond - Build What's Next
Trend Analysis

Google Research: Themes from 2021 and Beyond

2875

Of your peers have already read this article.

4:00 Minutes

The most insightful time you'll spend today!

Read the blogpost to catch up on Google's research on AI and five emerging ML-related trends that are poised to redefine the way systems interact around the world with new product features and accomplish more with ML models.


Posted by Jeff Dean, Senior Fellow and SVP of Google Research, on behalf of the entire Google Research community

Over the last several decades, I’ve witnessed a lot of change in the fields of machine learning (ML) and computer science. Early approaches, which often fell short, eventually gave rise to modern approaches that have been very successful. Following that long-arc pattern of progress, I think we’ll see a number of exciting advances over the next several years, advances that will ultimately benefit the lives of billions of people with greater impact than ever before. In this post, I’ll highlight five areas where ML is poised to have such impact. For each, I’ll discuss related research (mostly from 2021) and the directions and progress we’ll likely see in the next few years.


· Trend 1: More Capable, General-Purpose ML Models
· Trend 2: Continued Efficiency Improvements for ML
· Trend 3: ML Is Becoming More Personally and Communally Beneficial
· Trend 4: Growing Benefits of ML in Science, Health and Sustainability
· Trend 5: Deeper and Broader Understanding of ML

Trend 1: More Capable, General-Purpose ML Models


Researchers are training larger, more capable machine learning models than ever before. For example, just in the last couple of years models in the language domain have grown from billions of parameters trained on tens of billions of tokens of data (e.g., the 11B parameter T5 model), to hundreds of billions or trillions of parameters trained on trillions of tokens of data (e.g., dense models such as OpenAI’s 175B parameter GPT-3 model and DeepMind’s 280B parameter Gopher model, and sparse models such as Google’s 600B parameter GShard model and 1.2T parameter GLaM model). These increases in dataset and model size have led to significant increases in accuracy for a wide variety of language tasks, as shown by across-the-board improvements on standard natural language processing (NLP) benchmark tasks (as predicted by work on neural scaling laws for language models and machine translation models).

Many of these advanced models are focused on the single but important modality of written language and have shown state-of-the-art results in language understanding benchmarks and open-ended conversational abilities, even across multiple tasks in a domain. They have also shown exciting capabilities to generalize to new language tasks with relatively little training data, in some cases, with few to no training examples for a new task. A couple of examples include improved long-form question answering, zero-label learning in NLP, and our LaMDA model, which demonstrates a sophisticated ability to carry on open-ended conversations that maintain significant context across multiple turns of dialog.

A dialog with LaMDA mimicking a Weddell seal with the preset grounding prompt, “Hi I’m a weddell seal. Do you have any questions for me?” The model largely holds down a dialog in character.
(Weddell Seal image cropped from Wikimedia CC licensed image.)

Transformer models are also having a major impact in image, video, and speech models, all of which also benefit significantly from scale, as predicted by work on scaling laws for visual transformer models. Transformers for image recognition and for video classification are achieving state-of-the-art results on many benchmarks, and we’ve also demonstrated that co-training models on both image data and video data can improve performance on video tasks compared with video data alone. We’ve developed sparse, axial attention mechanisms for image and video transformers that use computation more efficiently, found better ways of tokenizing images for visual transformer models, and improved our understanding of visual transformer methods by examining how they operate compared with convolutional neural networks. Combining transformer models with convolutional operations has shown significant benefits in visual as well as speech recognition tasks.

The outputs of generative models are also substantially improving. This is most apparent in generative models for images, which have made significant strides over the last few years. For example, recent models have demonstrated the ability to create realistic images given just a category (e.g., “irish setter” or “streetcar”, if you desire), can “fill in” a low-resolution image to create a natural-looking high-resolution counterpart (“computer, enhance!”), and can even create natural-looking aerial nature scenes of arbitrary length. As another example, images can be converted to a sequence of discrete tokens that can then be synthesized at high fidelity with an autoregressive generative model.

Example of a cascade diffusion models that generate novel images from a given category and then use those as the seed to create high-resolution examples: the first model generates a low resolution image, and the rest perform upsampling to the final high resolution image.
The SR3 super-resolution diffusion model takes as input a low-resolution image, and builds a corresponding high resolution image from pure noise.

Because these are powerful capabilities that come with great responsibility, we carefully vet potential applications of these sorts of models against our AI Principles.

Beyond advanced single-modality models, we are also starting to see large-scale multi-modal models. These are some of the most advanced models to date because they can accept multiple different input modalities (e.g., language, images, speech, video) and, in some cases, produce different output modalities, for example, generating images from descriptive sentences or paragraphs, or describing the visual content of images in human languages. This is an exciting direction because like the real world, some things are easier to learn in data that is multimodal (e.g., reading about something and seeing a demonstration is more useful than just reading about it). As such, pairing images and text can help with multi-lingual retrieval tasks, and better understanding of how to pair text and image inputs can yield improved results for image captioning tasks. Similarly, jointly training on visual and textual data can also help improve accuracy and robustness on visual classification tasks, while co-training on image, video, and audio tasks improves generalization performance for all modalities. There are also tantalizing hints that natural language can be used as an input for image manipulation, telling robots how to interact with the world and controlling other software systems, portending potential changes to how user interfaces are developed. Modalities handled by these models will include speech, sounds, images, video, and languages, and may even extend to structured data, knowledge graphs, and time series data.

Example of a vision-based robotic manipulation system that is able to generalize to novel tasks. Left: The robot is performing a task described in natural language to the robot as “place grapes in ceramic bowl”, without the model being trained on that specific task. Right: As on the left, but with the novel task description of “place bottle in tray”.

Often these models are trained using self-supervised learning approaches, where the model learns from observations of “raw” data that has not been curated or labeled, e.g., language models used in GPT-3 and GLaM, the self-supervised speech model BigSSL, the visual contrastive learning model SimCLR, and the multimodal contrastive model VATTSelf-supervised learning allows a large speech recognition model to match the previous Voice Search automatic speech recognition (ASR) benchmark accuracy while using only 3% of the annotated training data. These trends are exciting because they can substantially reduce the effort required to enable ML for a particular task, and because they make it easier (though by no means trivial) to train models on more representative data that better reflects different subpopulations, regions, languages, or other important dimensions of representation.

All of these trends are pointing in the direction of training highly capable general-purpose models that can handle multiple modalities of data and solve thousands or millions of tasks. By building in sparsity, so that the only parts of a model that are activated for a given task are those that have been optimized for it, these multimodal models can be made highly efficient. Over the next few years, we are pursuing this vision in a next-generation architecture and umbrella effort called Pathways. We expect to see substantial progress in this area, as we combine together many ideas that to date have been pursued relatively independently.

Pathways: a depiction of a single model we are working towards that can generalize across millions of tasks.

Explainer

Delivering 10X Improvement to Risk and Regulatory Reporting Through Cloud and AI

4984

Of your peers have already read this article.

2:30 Minutes

The most insightful time you'll spend today!

Respond quickly to the ever-changing risk and regulatory landscape by adopting cloud and machine learning, derive new insights and allow risk management to become more embedded into operational processes.

Enterprise agility and the ability to innovate, adapt and respond quickly to the ever-changing risk and regulatory landscape is no longer a choice, but the cornerstone of successful digital transformation and commercial growth. Traditional access to and ways of managing data invariably create challenges in dealing with multiple data repositories, reconciliations, fire-drills, etc.

In response, the move to cloud is increasing significantly. It enables risk analytics and regulatory reporting at scale in a secure environment with data storage, management and encryption capabilities as a standard. In addition, as regulatory reporting requirements become more granular, machine learning can help facilitate new insights and allow for risk management to become more embedded into operational processes.

This webinar will address the day-to-day challenges in risk management and regulatory compliance, while also exploring how technological innovations can provide massive improvements and potential.

Key themes

  • Real-life data challenges in the eyes of risk managers: can compliance, fraud detection and identifying liquidity positions be improved through the use of AI?
  • Innovative approaches to streamline regulatory reporting to derive deeper customer insights from data at the moment of truth.
  • Reimagining operations: how to modernise the data infrastructure to accommodate data explosion, drive flexibility and deliver a more cost effective outcome.
VIEW WEBINAR

4937

Of your peers have already watched this video.

21:10 Minutes

The most insightful time you'll spend today!

How-to

Learn Modern App Development Practices to Ship Software Faster

Cloud-native, Kubernetes, Serverless have been the hottest and most widely discussed topics given the velocity and agility benefits.

Learn more about how you can leverage these modern app development practices to ship software faster, while reducing costs and improving security and compliance.

Learn how Google Cloud lets you modernize existing applications at your own pace using these technologies. Regardless of where you are in your app modernization journey, watch this video to learn how to improve the developer experience and deliver software faster.

Case Study

Mercari’s Big Leap: Supercharging Growth with Google Cloud’s BigQuery

1118

Of your peers have already read this article.

6:30 Minutes

The most insightful time you'll spend today!

Explore how Mercari, an online marketplace, accelerated its growth using BigQuery and GrowthLoop, unlocking unprecedented customer understanding and marketing personalization. Learn more...

When peer-to-peer marketplace Mercari came to the US in 2014, it had its work cut out for it. Surrounded by market giants like eBay, Craigslist, and Wish, Mercari needed to carve out an approach to compete for new users. Furthermore, Mercari wanted to build a network for buyers and sellers to return to, rather than a site for individual specialty purchases. 

As an online marketplace that connects millions of people across the U.S. to shop and sell items of value no longer being used, Mercari is built for the everyday shopper and casual seller. Two teams, Machine Learning (ML) team and Marketing Technology specialists, both led by Masumi Nakamura, Mercari VP of Engineering, saw an opportunity to supercharge Mercari’s growth in the US by leveraging their first-party data in BigQuery and connecting predictive models built in Google Cloud directly to marketing channels, such as churn predictions and item recommendations for email campaigns, and LTV predictions to optimize paid media. Churn predictions could be used to target marketing communications, and item recommendations could be used to personalize the content of those communications at the user level. By fully utilizing cloud computing services, they could grow sustainably and flexibly, focusing their team’s efforts where they belonged — user understanding and personalized marketing.

In 2018, the Mercari US team engaged GrowthLoop, formerly Flywheel Software, experts in leveraging first-party customer data for business growth. Working exclusively in Google Cloud and BigQuery, GrowthLoop helped Masumi transform Marketing Technology at Mercari in the US.

Use cases: challenges

Masumi and the ML team’s primary goal aimed to reduce churn across buyers and sellers. Customers would make an initial purchase, but repurchase and resale rates were lower than the team hoped for. The ML team, led by Masumi, was confident that if they could get customers to make a second and third purchase, they could drive strong lifetime value (LTV). 

Despite the team’s robust data science capabilities and investments in a data warehouse (BigQuery), they were missing the ability to streamline efforts for efficient audience segmentation and targeting. Like most companies looking to utilize data for marketing, the team at Mercari had to engage with engineering in order to build out customer segments for campaign launches and testing. From start to finish, launching a single campaign could take three months. 

In short, Mercari needed a way to speed up the process across the teams at Mercari. How could they turn the team’s predictions into active marketing experiments with greater velocity and agility?

Solution: BigQuery and GrowthLoop supercharge growth across the customer lifecycle

With their strong data engineering foundation and BigQuery already in place, the Mercari team began addressing their needs step-by-step. First, they used predictions to identify retention features, then built out initial segment definitions based on those features. From there, the team designed and launched experiments and measured their performance, refining as they went. By providing Mercari’s marketing team with the ability to build their own customer lists that leveraged predictive models without requiring continuous support from other teams’ data engineers and business intelligence analysts, GrowthLoop enabled them to address churn and acquisition with a single, self-serve solution.

https://storage.googleapis.com/gweb-cloudblog-publish/images/1_Mercari.max-2000x2000.png
Mercari Architecture Diagram on Google Cloud with GrowthLoop

The dynamic duo: GrowthLoop and BigQuery

  • Customer 360: GrowthLoop enabled Mercari to combine their data sources into a single view of their customers in BigQuery, then connected them to marketing and sales channels via GrowthLoop’s platform. Notably, Mercari is able to leverage its own complex data model, which was ideal for a two-sided marketplace. This is shown in the “Collect & Transform” stage in the architecture diagram above.
  • Predictive models: GrowthLoop activated predictions that had been snapshotted by Mercari’s team in BigQuery. The Mercari ML team used Jupyter notebooks offering part of Google Vertex AI Workbench to build user churn and customer lifetime value (CLTV) prediction models, then productionized them using Cloud Composer to deploy Airflow DAGs, which wrote the predictions back to BigQuery for targeting, and triggered exports to destination channels using Pub/Sub. This is shown in the “Intelligence” stage of the architecture diagram above.
  • Extensible measurement and data visualization: Since GrowthLoop writes all audience data back to BigQuery, the Mercari analytics team can conduct performance analysis on metrics from revenue to retention. They are able to use GrowthLoop’s performance visualization in-app, but they are also able to create custom data visualizations with Looker Studio. This is also shown in the “Intelligence” stage of the architecture diagram.
  • Seamless routing and activation: With GrowthLoop’s audience platform connected directly to Customer 360 and the predictive model’s results in BigQuery, the marketing team is able to launch and sync audiences and their personalization attributes across all of Mercari’s major marketing, sales and product channels, such as Braze, Google Ads and other destinations. This is shown in the “Routing” and “Activate” stage of the architecture diagram.

“Being able to measure what you’re doing – that results-based orientation – is key. The thing that I like most about GrowthLoop is that you brought a really fundamental way of thinking which was very feedback-based and open to experimenting but within reason. With other products, that feedback loop isn’t so built in that it’s very easy to get lost.”– Masumi Nakamura, VP of Engineering at Mercari

Predictive modeling puts the burn on churn

https://storage.googleapis.com/gweb-cloudblog-publish/images/2_Mercari.max-1200x1200.png
Machine Learning model visualization as a decision tree to predict customer churn

In collaboration with GrowthLoop, Mercari began analyzing user data in BigQuery via Vertex AI Workbench to identify patterns across churned customers. The teams evaluated a range of attributes like the customer acquisition channel, categories browsed or purchased from, and whether or not they had any saved searches while shopping. Comparing various models and performance metrics, the teams selected the best model for accurately predicting when a buyer or seller was at risk to churn. For sellers, they evaluated audience members by the time elapsed since their last sale – for buyers, the time since their last purchase. 

These churn prediction scores could then be applied to data pipelines that would feed into GrowthLoop’s audience builder. Audience members with a high likelihood to churn would be segmented into their own group and from there, Mercari could target those users with relevant paid media and email campaigns. 

By partnering with GrowthLoop, Mercari was able to simultaneously bridge the gap between the data and marketing teams – and reduce the time between segmentation and campaign launch from months to just a few days.

https://storage.googleapis.com/gweb-cloudblog-publish/images/3_Mercari.max-1600x1600.png
A view of the user-friendly the GrowthLoop first-party data platform to build an audience

“One of the big areas of benefit of working with GrowthLoop was the increased integration of marketing channels such as the CRM, User Acquisition, as well as more traditional marketing channels.” – Masumi Nakamura, VP of Engineering at Mercari

Creating the audience within the audience

Once the team had successfully created a model to predict churn across buyers and sellers, Mercari needed to launch retargeting campaigns to measure their ability to reduce churn. Each of their ongoing experiments features tailored segments along with automatic A/B testing. With analytics and activation all under one roof, the marketing team at Mercari could craft audiences and begin measuring the impact of their targeted campaigns. Since starting their work with GrowthLoop, the Mercari team has created over 120 audiences.

“Our marketing teams are more sophisticated with in-house knowledge, but GrowthLoop provides a more user-friendly way to build audiences for campaigns.” – Masumi Nakamura, VP of Engineering, Mercari

https://storage.googleapis.com/gweb-cloudblog-publish/images/4_Mercari.max-1600x1600.png
A summary view of the central “Audience Hub” on the GrowthLoop first party data platform

“GrowthLoop brings a very fundamental way of thinking about problems, including experimentation…. The ability to organize experiments and results was key. The number of variables is too high for most people without good organization.”– Masumi Nakamura, VP of Engineering at Mercari 

Making segmentation smarter

Mercari’s first audiences leveraging GrowthLoop were sent to Braze to supercharge email campaigns and coupons with churn predictions and automated campaign performance evaluations. Then, Mercari shifted its focus to Facebook for paid media retargeting, using GrowthLoop’s lifecycle segmentation framework to target customers at the right step in their user journey. Lastly, Mercari moved its focus to Google Ads, where they used GrowthLoop to implement new segmentation models based on product category propensity. Mercari had long used Google Ads for product listing ads, and with GrowthLoop, Mercari was able to define more powerful product propensity segments and measure custom incremental lift metrics.

https://storage.googleapis.com/gweb-cloudblog-publish/images/5_Mercari.max-1100x1100.png

Finding new users in the haystack

Finally, in addition to preventing churn and driving retention, the Mercari team also wanted to boost user acquisition. They were having trouble measuring performance of UA campaigns due to new iOS and Facebook data privacy restrictions that made measuring campaign attribution impossible for many users. Using the familiar stack of Vertex AI Workbench for analysis, performance analysis on campaign data in BigQuery, and Airflow DAGs deployed via Cloud Composer to productionize the data pipelines, GrowthLoop enabled the team to activate targeted campaigns based on a user’s geographical location. In this way, Mercari could make decisions about their UA campaigns using incrementality analysis between geographic regions rather than attribution data, thus preserving user privacy. 

The Mercari approach to customer data activation and acquisition

Other marketplace retailers can learn from Mercari’s successes activating data from BigQuery with GrowthLoop. Here are a few best practices to apply:

Identify your team’s needs and existing strengths

Mercari knew that their team had built out a strong foundation for data analysis within BigQuery. They also knew that their process was missing a key component that would allow them to activate that data. In order to achieve similar results, work to evaluate the strength of your team and your data – and define exactly what you aim to achieve with customer segmentation.

Partner with the right providers

With BigQuery, the Mercari team had all of their data centralized in one single location, simplifying the process for predictive modeling, segmentation, and activation. By partnering with GrowthLoop, this centralized data could be activated with ease across Mercari’s marketing teams. When evaluating providers for data warehousing, segmentation, and activation, be sure to partner with a provider that ensures you can get the most out of your data.

Know your audience

With a deeper understanding of their customers, Mercari was able to see nearly immediate value. By investing in the proper tools to accurately predict customer behavior, Mercari delivered impact in exactly the right areas. Using the data you’ve already compiled on your customers, consider partnering with a customer segmentation platform provider like GrowthLoop. In fact, Masumi went so far as to organize his Machine Learning team around these concepts: “We split the ML team into two areas – one to augment and work with GrowthLoop, the other team was to augment and orient around item data.”

https://storage.googleapis.com/gweb-cloudblog-publish/images/6_Mercari.max-1200x1200.png
Incremental lift in sales on Mercari’s platform by audience
Scale has been modified to intentionally obfuscate actual results.

How to boost growth like Mercari in three steps

Today, many leading brands leverage GrowthLoop and BigQuery to drive marketing and sales wins. Whether your company is in retail, financial services, travel, software, or another industry entirely, you can join the growing number of companies driving sustainable growth through real-time analytics by connecting BigQuery from Google Cloud to GrowthLoop. Here’s how:

If you have customer data in BigQuery…

  1. Book a GrowthLoop + BigQuery demo customized to your use cases.
  2. Link your BigQuery tables and marketing and sales destinations to the GrowthLoop platform.
  3. Launch your first GrowthLoop audience in less than one week.

If you are getting started with BigQuery…

  1. Get a Data Strategy Session with a GrowthLoop Solutions Architect at no cost.
  2. Use our Quick Start Program to get started with BigQuery in 4 to 8 weeks.
  3. Launch your first GrowthLoop audience in less than one week thereafter.

GrowthLoop and Google: Better together

The key question for many marketers today is, “How do you best leverage all you know about your customers to drive more intelligent and effective marketing engagement?” When Mercari set out to answer this question in 2019, they applied an innovative BigQuery data strategy that leveraged machine learning models. However, they achieved remarkable marketing results because they were among the first companies to discover and apply GrowthLoop to enable the marketing team to launch audiences with a first party data platform directly connected to their datasets and predictions in BigQuery. This greatly accelerated the design-launch-measure feedback loop to generate repeatable growth in customer lifetime value.

The Built with BigQuery advantage for ISVs and Data Providers

Google is helping companies like GrowthLoop build innovative applications on Google’s data cloud with simplified access to technology, helpful and dedicated engineering support, and joint go-to-market programs through the Built with BigQuery initiative. Participating companies can: 

  • Accelerate product design and architecture through access to designated experts who can provide insight into key use cases, architectural patterns, and best practices. 
  • Amplify success with joint marketing programs to drive awareness, generate demand, and increase adoption.

BigQuery gives ISVs the advantage of a powerful, highly scalable data warehouse that’s integrated with Google Cloud’s open, secure, sustainable platform. And with a huge partner ecosystem and support for multi-cloud, open source tools and APIs, Google provides technology companies the portability and extensibility they need to avoid data lock-in.

Click here to learn more about Built with BigQuery.


We thank the Mercari, GrowthLoop and Google Cloud team members who collaborated on the blog:
Mercari: Masumi Nakamura, VP of Engineering
GrowthLoop: Julia Parker, Product Marketing Manager; Alex Cuevas, Head of Analytics
Google: Sujit Khasnis, Solutions Architect

3127

Of your peers have already watched this video.

17:00 Minutes

The most insightful time you'll spend today!

Blog

How FLYR and Google Cloud Help Airlines Forecast Demand and Set Prices

FLYR Labs is an international team of industry experts and specialists in revenue management that works to bring in intelligence to the airlines companies. FLYR uses machine learning and AI to help predict demand and optimize price so that every airline is operating its complete capacity. Watch the video from Architecting with Google Cloud to deep-dive into a use case with FLYR involving the use of historic data, competitors data and future information to build model for outputting demand, set prices and optimize revenue. You can can even have a quick view of the FLYR ML platform!

Blog

Predict Protein Structures with AlphaFold on Vertex AI

2818

Of your peers have already read this article.

3:30 Minutes

The most insightful time you'll spend today!

To accelerate research in the bio-pharma space, we are announcing a new Vertex AI solution that demonstrates how to use Vertex AI Pipelines to run DeepMind’s AlphaFold protein structure predictions at scale.

Today, to accelerate research in the bio-pharma space, from the creation of treatments for diseases to the production of new synthetic biomaterials, we are announcing a new Vertex AI solution that demonstrates how to use Vertex AI Pipelines to run DeepMind’s AlphaFold protein structure predictions at scale.

Once a protein’s structure is determined and its role within the cell is understood, scientists can develop drugs that can modulate the protein function based on its role in the cell. DeepMind, an AI research organization within Alphabet, created the AlphaFold system to advance this area of research by helping data scientists and other researchers to accurately predict protein geometries at scale.

In 2020, in the Critical Assessment of Techniques for Protein Structure Prediction (CASP14) experiment, DeepMind presented a version of AlphaFold that predicted protein structures so accurately, experts declared the “protein-folding problem” solved. The next year, DeepMind open sourced the AlphaFold 2.0 system. Soon after, Google Cloud released a solution that integrated AlphaFold with Vertex AI Workbench to facilitate interactive experimentation. This made it easier for many data scientists to efficiently work with AlphaFold, and today’s announcement builds on that foundation.

Last week, AlphaFold took another significant step forward when DeepMind, in partnership with the European Bioinformatics Institute (EMBL-EBI), released predicted structures for nearly all cataloged proteins known to science. This release expands the AlphaFold database from nearly 1 million structures to over 200 million structures—and potentially increases our understanding of biology to a profound degree. Between this continued growth in the AlphaFold database and the efficiency of Vertex AI, we look forward to the discoveries researchers around the world will make.

In this article, we’ll explain how you can start experimenting with this solution, and we’ll also survey its benefits, which include offering lower costs through optimized selection of hardware, reproducibility through experiment tracking, lineage and metadata management, and faster run time through parallelization.

Background for running AlphaFold on Vertex AI

Generating a protein structure prediction is a computationally intensive task. It requires significant CPU and ML accelerator resources and can take hours or even days to compute. Running inference workflows at scale can be challenging—these challenges include optimizing inference elapsed time, optimizing hardware resource utilization, and managing experiments.Our new Vertex AI solution is meant to address these challenges.

To better understand how the solution addresses these challenges, let’s review the AlphaFold inference workflow:

  1. Feature preprocessing. You use the input protein sequence (in the FASTA format) to search through genetic sequences across organisms and protein template databases using common open source tools. These tools include JackHMMER with MGnify and UniRef90, HHBlits with Uniclust30 and BFD, and HHSearch with PDB70. The outputs of the search (which consist of multiple sequence alignments (MSAs) and structural templates) and the input sequences are processed as inputs to an inference model. You can run the feature preprocessing steps only on a CPU platform. If you’re using full-size databases, the process can take a few hours to complete.
  2. Model inference. The AlphaFold structure prediction system includes a set of pretrained models, including models for predicting monomer structures, models for predicting multimer structures, and models that have been fine-tuned for CASP. At inference time, you independently run the five models of a given type (such as monomer models) on the same set of inputs. By default, one prediction is generated per model when folding monomer models, and five predictions are generated per model when folding multimers. This step of the inference workflow is computationally very intensive and requires GPU or TPU acceleration.
  3. (Optional) Structure relaxation. In order to resolve any structural violations and clashes that are in the structure returned by the inference models, you can perform a structure relaxation step. In the AlphaFold system, you use the OpenMM molecular mechanics simulation package to perform a restrained energy minimization procedure. Relaxation is also very computationally intensive, and although you can run the step on a CPU-only platform, you can also accelerate the process by using GPUs.

The Vertex AI solution

The AlphaFold batch inference with the Vertex AI solution lets you efficiently run AlphaFold inference at scale by focusing on the following optimizations:

  • Optimizing inference workflow by parallelizing independent steps.
  • Optimizing hardware utilization (and as a result, costs) by running each step on the optimal hardware platform. As part of this optimization, the solution automatically provisions and deprovisions the compute resources required for a step.
  • Describing a robust and flexible experiment tracking approach that simplifies the process of running and analyzing hundreds of concurrent inference workflows.

The following diagram shows the architecture of the solution.


The solution encompasses the following:

A strategy for managing genetic databases. The solution includes high-performance, fully managed file storage. In this solution, Cloud Filestore is used to manage multiple versions of the databases and to provide high throughput and low-latency access.

An orchestrator to parallelize, orchestrate, and efficiently run steps in the workflow. Predictions, relaxations, and some feature engineering can be parallelized. In this solution, Vertex AI Pipelines is used as the orchestrator and runtime execution engine for the workflow steps.

Optimized hardware platform selection for each step. The prediction and relaxation steps run on GPUs, and feature engineering runs on CPUs. The prediction and relaxation steps can use multi-GPU node configurations. This is especially important for the prediction step because the memory usage is approximately quadratic with the number of residues. Therefore, predicting a large protein structure can exceed the memory of a single GPU device.

Metadata and artifact management. The solution includes management for running and analyzing experiments at scale. In this solution, Vertex AI Metadata is used to manage metadata and artifacts.

The basis of the solution is a set of reusable Vertex AI Pipelines components that encapsulate core steps in the AlphaFold inference workflow: feature preprocessing, prediction, and relaxation. In addition to those components, there are auxiliary components that break down the feature engineering step into tools, and helper components that aid in the organization and orchestration of the workflow.

The solution includes two sample pipelines: the universal pipeline and a monomer pipeline. The universal pipeline mirrors the settings and functionality of the inference script in the AlphaFold Github repository. It tracks elapsed time and optimizes compute resources utilization. The monomer pipeline further optimizes the workflow by making feature engineering more efficient. You can customize the pipeline by plugging in your own databases.

Next steps

To learn more and to try out this solution, check our GitHub repository, which contains the components and universal and monomer pipelines. The artifacts in the repository are designed so that you can customize them. In addition, you can integrate this solution into your upstream and downstream workflows for further analysis. To learn more about Vertex AI, visit our product page.

Acknowledgements

We would like to thank the following people for their collaboration: Shweta Maniar, Sampath Koppole, Mikhail Chrestkha, Jasper Wang, Alex Burdenko, Meera Lakhavani, Joan Kallogjeri, Dong Meng (NVIDIA), Mike Thomas (NVIDIA), and Jill Milton (NVIDIA).

Finally and most importantly, we would like to thank our Solution Manager Donna Schut for managing this solution from start to finish. This would not have been possible without Donna.

More Relevant Stories for Your Company

Case Study

Personalization with Recommendation AI Improves Reader Experience on Newsweek Platform

Newsweek provides the latest news, in-depth analysis, and ideas about international issues, technology, business, culture, and politics to its readers around the world. While editors pick the best articles to display on the home page and topic pages, it is also critical for Newsweek to offer a personalized experience by

Blog

DocAI Lowers Customer’s Document Processing Cost by 60 Percent. Learn How

Some of the most important data at your company isn’t living in databases, but in documents, and most business processes begin, involve or end with a document.  Yet most companies are still manually entering data and reliant on guesswork to make sense of it all as the volume and variety

Blog

Google Cloud VMware Engine Achieves HIPAA Compliance

We are excited to announce that as of April 1, 2021, Google Cloud VMware Engine is covered under the Google Cloud Business Associate Agreement (BAA), meaning it has achieved HIPAA compliance. Healthcare organizations can now migrate and run their HIPAA-compliant VMware workloads in a fully compatible VMware Cloud Verified stack running natively

Explainer

Document AI

Most business transactions begin, involve, or end with a document. But working with documents can be tricky, as leaders across industries seeking digital transformation can attest to. These enterprises face similar challenges as they seek to extract information from documents. The process can be costly, time consuming, and prone to

SHOW MORE STORIES