Lending DocAI Shortens Borrowers' Journey on Roostify - Build What's Next
Case Study

Lending DocAI Shortens Borrowers’ Journey on Roostify

11820

Of your peers have already read this article.

3:00 Minutes

The most insightful time you'll spend today!

Google Cloud's Lending DocAI automated Roostify's document processing for home application with multi-language support, allowing the provider of enterprise cloud apps for mortgage and home lenders manage upto thousands of borrowers on daily basis.

The home lending journey entails processing an immense number of documents daily from hundreds of thousands of borrowers. Currently, home lending document processing relies on some outdated digital models and a high dependency on manual labor, resulting in slow processing times and higher origination costs. Scaling a business that sorts through millions of documents daily, while increasing efficacy and accuracy, is no small feat. When it comes to applying for a mortgage loan, consumers expect a digital experience that’s as good as the in-person one. Roostify simplifies the home lending journey for lenders and their customers.

No time to spare: Overcoming document processing challenges with AI

Roostify provides enterprise cloud applications for mortgage and home lenders. In order to empower its customers to deliver a better, more personalized lending experience, they needed to automate and scale their in-house document parsing functionality.

As a key component of its document intelligence service, Roostify is leveraging Google Cloud’s Lending DocAI machine learning platform to automate processing documents required during a home loan application process, such as tax returns or bank statements with multi-language support. This partnership delivers data capture at scale, enabling Roostify customers to automatically identify document types from the uploaded file and to extract relevant entities such as wages, tax liabilities, names, and ID numbers for further processing, and make things move faster in the cumbersome lending process. 

Roostify’s solutions leverage Google Cloud’s Lending DocAI, which is built on the recently announced Document AI platform, a unified console for document processing. Customers can easily create and customize all the specialized parsers (e.g., mortgage lending documents and tax returns parsers) on the platform without the need to perform additional data mapping or training. All Google Cloud’s specialized parsers are fine-tuned to achieve industry-leading accuracy, helping customers and partners confidently unlock insights from documents with machine learning. Learn more about the solution from the GA launch blog and the overview video.

Integrating Lending DocAI’s intelligent document processing capabilities into the Roostify platform means more innovation for their customers and tangible results: faster loan processing times, fewer document intake errors, and lower origination costs. Additional support in Google Lending DAI for other languages and more documents like global Know Your Customer (KYC) documents or payroll reports is in the near future.

Full integration of AI solutions

Working together with Roostify’s platform team, we were able to help them solve their document processing challenge through integration of various GCP products such as Lending DocAI (LDAI), Data Loss Prevention (DLP) for redacting sensitive data, BigQuery for data warehousing and analytics, and Firestore for API status. To make it very safe and secure, all data was encrypted end-to-end at Rest and in Transit. LDAI won’t require any training data to process. It is an easy plug and play API.

Here is a sneak peek in the high level deployment architecture for LDAI in Roostify environment:

LDAI in Roostify.jpg

Here are the steps for processing data:

  1. Receives document processing request from the client.
  2. API Function directs requests to the pre-processing service. For Async requests a processing ID is generated and returned to the caller.
  3. Pre-processing service sends the request for further processing (Long/short PDF conversion), calling other microservices and receives back the responses. Any error in the response received is then sent to the response processing service. 
  4. If the response is synchronous, the pre-processing service directs it to the LDAI Invoker service. 
    1. If the response is asynchronous, the pre-processing service feeds it into the Cloud Pub/Sub service.
  5. Cloud Pub/Sub service feeds the response back to the LDAI Invoker service.
  6. LDAI Invoker service routes the request to the Google LDAI API for classification if there are multiple pages in the document.
  7. Document will be split based on LDAI response and then saved in a GCS bucket for temporary storage.
  8. LDAI entity interface for single page processing and then LDAI Invoker sends LDAI results to LDAI Response Processing
  9. If a request is a synchronous request the LDAI Response Processor sends results to the API Function so that it can complete the synchronous call and respond to the rConnect caller.
    1. If the request is an asynchronous request the LDAI Response Processor will respond to the caller’s webhook and complete the transaction.
  10. Finally, Data stored in the GCP bucket will be deleted.

All the responses that come from the LDAI API can optionally feed into BigQuery via the Response Processor, after parsing it through Data Loss Prevention (DLP) API to redact the PII/sensitive information.  Throughout the processing of both asynchronous and synchronous requests all transactions are logged using Cloud Logging.  For asynchronous transactions, the state is maintained throughout the process using Cloud Firestore.

Roostify currently uses this technology to power two different solutions: Roostify Document Intelligence and Roostify Beyond™. Roostify Document Intelligence is a real-time document capture, classification, and data extraction solution built for home lenders. It ingests documents uploaded by borrowers and loan officers, identifies the relevant documents, and extracts and classifies key information. Roostify Document Intelligence is available as a standalone API service to any home lender with any digital lending infrastructure already in place. 

Roostify Beyond™ is a robust suite of AI-powered solutions that enables home lenders to create intelligent experiences from start to close. It combines powerful data, insightful analytics, and meaningful visualization to streamline the underwriting process. Roostify Beyond™ is currently available only to Roostify customers as part of an Early Adopter program and will be rolled out to the market later this year.

field confidence level.jpg
Lenders can set the desired field confidence level. An extracted field that does not meet the set field confidence will display a warning indicator to borrowers asking them to validate the uploaded document.
Beyond algorithms.jpg
If the Beyond algorithms aren’t sure about the document (i.e., with lower confidence in the classification result than that set by the admin), the user sees a message asking them to validate the task.

Through this partnership, Roostify has enabled its customers to adopt a data-first approach to their home lending processes, which will lead to improved user experiences and significantly reduced loan processing times.

Fast track end-to-end deployment with Google Cloud AI Services (AIS)

Google AIS (Professional Services Organization), in collaboration with our partner Quantiphi, helped Roostify deploy this system into production and fast-tracked the development multifold to generate the final business value.

The partnership between Google Cloud and Roostify is just one of the latest examples of how we’re providing AI-powered solutions to solve business problems.

Blog

Transform Your Marketing Strategy with Tinyclues and Google Cloud CDP

2623

Of your peers have already read this article.

4:00 Minutes

The most insightful time you'll spend today!

Tinyclues and Google Cloud's BigQuery help marketers organize and centralize customer data for targeted and personalized campaigns. Our CDP allows for a comprehensive view of customers and informed marketing decisions.

Editor’s note: The post is part of a series highlighting our awesome partners, and their solutions, that are Built with BigQuery.

What are Customer Data Platforms (CDPs) and why do we need them?

Today, customers utilize a wide array of devices when interacting with a brand. As an example, think about the last time you bought a shirt. You may start with a search on your phone as you take the subway to work. During that 20 minute ride, you narrow down the type of shirt . Later, as you take your lunch break, you spend a few more minutes refining your search on your work laptop and you are able to find two shirt models of interest. Pressed for time, you add both to your shopping cart at an online retailer to review at a later point. Finally, after you arrive back home and as you are checking your physical mail, you stumble across a sales advertisement for the type of shirt that you are looking for, available at your local brick and mortar store. The next day you visit that store during your lunch break and purchase the shirt.

Many marketers face the challenge of creating a consistent 360 customer view that captures the customer lifecycle, as illustrated in the example above – including their online/offline journey, interacting with multiple data points across multiple data sources.

The evolution of managing customer data reached a turning point in the late 90’s with CRM software that sought to match current and potential customers with their interactions. Later as a backbone of data-driven marketing, Data Management Platforms (DMPs) expanded the reach of data management to include second and third party datasets including anonymous IDs. A Customer Data Platform combines these two types of systems, creating a unified, persistent customer view across channels (mobile, web etc) that provide data visibility and granularity at individual level.

A new approach to empowering marketing heroes

Tinyclues is a company that specializes in empowering marketers to drive sustainable engagement from their customers and generate additional revenue, without damaging customer equity. The company was founded in 2010 on a simple hunch: B2C marketing databases contain sufficient amounts of implicit information (data unrelated to explicit actions) to transform the way marketers interact with customers, and a new class of algorithms based on Deep Learning (sophisticated machine learning that mimics the way humans learn) holds the power to unlock this data’s potential. Where other players in the space have historically relied – and continue to rely – on a handful of explicit past behaviors and more than a handful of assumptions, Tinyclues’ predictive engine uses all of the customer data that marketers have available in order to formulate deeply precise models, down even to the SKU level. Tinyclues’ algorithms are designed to detect changes in consumption patterns in real-time, and adapt predictions accordingly.

This technology allows marketers to find precisely the right audiences for any offer during any timeframe, increasing engagement with those offers and, ultimately, revenue; additionally, marketers are able to increase campaign volume while decreasing customer fatigue and opt-outs, knowing that audiences are receiving only the most relevant messages. Tinyclues’ technology also reduces time spent building and planning campaigns by upwards of 80%, as valuable internal resources can be diverted away from manual audience-building.

Google Cloud’s Data Platform, spearheaded by BigQuery, provides a serverless, highly scalable, and cost-effective foundation to build this next generation of CDPs.

Tinyclues Architecture:


To enable this scalable solution for clients, Tinyclues receives purchase and interaction logs from clients in addition to product and user tables. In most cases, this data is already in the client’s BigQuery instance, in which case they can be easily shared with Tinyclues utilizing BigQuery authorized views.

In cases where the data is not in BigQuery, flat files are sent to Tinyclues via GCS and are ingested in the client’s data set via a lightweight Cloud Function. The orchestration of all pipelines is implemented via Cloud Composer (Google’s managed Airflow). The transformation of data is accomplished by utilizing simple select statements in the Data Built Tool (DBT), which is wrapped inside an airflow DAG that powers all data normalization and transformations. There are several other DAGs to fulfill more functionalities, including:

  • Indexing the product catalog on Elastic Cloud (Elasticsearch managed service) on GCP to provide auto-complete search capabilities to TCs clients as shown below:
  • The export of Tinyclues-powered audiences to the clients’ activation channels, whether they are using SFMC, Braze, Adobe, GMP, or Meta.

Tinyclues AI/ML Pipeline powered by Google Vertex AI

TCs ML Training pipelines are used to train models that calculate propensity scores. They are composed using Airflow DAGs, powered by Tensorflow & Vertex AI Pipelines. BigQuery is used natively, without data movement, to perform as much feature engineering as possible in-place.

TC uses the TFX library to run ML Pipelines in Vertex AI. Building on top of Tensorflow as their main deep learning framework of choice due to its maturity, open source platform, scalability and support for complex data structures (Ragged and Sparse Tensors).

Below is a partial example of TC’s Vertex AI Pipeline graph, illustrating the workflow steps in the training pipeline. This pipeline allows for the modularization & standardization of functionality into easily manageable building blocks. These blocks are composed of TFX components (TC reuses most of the standard components in addition to customizing some such as a proprietary implementation of the Evaluator to compute both ML Metrics (which is part of the standard implementation) but also more Business Metrics like Overlap of clickers etc. The individual components/steps are chained with DSL to form a pipeline that is modular and easily orchestrated or updated as needed.

With the trained Tensorflow models available in GCS, TCs exposes these in BigQuery ML (BQML) to enable their clients to score millions of users for their propensity to buy X or Y within minutes. This would not be possible without the power of BigQuery and also frees TC from previously experienced scalability issues.

As an illustration, TC has the need to score thousands of topics among millions of users. This used to take north of 20 hours on their previous stack, and now takes less than 20 minutes thanks to the optimization work that TC has implemented in their custom algorithm and the sheer power of BQ to scale to any workload accordingly.

Data Gravity: Breaking the Paradigm – Bringing the Model to your Data

BQML enables TC to call pre-trained TensorFlow models within an SQL environment, thus avoiding exporting data in and out of BQ using already provisioned BQ serverless processing power. Using BQML removes the layers between the models and the data warehouse and allows them to express the entire inference pipe as a number of SQL requests. TC no longer has to export data to load it into their models. Instead, they are bringing their models to the data.

Avoiding the export of data in and out of BQ and the serverless provisioning and start of machines saves significant time. As an example, exporting an 11M lines campaign for a large client previously took 15 min or more to process. Deployed on BQML it now takes minutes with more than half of the processing time attributed to network transfers to our client system.

Inference times in BQML compared to TCs legacy stack:

As can be seen, using this approach enabled by BQML, the reduction in the number of steps leads to a 50% decrease in overall inference time, improving upon each step of the prediction.

The Proof is in the pudding

Tinyclues has consistently delivered on its promises of increased autonomy for CRM teams, rapid audience building, superior performance against in-house segmentation, identification of untapped messaging and revenue opportunities, fatigue management, and more, working with partners like Tiffany & Co, Rakuten, and Samsung, among many others.

Conclusion

Google’s data cloud provides a complete platform for building data-driven applications like the headless CDP solution developed by Tinyclues — from simplified data ingestion, processing, and storage to powerful analytics, AI, ML, and data sharing capabilities — all integrated with the open, secure, and sustainable Google Cloud platform. With a diverse partner ecosystem, open-source tools, and APIs, Google Cloud can provide technology companies the portability and differentiators they need to serve the next generation of marketing customers.

To learn more about Tinyclues on Google Cloud, visit Tinyclues. Click here to learn more about Google Cloud’s Built with BigQuery initiative.

We thank the many Google Cloud team members who contributed to this ongoing data platform collaboration and review, especially Dr. Ali Arsanjani in Partner Engineering.

Case Study

Google Cloud Partnership Fuels ListenField’s Agriculture Revolution

917

Of your peers have already read this article.

3:30 Minutes

The most insightful time you'll spend today!

Discover how ListenField leverages AI, machine learning, and Google Cloud to empower farmers, optimize agriculture, and enhance sustainability, revolutionizing the future of farming in Southeast Asia.

When I was growing up in Thailand, I witnessed the challenges facing farmers including rising food demand, shortage of labor, and uneven crop yields caused by climate change. As a result, many smallholder farmers found themselves trapped in a vicious circle, unable to reduce food insecurity due to low yields, but lacking the resources to invest for a more profitable future.

I was determined to make a difference. After earning my master’s degree in Information Management I joined a research project with the University of Tokyo where we used sensors to monitor spinach fields in Thailand. The results of this early experiment in precision farming were impressive. We proved that it was possible to grow organic crops with minimal use of fertilizers, while consumers benefited from higher quality spinach grown in Thailand.  

This inspired me to found ListenField in 2017. Our mission is to transform farm management by collecting data from multiple sources, including field sensors, soil scanning, weather data, seasonal forecasts, and satellite imagery. By modeling this data, we provide farmers with insights that enable them to optimize production ‘from soil to harvest’.

A bumper crop of farming data

Artificial intelligence and machine learning play a central role in our prediction platform, combining crop health monitoring, growth prediction, and soil nutrition analysis. Farmers can apply real-time insights to their schedule from our FarmAI Mobile App, while our FarmAI Dashboard enables agri-food businesses to collaborate with agronomists and farmers so that all parties benefit from higher profit margins and more sustainable growing strategies.

Another important feature of our business model is AgroAPI, which makes our analytics available to third parties. Clients can embed deep analytics in their applications, including crop growth prediction and remote sensing analysis, without needing to develop complicated algorithms and data pipelines by themselves.

We are also excited about our research into genomic prediction in collaboration with the Japanese government and several research companies. Using our Data-Driven Breeding Platform, breeders and seed companies can upload their genomic data and gain practical insights that help accelerate the reproduction of high-quality seeds and plants.

Today, more than 30,000 farmers use our technology especially in Vietnam and Thailand where it is used to improve rice, cassava, and sugar cane yields. The technology is also being rolled out for orange and mango farmers enabling them to monitor individual trees and adjust irrigation to improve the sweetness of the fruit at harvest time.

Responding fast to changing conditions

We were using another cloud provider to run our business, but one of the ListenField team members drew our attention to Google Cloud. As well as the technology, we were also attracted by the Google for Startups Cloud Program which provides us with Google Cloud credits that cover our first and second years of Google Cloud usage. We also met our Account Representative, who provides us with training, business and tech support, and Google-wide discounts. It’s great to have a point of contact that can help us on our startup journey and make the most of Google’s resources.

By reducing the pressure on our finances and human resources, we were able to experiment and adapt in response to early experiments. This flexibility also enabled us to demonstrate a compelling business case to new and existing investors.

Our Google Cloud Platform Partner, Navagis, gave us additional momentum thanks to their expertise in mapping and geospatial data. They also played a crucial role in the integration of Google Earth Engine, which we use to map agricultural areas.

We also use Firebase for application development and Google Workspace for team collaboration. Colab and Vertex AI enable us to build, deploy, and scale our machine learning models quickly, ensuring that we remain competitive and attractive to new customers.

Giving female entrepreneurs the opportunity to flourish

Both the Google Cloud team and our colleagues at Navagis helped us to navigate the challenges many early-stage startups face. My background is in science and academia, so I appreciated the business mindset offered by both organizations to help us continue to grow and scale.

Being a female entrepreneur leading a startup can also be tough, but Google Cloud and Navagis helped me to build a strong network, access funding, and make my voice heard. Today, ListenField has several female executives, while 50% of our researchers and many of the farmers on our platform are women.  

Above all, Google Cloud helps us to power an agriculture revolution in south-east Asia. Smallholder farmers can transition from analog to digital farming, improving their yields and reducing waste. The benefits to the economy are also significant including greater food security and reducing the use of industrial fertilizers that generate potent greenhouse gasses.

And that’s just the beginning of what we can do. Our next milestone is to reach 50,000 farmers and cut one million tonnes of greenhouse gas emissions. It sounds ambitious, but with the Google Cloud and Navagis teams behind us, I’m confident that we will reach these targets.

https://storage.googleapis.com/gweb-cloudblog-publish/images/ListenField.max-2000x2000.jpg

ListenField team members

If you want to learn more about how Google Cloud can help your startup, visit our page here to get more information about our program, and sign up for our communications to get a look at our community activities, digital events, special offers, and more.

Blog

Transforming the Contact Center Experience with Artificial Intelligence

2506

Of your peers have already read this article.

1:30 Minutes

The most insightful time you'll spend today!

Unlock the potential of AI in contact centers to enhance customer experience. Learn how far we've come and what's possible in our retrospective and future outlook of using AI in contact centers.

We meet daily with contact center owners and customer experience (CX) execs across all industries, geographies, and business sizes. Looking back at these conversations, it’s crystal clear that 2022 was a high-stakes year for call centers, with three primary challenges trending across all customers and continuing in 2023:

  • Many organizations feel pressure to rapidly scale up their call center operations in response to macroeconomic changes. Uncertain conditions are forcing Contact Centers to be ever more cost-effective, and to find ways to generate revenue for the business.
  • End users are increasingly demanding and less forgiving when it comes to CX. Users have a choice, and they expect brands to meet them where they are with superior experiences. Connecting with customers where they engage is one of the key components of superior CX—customers should not have to go through elaborate processes or unhelpful phone trees to get help but should rather have service available quickly and easily in their preferred channels. Consumers demand more intimate ways of connecting with brands and Conversational AI can create that critical interaction medium. 
  • Organizations understand that AI can help address these challenges. However, many business leaders remain unsure how to successfully make the journey. A growing number of offerings are on the market, but many don’t deliver on their promise, with long and expensive integration requirements and unpredictable and underwhelming outcomes.

Helping our customers successfully address these challenges and opportunities was one of our top priorities last year and will continue to be a significant focus in coming months. In this blog post, we’ll review our Contact Center AI (CCAI) news from last year, as a primer for 2023. 

Looking back: Why 2022 was a big year for Contact Center AI

In 2022, we increased our strategic investment in CCAI, including expanding it to include a comprehensive, end-to-end contact center solution suite that is user-first, AI-first, and cloud-first. We launched Contact Center AI Platform, our Contact Center as a Service (CCaaS) offering, as part of the CCAI product suite that offers a modern, turnkey solution, designed with user-first, AI-first, and cloud-first design. During Google Cloud Next ‘22, we shared lots of great content on how organizations can use CCAI to improve customer experiences, including these breakout sessions:

We also got a chance to hear how customers are using CCAI to better reach their own customers, including Wells Fargo and TIAA. We partnered with CDW to discuss Providing Better Customer Experiences and with Quantiphi in a webinar called “Elevating the Banking Experience with CCAI Platform.” Just recently, our customer Segra shared their success story.

Through these customer interactions, three key priorities have surfaced as we look forward to 2023: Elevate the customer experience, bring new forms of AI to drive new automation and accelerate time to value.

Looking forward: Elevate CX, integrate new forms of AI, accelerate time to value 

1. User-first: Meet them where they are with elevated Customer Experience.

As we have learned, users expect that brands meet them where they are and on their own terms and expectations. To do that, brands must integrate with and adopt the latest user-centric technologies and product best practices from consumer mobile and web apps. Enterprise B2C can’t exist anymore in a parallel world of different and often inferior user experience. Google has over 20 years of experience in building such consumer experiences, with multiple products successfully serving billions of users. Bringing these capabilities and experiences from our consumer products and research teams to our cloud offerings was a key component for our product offerings in 2022 and is a big part of our key investments in 2023. Moreover, a vast majority of CX user journeys start with a query on Google Search or YouTube. Connecting with the users at that point, even before they reach out directly to the contact center is a win-win, saving money for the brand and delivering immediate value to the user. By focusing on the user we created a superior  integrated omnichannel experience.

2. AI-first and cloud-first: Quality contact center growth depends on transforming to modern, Cloud, AI solutions.

For contact centers to evolve, they need to transform from cost centers to revenue generators. That requires modern Cloud and AI solutions. Conversational data spans across all parts of the contact center, opening new ways to generate value. Cloud capabilities of privacy, security and scale can enable personalized CX across channels, enabling key omnichannel experiences. From a study by McKinsey: “Cross-channel integration and migration issues continue to hamper progress. For example, 77 percent of survey respondents report that their organizations have built digital platforms, but only 10 percent report that those platforms are fully scaled and adopted by customers. Only 12 percent of digital platforms are highly integrated, and, for most organizations, only 20 percent of digital contacts are unassisted.” Traditional telephony technologies are becoming commoditized and struggle to keep up with ever more complex rule based systems. Leaders in applicative AI and Cloud technology are stepping up as the new partners for brands who understand they need to take the leap to the next generation CX solutions. .  

3. Accelerating time to value while future proofing investments with predictable and measurable value

Reducing upfront implementation investment and accelerating time to value can be a challenge for contact center solutions. Scaling Cloud and AI can provide a faster path advanced conversational AI, can help address these challenges. Let’s look at three examples:

  • Out of the Box(OOTB) integrated transcription, chat and voice summarization, and topic modeling — This saves customers money by reducing agent handling time for every chat and call, as well as providing valuable insights that can be used for quality management, contact center optimization and automation, agent and user churn prediction, business insights, and revenue opportunities.
  • AI based chat and voice calls steering  paired with info-seeking virtual agents — Together these deliver higher Customer Satisfaction at scale while reducing cost –  by significantly reducing waiting queues and being routed to the wrong agent, as well as automating away total handling time.
  • Reduced time to full automation — Reduce the complexity of conversation modeling, prebuilt components and APIs for shorter time to value and more predictable outcomes, and metrics driven ML-Dev & QA tools and playbooks.

With these new capabilities, our customers can now see results as soon as they implement CCAI. We’re excited to get our customers to where they want to be faster!

And there you have it: a quick overview of CCAI and its progress in 2022 and what’s coming in 2023. For more details, check out the documentation or our CCAI solutions page.

Trend Analysis

Google Research: Themes from 2021 and Beyond

2883

Of your peers have already read this article.

4:00 Minutes

The most insightful time you'll spend today!

Read the blogpost to catch up on Google's research on AI and five emerging ML-related trends that are poised to redefine the way systems interact around the world with new product features and accomplish more with ML models.


Posted by Jeff Dean, Senior Fellow and SVP of Google Research, on behalf of the entire Google Research community

Over the last several decades, I’ve witnessed a lot of change in the fields of machine learning (ML) and computer science. Early approaches, which often fell short, eventually gave rise to modern approaches that have been very successful. Following that long-arc pattern of progress, I think we’ll see a number of exciting advances over the next several years, advances that will ultimately benefit the lives of billions of people with greater impact than ever before. In this post, I’ll highlight five areas where ML is poised to have such impact. For each, I’ll discuss related research (mostly from 2021) and the directions and progress we’ll likely see in the next few years.


· Trend 1: More Capable, General-Purpose ML Models
· Trend 2: Continued Efficiency Improvements for ML
· Trend 3: ML Is Becoming More Personally and Communally Beneficial
· Trend 4: Growing Benefits of ML in Science, Health and Sustainability
· Trend 5: Deeper and Broader Understanding of ML

Trend 1: More Capable, General-Purpose ML Models


Researchers are training larger, more capable machine learning models than ever before. For example, just in the last couple of years models in the language domain have grown from billions of parameters trained on tens of billions of tokens of data (e.g., the 11B parameter T5 model), to hundreds of billions or trillions of parameters trained on trillions of tokens of data (e.g., dense models such as OpenAI’s 175B parameter GPT-3 model and DeepMind’s 280B parameter Gopher model, and sparse models such as Google’s 600B parameter GShard model and 1.2T parameter GLaM model). These increases in dataset and model size have led to significant increases in accuracy for a wide variety of language tasks, as shown by across-the-board improvements on standard natural language processing (NLP) benchmark tasks (as predicted by work on neural scaling laws for language models and machine translation models).

Many of these advanced models are focused on the single but important modality of written language and have shown state-of-the-art results in language understanding benchmarks and open-ended conversational abilities, even across multiple tasks in a domain. They have also shown exciting capabilities to generalize to new language tasks with relatively little training data, in some cases, with few to no training examples for a new task. A couple of examples include improved long-form question answering, zero-label learning in NLP, and our LaMDA model, which demonstrates a sophisticated ability to carry on open-ended conversations that maintain significant context across multiple turns of dialog.

A dialog with LaMDA mimicking a Weddell seal with the preset grounding prompt, “Hi I’m a weddell seal. Do you have any questions for me?” The model largely holds down a dialog in character.
(Weddell Seal image cropped from Wikimedia CC licensed image.)

Transformer models are also having a major impact in image, video, and speech models, all of which also benefit significantly from scale, as predicted by work on scaling laws for visual transformer models. Transformers for image recognition and for video classification are achieving state-of-the-art results on many benchmarks, and we’ve also demonstrated that co-training models on both image data and video data can improve performance on video tasks compared with video data alone. We’ve developed sparse, axial attention mechanisms for image and video transformers that use computation more efficiently, found better ways of tokenizing images for visual transformer models, and improved our understanding of visual transformer methods by examining how they operate compared with convolutional neural networks. Combining transformer models with convolutional operations has shown significant benefits in visual as well as speech recognition tasks.

The outputs of generative models are also substantially improving. This is most apparent in generative models for images, which have made significant strides over the last few years. For example, recent models have demonstrated the ability to create realistic images given just a category (e.g., “irish setter” or “streetcar”, if you desire), can “fill in” a low-resolution image to create a natural-looking high-resolution counterpart (“computer, enhance!”), and can even create natural-looking aerial nature scenes of arbitrary length. As another example, images can be converted to a sequence of discrete tokens that can then be synthesized at high fidelity with an autoregressive generative model.

Example of a cascade diffusion models that generate novel images from a given category and then use those as the seed to create high-resolution examples: the first model generates a low resolution image, and the rest perform upsampling to the final high resolution image.
The SR3 super-resolution diffusion model takes as input a low-resolution image, and builds a corresponding high resolution image from pure noise.

Because these are powerful capabilities that come with great responsibility, we carefully vet potential applications of these sorts of models against our AI Principles.

Beyond advanced single-modality models, we are also starting to see large-scale multi-modal models. These are some of the most advanced models to date because they can accept multiple different input modalities (e.g., language, images, speech, video) and, in some cases, produce different output modalities, for example, generating images from descriptive sentences or paragraphs, or describing the visual content of images in human languages. This is an exciting direction because like the real world, some things are easier to learn in data that is multimodal (e.g., reading about something and seeing a demonstration is more useful than just reading about it). As such, pairing images and text can help with multi-lingual retrieval tasks, and better understanding of how to pair text and image inputs can yield improved results for image captioning tasks. Similarly, jointly training on visual and textual data can also help improve accuracy and robustness on visual classification tasks, while co-training on image, video, and audio tasks improves generalization performance for all modalities. There are also tantalizing hints that natural language can be used as an input for image manipulation, telling robots how to interact with the world and controlling other software systems, portending potential changes to how user interfaces are developed. Modalities handled by these models will include speech, sounds, images, video, and languages, and may even extend to structured data, knowledge graphs, and time series data.

Example of a vision-based robotic manipulation system that is able to generalize to novel tasks. Left: The robot is performing a task described in natural language to the robot as “place grapes in ceramic bowl”, without the model being trained on that specific task. Right: As on the left, but with the novel task description of “place bottle in tray”.

Often these models are trained using self-supervised learning approaches, where the model learns from observations of “raw” data that has not been curated or labeled, e.g., language models used in GPT-3 and GLaM, the self-supervised speech model BigSSL, the visual contrastive learning model SimCLR, and the multimodal contrastive model VATTSelf-supervised learning allows a large speech recognition model to match the previous Voice Search automatic speech recognition (ASR) benchmark accuracy while using only 3% of the annotated training data. These trends are exciting because they can substantially reduce the effort required to enable ML for a particular task, and because they make it easier (though by no means trivial) to train models on more representative data that better reflects different subpopulations, regions, languages, or other important dimensions of representation.

All of these trends are pointing in the direction of training highly capable general-purpose models that can handle multiple modalities of data and solve thousands or millions of tasks. By building in sparsity, so that the only parts of a model that are activated for a given task are those that have been optimized for it, these multimodal models can be made highly efficient. Over the next few years, we are pursuing this vision in a next-generation architecture and umbrella effort called Pathways. We expect to see substantial progress in this area, as we combine together many ideas that to date have been pursued relatively independently.

Pathways: a depiction of a single model we are working towards that can generalize across millions of tasks.

Blog

Headless e-Commerce is the Next Big Thing in Retail

5008

Of your peers have already read this article.

2:00 Minutes

The most insightful time you'll spend today!

Headless commerce (HC) for retail helps innovate, develop and launch with limited resources, decoupling the backend and frontend. Learn more about headless e-commerce as the future of retail and commerce tools on Google Cloud Marketplace.

Headless Ecommerce

In the last couple of years there has been a shift in the way retailers approach ecommerce: where in the past development efforts were prioritized around building a solid foundation for backend transactions and operations now it is clear that companies in this space are focusing on differentiating themselves by creating unique shopping experiences that increase engagement and reduce friction. 

But how can development teams spend the necessary time designing and writing code for this kind of interactions while also having to seamlessly maintain ecommerce vital components like online catalogs, shopping carts and checkout payment processes? Enter headless commerce.  

Headless commerce (HC) helps companies of all sizes to innovate, develop and launch in less time and using fewer resources by decoupling backend and frontend. Headless solution providers empower online retailers by offering a balance between flexibility and optimization through pre-built api-accessible modules and components that can be easily plugged into their frontend architecture. This translates into rapid development while keeping desired levels of security, compliance, integration and responsiveness. 

This composable approach enables dev teams not only to create new features but also connect other ecommerce components with less effort which is critical when responding to business trends. But above all, the main benefit retailers receive from HC, is owning and controlling the frontend for an engaging customer journey as well as quickly launching new experiences.  

Google Cloud + commercetools

commercetools, a leader in the headless commerce space, has partnered with Google Cloud to make their cloud-native SaaS platform available in the Google Cloud Marketplace. With a flexible API system (REST API and GraphQL), commercetools’ architecture has been designed to meet the needs of demanding omnichannel ecommerce projects while offering real flexibility to modify or extend its features. It supports a variety of storefront providers like Vue Storefront, offers a large set of integrations and supports microservice-based architectures. All this while providing access to multiple programming languages (PHP, JS, Java) via its SDK tools

commercetools and Google Cloud provide development teams with all the tools to build high-quality digital commerce systems. Google Cloud’s scalability, AI/ML components, API management capabilities and CI/CD tools are a perfect fit to build frontend shopping experiences that easily integrate with the commercetools stack. Developers can take advantage of this compatibility by:

Additionally, commercetools allows ecommerce solutions to tap into the wider Google Ecosystem by providing authoritative data via Merchant Center, advertising via product listing ads and selling via Google Shopping

Architecture Overview

As mentioned previously, headless commerce is increasingly preferred by retailers who want to own and control the ‘front-end’ for providing and enabling an engaging and differentiated user and shopping experiences. 

The approach involves a loosely coupled architecture that separates ‘front end’ from the ‘back end’ of a digital commerce application. The front end is typically built and managed by the retailer. They want to leverage an independent software vendor (ISV) offered, ready-to-use ‘back-end’ commerce building blocks for capabilities, such as product catalog, pricing, promotions, cart, shipping, account and others.

retail.jpg

Most retailers want to invest their time and resources in building a front end that requires an agile development model to introduce new and tweaking existing user experiences to acquire and retain customers. A few retailers that do not have an in-house web development team may choose an ISV that offers ready to use front end. The front end is a web app and designed as a progressive web application (PWA) on Google Cloud. The backend is a headless commerce offered by an ISV, such as commercetools. The backend commerce capabilities are built as a set of microservices, exposed as APIs, run cloud-native and implemented as headless. It is commonly referred to as the  “MACH” solution. The API-first approach of the architecture allows easily integrating ‘best of breed’ capabilities built internally and/or offered by 3rd party ISVs.

Leveraging Google Cloud Components

The architecture of the front end will be implemented on Google Cloud and will integrate with the ISV’s headless commerce back end that runs natively in Google Cloud. 

The front end will be designed using cloud-native services for 

Additionally, API management (Apigee on Google Cloud) can be used to orchestrate interactions of the front end with the APIs of the backend commerce services. The API management’s capability will be used for accessing the services of on-premises systems, such as ERP, order management system (OMS), warehouse management system (WMS) as needed to support the functioning of digital commerce application.  Alternatively, depending on the frontend capabilities, developers can use middleware to build custom services and route requests. 

What’s next?

A considerable number of retailers have adopted headless commerce and are now focusing on adopting best practices and leveraging the agility that comes with this approach. Just like commercetools offers robust components that meet the retailer’s backend operational needs (Product CatalogOrder ManagementCartsPayments, etc), Google Cloud’s Compute, Networking, Severless and AI/ML services  provide the agility and flexibility required by development teams to quickly and easily extend their frontend capabilities. 

commercetools and Google Cloud work seamlessly together because they both prioritize ease of integration, scalability, security and iterability while providing ready-to-use building blocks. It also helps that commercetools backend runs on Google Cloud. Once an initial foundation of Google Cloud and commercetools has been established, adding new commerce modules and extending functionally of the current ones becomes a straightforward process that allows to route efforts to innovation initiatives. In the end, the main beneficiaries of this technical synergy are the shoppers that enjoy experiences which increase engagement and minimize friction. 

Alternatively, retailers can also save time and resources by relying on frontend integrations. commercetools offers a variety of third-party solutions that can effortlessly be added to a headless commerce architecture. These integrations as well as other important headless commerce extensions will be explored in future blog entries.  In the meantime, all the necessary tools to leverage headless commerce can be found in just one place: 

Get started with commercetools on the Google Cloud Marketplace today!

More Relevant Stories for Your Company

Case Study

Google had Enough of Flooding in India. Here’s What It Did About It

Over the last century, floods have become the most common and deadly natural disaster on the planet. Many countries currently lack effective early warning systems and alerts. Today, 20 percent of flood fatalities occur in India. To help, Google sent a team to study the Ghaghara River, near Patna. “In

Blog

Data to Business Outcomes with Google’s Data Analytics Design Pattern

Companies today are inundated with vast amounts of data from various sources. This overwhelming amount of data is meant to benefit the company, but often leaves data teams feeling overwhelmed, which can create data bottlenecks and result in a slow time to value. In fact, only twenty seven percent of

Blog

How Google Cloud’s PSO Supports Customers’ Migration Goals

Google Cloud’s Professional Services Organization (PSO) engages with customers to ensure effective and efficient operations in the cloud, from the time they begin considering how cloud can help them overcome their operational, business or technical challenges, to the time they’re looking to optimize their cloud workloads.  We know that all parts of

Research Reports

Download the Forrester Study to Explore the Benefits of AI for IT Operations in Cloud Environment

Organizations are currently modernizing their businesses in order to meet the increasing complexity of today’s business landscape. In effect, business leaders must evaluate the best way to mitigate the challenges which plague their cloud operations, all while meeting customers’ growing expectations around digital experience (DX) through agility, automation, and proactive

SHOW MORE STORIES