Custom Voice Feature Can Help Brands Tweak IVR for Better Customer Experiences

5146
Of your peers have already read this article.
1:30 Minutes
The most insightful time you'll spend today!
With the rise of digital assistants and conversational interfaces, people have grown accustomed to hearing and speaking to synthetic voices. But what do those voices sound like? Often, pretty repetitive. We’re all familiar with the Google Assistant voice, for example.
That’s why we are excited to announce the general availability of Custom Voice in our Cloud Text-to-Speech (TTS) API, a new feature that lets you train custom voice models with your own audio recordings to create unique experiences.
For businesses looking to build a strong brand identity, establishing a unique voice can help turn mobile app interactions or customer service based on interactive voice responses (IVR) into differentiated customer experiences. Our TTS API has included a speech synthesis service with a static list of voices for some time, but now, with Custom Voice, moving beyond these predefined options is easier than ever.
Custom Voice lets you simply submit your audio recordings to get access to the new voice directly in the TTS API. Custom Voice TTS includes guidance on the audio requirements to help make sure you generate a high quality custom TTS voice model. Once this new model is trained, all you have to do to start using the newly trained voice is reference the model ID in your calls to the Cloud TTS API.
At Google, we are committed to building safe and accountable AI products, not only because it’s the right thing to do, but because it is a critical step in ensuring successful use in production. As part of Google Cloud’s Responsible AI governance process, we conducted a deep ethical evaluation of Custom Voice TTS, and its relation to synthetic media, in order to surface and mitigate potential harms that it may create. If you are interested in Custom Voice TTS, there is a review process to help ensure each use case is aligned with our AI Principles and adequate voice actor consent is given.
Additionally, to verify that voice actors are actually the ones producing the audio, you will need to submit an audio file producing a sentence that Google Cloud chooses (for example: “I agree that my voice will be used to create a synthetic custom Text-to-Speech voice).
We’re looking forward to seeing this API help businesses solve problems in an easy, fast, and scalable way. TTS Custom Voice is now GA in these languages:
English (US)
English (AU)
English (UK)
Spanish (US)
Spanish (Spain)
French (France)
French (Canada)
Italian (Italy)
German (Germany)
Portugues (Brazil)
Japanese (Japan)
We plan to continue expanding this lineup in order to meet your needs. Ready to try for yourself? Contact your seller to get started on your use case evaluation today!
Headless e-Commerce is the Next Big Thing in Retail

5014
Of your peers have already read this article.
2:00 Minutes
The most insightful time you'll spend today!
Headless Ecommerce
In the last couple of years there has been a shift in the way retailers approach ecommerce: where in the past development efforts were prioritized around building a solid foundation for backend transactions and operations now it is clear that companies in this space are focusing on differentiating themselves by creating unique shopping experiences that increase engagement and reduce friction.
But how can development teams spend the necessary time designing and writing code for this kind of interactions while also having to seamlessly maintain ecommerce vital components like online catalogs, shopping carts and checkout payment processes? Enter headless commerce.
Headless commerce (HC) helps companies of all sizes to innovate, develop and launch in less time and using fewer resources by decoupling backend and frontend. Headless solution providers empower online retailers by offering a balance between flexibility and optimization through pre-built api-accessible modules and components that can be easily plugged into their frontend architecture. This translates into rapid development while keeping desired levels of security, compliance, integration and responsiveness.
This composable approach enables dev teams not only to create new features but also connect other ecommerce components with less effort which is critical when responding to business trends. But above all, the main benefit retailers receive from HC, is owning and controlling the frontend for an engaging customer journey as well as quickly launching new experiences.
Google Cloud + commercetools
commercetools, a leader in the headless commerce space, has partnered with Google Cloud to make their cloud-native SaaS platform available in the Google Cloud Marketplace. With a flexible API system (REST API and GraphQL), commercetools’ architecture has been designed to meet the needs of demanding omnichannel ecommerce projects while offering real flexibility to modify or extend its features. It supports a variety of storefront providers like Vue Storefront, offers a large set of integrations and supports microservice-based architectures. All this while providing access to multiple programming languages (PHP, JS, Java) via its SDK tools.
commercetools and Google Cloud provide development teams with all the tools to build high-quality digital commerce systems. Google Cloud’s scalability, AI/ML components, API management capabilities and CI/CD tools are a perfect fit to build frontend shopping experiences that easily integrate with the commercetools stack. Developers can take advantage of this compatibility by:
- Integrating systems with Google’s Retail Search, Recommendations AI and Vision Product Search
- Injecting serverless functions into commercetools using Google Cloud Functions
- Extending and integrating commercetools via Events handled by Pub/Sub
- Managing 3rd party, legacy, microservices and commercetools APIs with Apigee
- Selecting the Google Cloud region commercetools uses for zero latency for custom apps
- Expanding their microservice ecosystem with components like Cloud Storage, Cloud SQL, Firestore and BigQuery
Additionally, commercetools allows ecommerce solutions to tap into the wider Google Ecosystem by providing authoritative data via Merchant Center, advertising via product listing ads and selling via Google Shopping.
Architecture Overview
As mentioned previously, headless commerce is increasingly preferred by retailers who want to own and control the ‘front-end’ for providing and enabling an engaging and differentiated user and shopping experiences.
The approach involves a loosely coupled architecture that separates ‘front end’ from the ‘back end’ of a digital commerce application. The front end is typically built and managed by the retailer. They want to leverage an independent software vendor (ISV) offered, ready-to-use ‘back-end’ commerce building blocks for capabilities, such as product catalog, pricing, promotions, cart, shipping, account and others.

Most retailers want to invest their time and resources in building a front end that requires an agile development model to introduce new and tweaking existing user experiences to acquire and retain customers. A few retailers that do not have an in-house web development team may choose an ISV that offers ready to use front end. The front end is a web app and designed as a progressive web application (PWA) on Google Cloud. The backend is a headless commerce offered by an ISV, such as commercetools. The backend commerce capabilities are built as a set of microservices, exposed as APIs, run cloud-native and implemented as headless. It is commonly referred to as the “MACH” solution. The API-first approach of the architecture allows easily integrating ‘best of breed’ capabilities built internally and/or offered by 3rd party ISVs.
Leveraging Google Cloud Components
The architecture of the front end will be implemented on Google Cloud and will integrate with the ISV’s headless commerce back end that runs natively in Google Cloud.
The front end will be designed using cloud-native services for
- PWA web app development (Google Kubernetes Engine, CI/CD services),
- Google Product Discovery solution that includes Retail Search and Vision API Product Search for serving product search (text and image) and Recommendations AI for serving recommendations.
- Storage (Cloud Storage), Database (Cloud SQL, Cloud Firestore), and edge caching for content delivery (Cloud CDN)
- Networking (Cloud DNS, Global Load Balancing), and
- Security (Cloud Armor for DDoS, API Defense for API protection)
Additionally, API management (Apigee on Google Cloud) can be used to orchestrate interactions of the front end with the APIs of the backend commerce services. The API management’s capability will be used for accessing the services of on-premises systems, such as ERP, order management system (OMS), warehouse management system (WMS) as needed to support the functioning of digital commerce application. Alternatively, depending on the frontend capabilities, developers can use middleware to build custom services and route requests.
What’s next?
A considerable number of retailers have adopted headless commerce and are now focusing on adopting best practices and leveraging the agility that comes with this approach. Just like commercetools offers robust components that meet the retailer’s backend operational needs (Product Catalog, Order Management, Carts, Payments, etc), Google Cloud’s Compute, Networking, Severless and AI/ML services provide the agility and flexibility required by development teams to quickly and easily extend their frontend capabilities.
commercetools and Google Cloud work seamlessly together because they both prioritize ease of integration, scalability, security and iterability while providing ready-to-use building blocks. It also helps that commercetools backend runs on Google Cloud. Once an initial foundation of Google Cloud and commercetools has been established, adding new commerce modules and extending functionally of the current ones becomes a straightforward process that allows to route efforts to innovation initiatives. In the end, the main beneficiaries of this technical synergy are the shoppers that enjoy experiences which increase engagement and minimize friction.
Alternatively, retailers can also save time and resources by relying on frontend integrations. commercetools offers a variety of third-party solutions that can effortlessly be added to a headless commerce architecture. These integrations as well as other important headless commerce extensions will be explored in future blog entries. In the meantime, all the necessary tools to leverage headless commerce can be found in just one place:
Get started with commercetools on the Google Cloud Marketplace today!
Google Research: Themes from 2021 and Beyond

2887
Of your peers have already read this article.
4:00 Minutes
The most insightful time you'll spend today!
Posted by Jeff Dean, Senior Fellow and SVP of Google Research, on behalf of the entire Google Research community
Over the last several decades, I’ve witnessed a lot of change in the fields of machine learning (ML) and computer science. Early approaches, which often fell short, eventually gave rise to modern approaches that have been very successful. Following that long-arc pattern of progress, I think we’ll see a number of exciting advances over the next several years, advances that will ultimately benefit the lives of billions of people with greater impact than ever before. In this post, I’ll highlight five areas where ML is poised to have such impact. For each, I’ll discuss related research (mostly from 2021) and the directions and progress we’ll likely see in the next few years.
· Trend 1: More Capable, General-Purpose ML Models
· Trend 2: Continued Efficiency Improvements for ML
· Trend 3: ML Is Becoming More Personally and Communally Beneficial
· Trend 4: Growing Benefits of ML in Science, Health and Sustainability
· Trend 5: Deeper and Broader Understanding of ML
Trend 1: More Capable, General-Purpose ML Models
Researchers are training larger, more capable machine learning models than ever before. For example, just in the last couple of years models in the language domain have grown from billions of parameters trained on tens of billions of tokens of data (e.g., the 11B parameter T5 model), to hundreds of billions or trillions of parameters trained on trillions of tokens of data (e.g., dense models such as OpenAI’s 175B parameter GPT-3 model and DeepMind’s 280B parameter Gopher model, and sparse models such as Google’s 600B parameter GShard model and 1.2T parameter GLaM model). These increases in dataset and model size have led to significant increases in accuracy for a wide variety of language tasks, as shown by across-the-board improvements on standard natural language processing (NLP) benchmark tasks (as predicted by work on neural scaling laws for language models and machine translation models).
Many of these advanced models are focused on the single but important modality of written language and have shown state-of-the-art results in language understanding benchmarks and open-ended conversational abilities, even across multiple tasks in a domain. They have also shown exciting capabilities to generalize to new language tasks with relatively little training data, in some cases, with few to no training examples for a new task. A couple of examples include improved long-form question answering, zero-label learning in NLP, and our LaMDA model, which demonstrates a sophisticated ability to carry on open-ended conversations that maintain significant context across multiple turns of dialog.


(Weddell Seal image cropped from Wikimedia CC licensed image.)
Transformer models are also having a major impact in image, video, and speech models, all of which also benefit significantly from scale, as predicted by work on scaling laws for visual transformer models. Transformers for image recognition and for video classification are achieving state-of-the-art results on many benchmarks, and we’ve also demonstrated that co-training models on both image data and video data can improve performance on video tasks compared with video data alone. We’ve developed sparse, axial attention mechanisms for image and video transformers that use computation more efficiently, found better ways of tokenizing images for visual transformer models, and improved our understanding of visual transformer methods by examining how they operate compared with convolutional neural networks. Combining transformer models with convolutional operations has shown significant benefits in visual as well as speech recognition tasks.
The outputs of generative models are also substantially improving. This is most apparent in generative models for images, which have made significant strides over the last few years. For example, recent models have demonstrated the ability to create realistic images given just a category (e.g., “irish setter” or “streetcar”, if you desire), can “fill in” a low-resolution image to create a natural-looking high-resolution counterpart (“computer, enhance!”), and can even create natural-looking aerial nature scenes of arbitrary length. As another example, images can be converted to a sequence of discrete tokens that can then be synthesized at high fidelity with an autoregressive generative model.

Because these are powerful capabilities that come with great responsibility, we carefully vet potential applications of these sorts of models against our AI Principles.
Beyond advanced single-modality models, we are also starting to see large-scale multi-modal models. These are some of the most advanced models to date because they can accept multiple different input modalities (e.g., language, images, speech, video) and, in some cases, produce different output modalities, for example, generating images from descriptive sentences or paragraphs, or describing the visual content of images in human languages. This is an exciting direction because like the real world, some things are easier to learn in data that is multimodal (e.g., reading about something and seeing a demonstration is more useful than just reading about it). As such, pairing images and text can help with multi-lingual retrieval tasks, and better understanding of how to pair text and image inputs can yield improved results for image captioning tasks. Similarly, jointly training on visual and textual data can also help improve accuracy and robustness on visual classification tasks, while co-training on image, video, and audio tasks improves generalization performance for all modalities. There are also tantalizing hints that natural language can be used as an input for image manipulation, telling robots how to interact with the world and controlling other software systems, portending potential changes to how user interfaces are developed. Modalities handled by these models will include speech, sounds, images, video, and languages, and may even extend to structured data, knowledge graphs, and time series data.

Often these models are trained using self-supervised learning approaches, where the model learns from observations of “raw” data that has not been curated or labeled, e.g., language models used in GPT-3 and GLaM, the self-supervised speech model BigSSL, the visual contrastive learning model SimCLR, and the multimodal contrastive model VATT. Self-supervised learning allows a large speech recognition model to match the previous Voice Search automatic speech recognition (ASR) benchmark accuracy while using only 3% of the annotated training data. These trends are exciting because they can substantially reduce the effort required to enable ML for a particular task, and because they make it easier (though by no means trivial) to train models on more representative data that better reflects different subpopulations, regions, languages, or other important dimensions of representation.
All of these trends are pointing in the direction of training highly capable general-purpose models that can handle multiple modalities of data and solve thousands or millions of tasks. By building in sparsity, so that the only parts of a model that are activated for a given task are those that have been optimized for it, these multimodal models can be made highly efficient. Over the next few years, we are pursuing this vision in a next-generation architecture and umbrella effort called Pathways. We expect to see substantial progress in this area, as we combine together many ideas that to date have been pursued relatively independently.

IBL Education’s GenAI-based chat mentor with Google

900
Of your peers have already read this article.
3:30 Minutes
The most insightful time you'll spend today!
With more than 6 years of experience in building open source and Generative AI in education at scale, ibleducation.com continues to evolve its approach and services. More recently, the company became an education-focused Vertex AI integrator. They provide enterprises and academic institutions with a platform to build, train and securely customize large language models (LLMs) for interactive mentors using Google’s Vertex AI.
“The education industry only recently began understanding the value of Generative AI with open source and the opportunities it presents,” says Miguel Amigot II, Chief Technology Officer at ibleducation.com. “We’re providing the chance to build a new type of learning, mentoring, and analytics platform that gives organizations total control over their training methods, data, and more, without locking into one vendor.”
Organizations large and small, including Fortune 500 companies and leading universities institutions, use ibleducation.com to support their learners, allow educators to develop coursework, and provide engineers with a solid platform to build on.rely on it to power their learning strategies.
Recently, ibleducation.com chose to partner with Google Cloud to improve its platform and scale beyond its base of millions of users.
Harnessing the power of GenAI in education and training
ibleducation.com initially began working with Google Cloud to unify its education data strategy to drive real-time and predictive analytics.
“Google Cloud outperforms others when it comes to maintaining control over algorithms and general data governance,” says Amigot. “We’re fortunate to be able to maintain ownership of our models while maintaining a pay-per-use pricing model. This is critical for our business. It allows our clients to avoid lower-value maintenance tasks, focus on innovation and user experience, and not worry about excessive costs.”
In the past, because of the high costs and need for AI and engineering talent to produce large language models, schools have been unable to take full advantage of open-source education technologies. Together, ibleducation.com and Google Cloud democratize access to AI-enabled tools for educators worldwide.
With a strong analytics and AI foundation, ibleducation.com has been able to personalize, translate, and create education content faster than before. This has been especially effective in improving accessibility, as automated translation and repurposing ensures those with visual and hearing impairments can interact with the content.
“Google Cloud enables us to reduce the total cost of building, running, and maintaining large language models by more than 70%,” says Amigot. “We’ve seen language model hosting costs decrease before, but nothing like this.”
Bringing AI-powered content creation to educators
Vertex AI has been especially useful in creating content, helping ibleducation.com to dramatically scale out its learning materials. Usage-based pricing mixed with an open-source model gives ibleducation.com the resources to support its users in developing evolving, valuable learning experiences to serve people worldwide.
Additionally, Google Cloud powers ibleducation.com’s Generative AI-based mentors, which provides a one-to-one learning experience to students and professionals looking for tailored skills development.
“The mission is to provide educators with a centralized system wherein they can manage everything from indexing data about customer courses to designing and facilitating UI/UX for users’ digital mentors,” says Amigot. “Google Cloud has helped us cover all of these bases.”
ibleducation.com allows users to create virtual mentors that provide personalized teaching, assess student knowledge, guide students through skills and learning paths, and offer robust learning analytics. Text-to-text and speech-to-speech interfaces can be offered over Slack, Discord, text message-based bots, web scripts, and LTI integrations.
“AI Mentor has been a very effective tool for our users because it adapts to the needs of educators and organizations very quickly,” says Amigot. “We’re seeing people use it for teaching assistance, marketing and administration, and a lot more. Vertex AI allows us to have many models for any variety of purposes.”
Looking forward, ibleducation.com is looking to continue taking personalized learning to new heights.
“From a mission standpoint, we want to expand to reach all the organizations not currently served by online education through our platform, analytics, and other AI-powered services,” says Amigot. “Our relationship with Google Cloud is instrumental in achieving our goal to democratize access to next-generation educational capabilities to more communities and organizations around the globe.”
Ready to take the next step? Join our upcoming GenAI workshop for higher education to learn how Google’s AI tools can help education institutions or explore our EdTech solutions to see how you can make education more personal, safe, and accessible with 100+ cutting-edge products.
Building Unique Customer Experiences with Speed & Scale: Sprinklr & Google Cloud

8243
Of your peers have already read this article.
2:00 Minutes
The most insightful time you'll spend today!
Enterprises are increasingly seeking out technologies that help them create unique experiences for customers with speed and at scale. At the same time, customers want flexibility when deciding where to manage their enterprise data, particularly when it comes to business-critical applications.
That’s why I’m thrilled that Sprinklr, the unified customer experience management (Unified-CXM) platform for modern enterprises, has partnered with Google Cloud to accelerate its go-to-market strategy and grow awareness among our joint customers. Sprinklr will work closely with our global salesforce, benefitting from our deep relationships with enterprises that have chosen to build on Google Cloud.
Akin to Google Cloud’s mission to accelerate every organization’s ability to digitally transform their business through data-powered innovation, Sprinklr’s primary objective is to empower the world’s largest and most loved brands to make their customers happier by listening, learning, and taking action through insights. With this strategic partnership now in place, Sprinklr and Google Cloud will go-to-market together with the end-customer as our sole focus.
Traditionally, brands have adopted point solutions to manage segments of the customer journey. In isolation, these may work — but they rarely work collaboratively, even when vendors build “Frankenstacks” of disconnected products. These solutions can’t deliver a 360° view of the customer, and often reinforce departmental silos. All of which creates point-solution chaos.
Sprinklr’s approach is fundamentally different and is the way out of the aforementioned point-solution chaos. As the first platform purpose-built for unified customer experience management (Unified-CXM) and trusted by the enterprise, Sprinklr’s industry-leading AI and powerful Care, Marketing, Research, and Engagement solutions enable the world’s top brands to learn about their customers, understand the marketplace, and reach, engage, and serve customers on all channels to drive business growth.
Sprinklr was built from the ground up as a platform-first solution, designed to evolve and grow with the rapid expansion of digital channels and applications. The results? Faster innovation. Stronger performance. And a future-proof strategy for customer engagement on an enterprise scale.

“Sprinklr works with large, global companies that want flexibility when deciding where to manage their enterprise data and consider our platform a business-critical application,” said Doug Balut, Senior Vice President of Global Alliances, Sprinklr. “Giving our customers the opportunity to manage Sprinklr on Google Cloud empowers them to create engaging customer experiences while maintaining the high security, scalability, and performance they need to run their business.”
To learn more about this exciting partnership and the challenges we jointly solve for customers, check out the recent conversation between Google Cloud’s VP of Marketing, Sarah Kennedy, and Sprinklr’s Chief Experience Officer, Grad Conn. Or read the press release on the partnership.
3565
Of your peers have already watched this video.
44:30 Minutes
The most insightful time you'll spend today!
Bra Fit or Brad Pitt? Fun Moments of Using AI for Contact Centers
For many enterprises, the wish to tap into the power of AI is negated only by a lack of know-how: How AI works, what sort and how much data they will need to gather and structure correctly, and how to use the platforms that make AI adoption easier.
Google understands that. Which is why the AI solutions they have built “are basically plug and play, ready to go with our partners, so that you (businesses) can have immediate business value for your specific use case, and for your specific workflow without a deep investment of any kind into machine learning,” says Levent Besik, Group Product Manager, Google Cloud.
Among the more interesting use cases that’s seeing adoption is contact center AI.
“So one thing that we see quite often, especially from our B2C customers, is that they often have this growing pain in their call centers. They face a trade-off between operational efficiency and great customer service. With the advances in language and conversational AI, that doesn’t have to be the case anymore,” says Besik.
That’s exactly the problem in front of Akash Parmar, Enterprise Architect,. “One of the big challenge we had was: how do we effectively manage 14 million calls, which come into our contact center, stores and head office? In the past, these calls were managed by completely different platforms with their own IVRs, with their own routing and reporting solutions. It was expensive, plus the experience across them was very inconsistent.”
Digging deeper, Parmar and team figured that the challenge was in the fact that an IVR couldn’t really capture the hundreds of reasons customers call.
To get around the problem Parmar and team decided to let customer tell them why they were calling. They then used AI to decipher what the customer was saying, extract the customer’s intent from that, and then help them directly, or re-route them to the best department.
Overall, it proved to be a big success, says Parmer. Although there were funny moments.
“In our early days, we were getting some transcriptions from the Speech API, which were not what we expected. It was not word by word and we had to kind of train the model to make sense of it. For example, we started to get calls about Brad Pitt. We were like, we’ve really won the Oscar here. Brad Pitt is calling us. It was not actually call for Brad Pitt. These were calls about bra fit, which it’s a big business for M&S,” remembers Parmer.
Find out how three companies, including Marks and Spencer, are using AI today.
More Relevant Stories for Your Company

FAQs: Everything Your Need to Know About Cloud Computing
There are a number of terms and concepts in cloud computing, and not everyone is familiar with all of them. To help, we’ve put together a list of common questions, and the meanings of a few of those acronyms. You can find all these, and many more, in our learning resources.

How Kinguin Notched Up Shopping Experience with Google Recommendations AI
Over 2.14 billion people worldwide are expected to buy online this year, according to Statista. Online retail sales will account for 22% of all purchases by 2023. But in a competitive retail landscape, positive interactions can mean the difference between a sale and an abandoned shopping cart. One of the leading global

How Google Cloud is Fortifying the Social Safety Net
The COVID-19 crisis is triggering both immediate and longer-term challenges for state and local governments. In this video, Denise Winkler, Strategic Business Executive, Google Cloud and Jennifer Ricker, Assistant Secretary - Illinois Department of Innovation & Technology (DoIT), discuss a number of challenges and solutions around the pandemic. First, they

World’s Largest Online-only Grocery Retailer Uses AI to Figure Which Customers Need Most Attention
In the United Kingdom, the popularity of online grocery shopping is expected to surge from about 6% of the market today to 9% by 2021, according to market research firm Mintel. One of the pioneers of online-only grocery retailing is Ocado, based in Hatfield, Hertfordshire in the U.K. Since starting commercial deliveries






