Accelerating AI Inference at Scale: Introducing Google Cloud TPU v5e

944
Of your peers have already read this article.
4:30 Minutes
The most insightful time you'll spend today!
Google Cloud’s AI-optimized infrastructure makes it possible for businesses to train, fine-tune, and run inference on state-of-the-art AI models faster, at greater scale, and at lower cost. We are excited to announce the preview of inference on Cloud TPUs. The new Cloud TPU v5e enables high-performance and cost-effective inference for a broad range AI workloads, including the latest state-of-the-art large language models (LLMs) and generative AI models.
As new models are released and AI becomes more sophisticated, businesses require more powerful and cost efficient compute options. Google is an AI-first company, so our AI-optimized infrastructure is built to deliver the global scale and performance demanded by Google products like YouTube, Gmail, Google Maps, Google Play, and Android that serve billions of users — as well as our cloud customers.
LLM and generative AI breakthroughs require vast amounts of computation to train and serve AI models. We’ve custom-designed, built, and deployed Cloud TPU v5e to cost-efficiently meet this growing computational demand.
Cloud TPU v5e is a great choice for accelerating your AI inference workloads:
- Cost Efficient: Up to 2.5x more performance per dollar and up to 1.7x lower latency for inference compared to TPU v4.
- Scalable: Eight TPU shapes support the full range of LLM and generative AI model sizes, up to 2 trillion parameters.
- Versatile: Robust AI framework and orchestration support.
In this blog, we’ll dive deeper into how you can leverage TPU v5e effectively for AI inference.
Up to 2.5x more performance per dollar and up to 1.7x lower latency for inference
Each TPU v5e chip provides up to 393 trillion int8 operations per second (TOPS), allowing complex models to make fast predictions. A TPU v5e pod consists of 256 chips networked over ultra-fast links. Each TPU v5e pod delivers up to 100 quadrillion int8 operations per second, or 100 PetaOps, of compute power.
We optimized the Cloud TPU inference software stack to take full advantage of this powerful hardware. The inference stack leverages XLA, Google’s AI compiler, which generates highly-efficient code for TPUs to maximize performance and efficiency.
The combined hardware and software optimizations, including int8 quantization, enable Cloud TPU v5e to achieve up to 2.5x greater inference performance per dollar than Cloud TPU v4 on state-of-the-art LLM and generative AI models, including Llama 2, GPT-3, and Stable Diffusion 2.1:

Google Internal Data. August 2023. Normalized to single-chip throughput. Precision: Llama 2 7B, 13B, 70B, GPT-J 6B: int8; GPT-J 175B, Stable Diffusion 2.1: bf16.
On latency, Cloud TPU v5e achieves up to 1.7x speedup compared to TPU v4:

Google Internal Data. August 2023. Precision: Llama 2 7B, 13B and 70B: int8; GPT-3 175B: bf16.
Google Cloud customers have been running inference on Cloud TPU v5e, and some have seen even greater speedups on their particular workloads.
AssemblyAI offers dozens of AI models to their customers for speech recognition and understanding with over 25 million inference calls on a daily basis.
“Cloud TPU v5e consistently delivered up to 4X greater performance per dollar than comparable solutions in the market for running inference on our production model. The Google Cloud software stack is optimized for peak performance and efficiency, taking full advantage of the TPU v5e hardware that was purpose-built for accelerating the most advanced AI and ML models. This powerful and versatile combination of hardware and software dramatically accelerated our time to solution: instead of spending weeks hand-tuning custom kernels, within hours we optimized our model to meet and exceed our inference performance targets.” – Domenic Donato, VP of Technology, AssemblyAI
Scale to the full range of LLM and Generative AI model sizes
LLMs and generative AI models continue to grow in size and computational cost. The largest models require the combined compute and memory of hundreds of hardware accelerators. Cloud TPU v5e enables inference for a wide range of model sizes. A single v5e chip can run models with up to 13B parameters. From there, you can scale up to hundreds of chips and run models with up to 2 trillion parameters.

Google Internal Data. August 2023. Batch size = 1. Multi-head attention based decoder only language models: prefix length = 2048, decode steps = 256, beam size = 32 for sampling.
Gridspace leverages Google Cloud TPU infrastructure to power its full-stack conversational AI platform – building and integrating real-time conversational ASR, LLMs, semantic search, and neural TTS.
“We’re a huge fan of Google Cloud TPUs. Our benchmarks are demonstrating a 5X increase in the speed of AI models when training and running on Google Cloud TPU v5e. We are also seeing a 6x improvement in the scale of our inference metrics. We’ve scaled our AI models to billions of conversations per year across financial services, capital markets, and healthcare with Google Cloud’s AI infrastructure. Our Grace bots are powered by models trained using Cloud TPUs and served at scale on GKE with support for PCI, HITRUST, and SOC 2 compliance.” – Wonkyum Lee, Head of Machine Learning, Gridspace
Robust AI framework and orchestration support
Leading AI frameworks, including PyTorch, JAX, and TensorFlow, provide robust support for inference on Cloud TPU v5e. This means you can now train and serve models end-to-end on Cloud TPUs: what you train is what you serve.

Google Cloud offers you many choices to run inference on Cloud TPUs easily and reliably. From GKE and Vertex AI, to popular open-source frameworks such as Ray and Slurm, you can leverage Google Cloud TPUs in your preferred way to fit your development process.

Try Cloud TPU v5e for inference today
Cloud TPU v5e provides a high-performance, cost-efficient, scalable, and reliable inference platform for LLMs and generative AI models. Leading AI companies are leveraging the power of Cloud TPU v5e to serve AI models at scale:

To get started with inference on Cloud TPU, reach out to your Google Cloud account manager or contact Google Cloud sales.
withVR Uses the Power of VR to Prep People with Speech Disorders for Real-life Speaking Situations

3273
Of your peers have already read this article.
1:30 Minutes
The most insightful time you'll spend today!
Editor’s note: Meet Gareth Walkom, an entrepreneur dedicated to helping others with speech disorders.
Turning life experience into innovation
Did you know that 3% of Americans have a speech disorder, while 1% of the world’s population have a stutter? Just getting what they need in everyday interactions can be stressful, which intensifies when the stakes are raised during job interviews, presentations, public speaking, and other activities. As a result, some people with speech disorders may avoid conversations and relationships, and risk being denied jobs because of a difference in how they speak.
As a person who stutters, I know firsthand the ableism that people with a speech disorder can encounter in wanting to use their voice in a judgemental world: the frustration of sometimes not being able to say exactly what you want to say and therefore speaking less in speaking situations. And the educational and career opportunities are lost when doors remain closed to us, especially when employers advertise their jobs as requiring someone who speaks the language ‘fluently’.
While researchers still don’t definitively know what causes stuttering, emerging technologies are giving us new and promising pathways for improving the quality of life of people with speech disorders.
That’s why after years of researching and testing potential therapeutic uses of virtual reality, and with the support of the Google for Startups Cloud Program, I founded withVR on International Stuttering Awareness Day (October 22) in 2020. The mission of withVR is to prepare people with speech disorders for real-life speaking situations by utilizing the power of virtual reality.
Working through it
One of the difficulties in adapting to any disability is the opportunity to work through it in a safe and nonjudgmental environment. withVR provides a virtual space for individuals, in collaboration with their speech therapists, to practice real-world speaking scenarios in safe, controlled environments.
Imagine being able to raise your hand in class and give your opinion without hesitation, ordering the meal you want rather than something that’s easier to say, or sit across the table from an avatar of an employer and explain why you are the right person for your dream job. Then further customize the speaking situation and its surroundings to challenge yourself and be ready for anything. That’s what withVR offers individuals and their speech therapists.
Making VR come to life
To bring the withVR vision to life, we are developing applications using the Unity game engine on Google Cloud with integrated Firebase services including authentication, web hosting, storage, and database. It’s a powerful combination that’s enabled us to build industrial-strength applications that we’ve rapidly deployed on a global basis. Today we are already collaborating with 80+ labs, clinics, and hospitals in more than 20 different countries worldwide.
These organizations help us to test and refine a virtual reality application to support people in achieving their speech goals and build comfort through immersive VR experiences using easily available viewers like Google Cardboard.
The application works in conjunction with a web app through which speech therapists configure customized VR scenarios for their patients to use. As no real-life speaking situation is ever exactly the same, customization of VR scenarios is vital. They can construct different scenes, create and script avatars, and through the Google Text-to-Speech API can even choose from hundreds of different voices in a variety of languages. This gives them the flexibility to create many unique speaking situations for their clients no matter where they are in the world.
Progress from the practice sessions is presented through a dashboard that provides therapists with a tool to monitor their clients’ progress and provide feedback and encouragement.
No shortage of support
My founder’s journey has been supported by many passionate people. The Google for Startups Cloud Program has been instrumental in helping us come so far in the first year, and we’ve only scratched the surface of what’s possible. There are many capabilities in Firebase and Google Cloud that we have yet to explore, and through the startup program I now have a Google Mentor who can help guide that exploration.
We also joined the 2Gether-International (2GI) Tech Cohort, which is supported by Google for Startups and is built for and run by entrepreneurs with disabilities. At the end of the 10-week cohort, we finished with a pitch competition, where I was one of six selected founders to pitch in just three minutes. I was very fortunate to win the Best Overall Pitch Award, gaining 10,000 USD in seed funding. This award not only highlights the potential of withVR, but also showcases that anyone can pitch their idea in a short amount of time no matter their difference.
Working with 2GI also gave me the opportunity to collaborate and learn from other founders who have disabilities. It’s a safe space where I don’t have to explain my everyday challenges and can focus on the all-important task of advancing the vision of withVR, while seeing how others use technology in their domain.
Building on a strong foundation
I’m amazed to look back and see what we’ve accomplished in just one year and humbled by the thousands of lives we’ve touched. Every day we receive valuable real-life feedback from people in the field—both clinicians and those with speech differences who benefit from VR therapy. That knowledge tells us that we are heading in the right direction and opening our eyes to new possibilities for where to take withVR. And inspiring us to keep moving ahead.
If you’d like to participate in testing or if you are speech therapist or researcher, please feel free to reach out to us. We’d love to show you how you can contribute to a world where anyone with a speech disorder can truly use their voice in any situation. If you’d like to take part, contact us at hello@withvr.app.
If you want to learn more about how Google Cloud can help your startup, visit our page here to get more information about our program, and sign up for our communications to get a look at our community activities, digital events, special offers, and more.

4458
Of your peers have already downloaded this article
1:30 Minutes
The most insightful time you'll spend today!
Increasingly, technology, data, and human interactions will follow a new model of real-time inputs, enabling change without interruption and positive iteration. Cloud computing powers an important shift for information technology, and it will usher in new and better ways to serve customers, make discoveries, and build great products and services.
As part of this transformation, technology models will become more flexible, interoperable, and open; office cultures will shift toward transparency, collaboration, and constant learning. That means innovation and team building – the best parts of work – can happen more frequently and efficiently. Customer focus can sharpen and become more responsive. Corporate knowledge gained over decades of competing and innovating can be brought to bear more easily and meaningfully.
To compete effectively in this changing landscape, business and technology leaders must find ways to embrace the potential of cloud computing – whether in the tools and systems they choose, the institutional cultures they create, or the business strategies they prioritize. Today’s enterprises can set the stage for future success by infusing their unique assets – such as customer relationships, excellent service capabilities, strong partnerships, market knowhow, and technology investment – with the power of innovative IT.
Indian social media platform ShareChat saw a 50% reduction in costs after migrating to Google Cloud

4353
Of your peers have already read this article.
6:30 Minutes
The most insightful time you'll spend today!
One of the key trends that emerged from the lockdown brought about by the COVID-19 outbreak was the significant rise in people seeking entertainment online.
According to data released by Carat India, smartphone usage increased by 1.5 hours a week and social media consumption nearly doubled to 280 minutes a day. Of India’s 670 million internet users, a significant percentage live in rural areas, which saw a surge in demand for local language content.
As India’s premier social media platform, ShareChat saw significant engagement from its 60 million monthly active users (MAU) during the lockdown. The platform now has surpassed 120 million MAU post the government banning 59 Chinese apps in the last week of June.
ShareChat was founded by Ankush Sachdeva (co-founder and Chief Executive Officer), Farid Ahsan (co-founder and Chief Operating Officer), and Bhanu Pratap Singh (co-founder and Chief Technical Officer), in 2015.
In a recent conversation with YourStory, Bhanu spoke about content consumption during COVID-19, how they are reaching out to the country’s vernacular audience and their recent transition to Google Cloud.
Read the Full Story on YourStory
Google Cloud Announces Improvements in Private Catalog to Drive Terraform Deployments

3365
Of your peers have already read this article.
3:00 Minutes
The most insightful time you'll spend today!
As an enterprise admin, when you choose to use Google Cloud Private Catalog to enable curated, self-serve Google Cloud infrastructure provisioning, you need the ability to manage your organization’s deployments. Today, we’re pleased to announce support for several improvements to Terraform driven deployments through Private Catalog.
With this new release, you can update Terraform configurations and keep your end users informed about updates. At the same time, Private Catalog users have the ability to view new updates, note version highlights and then update the deployment. This gives you greater control over managing deployments for solutions provisioned through Private Catalog and ensuring compliance with organizational policies and standards.
Let’s take a closer look at the features you’ll find in this release.
Deployment change management
Terraform solutions use Cloud Storage’s Object Versioning to manage updates to configuration files. With this release, you may update configuration files using multiple approaches.
- Update the solution’s Cloud Storage object with a new configuration version
- Use a different Cloud Storage object that contains a new configuration file
Once you view and apply the changes to the solution in a Private Catalog, end users are immediately able to consume the new version of the deployment configuration.

Additionally, prior to applying any changes, you can evaluate the contents of an update by comparing versions to download and compare the current and latest versions of the configuration and use new version highlights to add a description about the updates.


Ease of consumption
Once Private Catalog detects a change to the deployment configuration, it automatically informs catalog users about the change. On the Solutions page, end users have the ability to:
- Get informed about solutions that have updates
- View version highlights published by the admin
- Apply the new version
Additionally, with this release, Catalog users can retry existing deployments by modifying deployment parameters.
Reporting improvements
The deployment reporting dashboards for Private Catalog-based deployments now show additional information about the version of a solution deployed. This enables deeper insights into the overall deployment status across all Private Catalog solution assets.


Get started today
These new features are available to all Private Catalog customers. To learn how to use these features, refer to our documentation:
- Create a Terraform configuration in Private Catalog
- Manage and update your Terraform configurations in Private Catalog
4697
Of your peers have already watched this video.
10:30 Minutes
The most insightful time you'll spend today!
How Google’s Customer Data Platform Helps Retail Brands Offer Data-driven , Personalized CX
Retail companies need customer insights to deliver personalized experiences that impact revenue generation and cost savings. Watch how Google Cloud’s customer data platform helps brands integrate and build holistic view of data in silos to drive marketing and customer service success.
More Relevant Stories for Your Company

Google Unveils New Cloud Region in Delhi NCR to Power India’s Digitization
In the past year, Google has worked to surface timely and reliable health information, amplify public health campaigns, and help nonprofits get urgent support to Indians in need. Now, we are continuing to focus on helping India’s businesses accelerate their digital transformation, deepening our commitment to India’s digitization and economic recovery.

An Indian Example of How to Really Up Your Customer Experience Game and Increase Conversion Rates With AI
How about selfie analysis of users to recommend them the right lipstick color? That’s just one of the many ideas folks at Purplle.com came up with to improve the buying experience of Indian consumers. And without the power of Google Cloud, it would probably have remained just that…an idea. But

Cloud FinOps Breaks Down Gaps in Finance, Tech and Business Teams, Accelerating Digital Transformation!
Accelerating digital transformation Digital transformation is what propels businesses and industries forward. Organizations of all sizes—from startups to global enterprises—focus on digital transformation not only to make scaled improvements, but also to drive significant change and fully embrace the digital age. The pandemic has jump started and pushed many organizations

This Chart, from Home Depot, Dramatically Demonstrates the Power of a Cloud Data Warehouse
The Home Depot (THD) is the world’s largest home-improvement chain, growing to more than 2,200 stores and 700,000 products in four decades. Much of that success was driven through the analysis of data. This included developing sales forecasts, replenishing inventory through the supply chain network, and providing timely performance scorecards. However,






