How Cleartrip.com is leveraging Google Cloud to survive the slump in the travel industry - Build What's Next
Case Study

How Cleartrip.com is leveraging Google Cloud to survive the slump in the travel industry

5152

Of your peers have already read this article.

5:30 Minutes

The most insightful time you'll spend today!

COVID-19 dealt a body blow to the travel and tourism industry. But as the effects of the pandemic start to wane, Manoj Sharma, CTO, Cleartrip.com, talks about how technology will be a game-changer and how it will pave the road to recovery.

With the novel coronavirus COVID-19 sweeping across continents and fatalities climbing every day, it was only a matter of time before countries closed their borders to contain its spread.

In the wake of this decision, travel and tourism, the linchpins of many economies, were among the worst affected.

According to the United Nations World Tourism Organization (UNWTO), the COVID-19 pandemic caused a 22 percent fall in international tourist arrivals during the first quarter of 2020, and could see an annual decline of between 60 percent and 80 percent when compared with 2019.

The impact on the economy and to livelihoods that are dependent on tourism and hospitality has been significant. Prior to the pandemic, the outlook was quite different. Research by UNWTO in 2018 estimated that India would have 50 million outbound tourists by 2020.

Technology was also set to play a huge part in that growth. A Google Travel study showed that 74 percent of travellers were planning their trips on the Internet, and technologies like AI, IoT and VR were all set to be key trends this year.

Now, as borders slowly reopen and travel restrictions are gradually lifted, technology could once again be the game-changer. Manoj Sharma CTO, Cleartrip.com spoke to YourStory about the industry’s road to recovery and how technology will aid that journey.

Read the Full Story on YourStory

Blog

Google Introduces ML-based Predictive Autoscaling to Forecast Capacity and Match Scaling Demands

4987

Of your peers have already read this article.

3:00 Minutes

The most insightful time you'll spend today!

Google Cloud's predictive autoscaling makes the infrastructure scaling process more proactive! End unpredictability by forecasting scaling capacity in advance, and match the demands, creating VMs with enough time for applications to initialize.

At Google Cloud, we believe you get most benefits from the cloud when you scale infrastructure based on changing demand. Compute Engine allows you to configure autoscaling to save costs during periods of low demand, and add capacity to support peak loads. 

When you use a managed instance group (MIG), you can have an autoscaler automatically create or delete virtual machine (VM) instances based on increases or decreases in load. However, if your application takes several minutes to initialize, creating VMs in response to growing load might not increase your application’s capacity quickly enough. For example, if there’s a large increase in load (like when users first wake up in the morning), some users might experience delays while your application is initializing on new instances.

A good way to solve this problem would be to create VMs ahead of demand so that your application has enough time to initialize beforehand. This requires knowing upcoming demand. If only we could predict the future… Well, now we can!

Introducing predictive autoscaling

Predictive autoscaling uses Google Cloud’s machine learning capabilities to forecast capacity needs. It creates VMs ahead of growing demand allowing enough time for your application to initialize.

Figure 1.jpg
Figure 1. Autoscaling creates VMs as demand grows leaving no buffer for application to initialize. Predictive autoscaling creates VMs ahead of demand allowing enough time for your application to initialize and start serving new load.

How does it work?

Predictive autoscaling uses your instance group’s CPU history to forecast future load and calculate how many VMs are needed to meet your target CPU utilization. Our machine learning adjusts the forecast based on recurring load patterns for each MIG. 

You can specify how far in advance you want autoscaler to create new VMs by configuring the application initialization period. For example, if your app takes 5 minutes to initialize, autoscaler will create new instances 5 minutes ahead of the anticipated load increase. This allows you to keep your CPU utilization within the target and keep your application responsive even when there’s high growth in demand. 

Many of our customers have different capacity needs during different times of the day or different days of the week. Our forecasting model understands weekly and daily patterns to cover for these differences. For example, if your app usually needs less capacity on the weekend our forecast will capture that. Or, if you have higher capacity needs during working hours, we also have you covered.

Why should you try it?

Predictive autoscaling continuously adapts forecasted capacity to best match upcoming demand. Autoscaler checks the forecast several times per minute and creates or deletes VMs to match its prediction. The forecast itself is updated every few minutes to match recent load trends so if your growth rate is higher or lower than usual we will adjust the forecast accordingly. This gives you capacity needed to cover peak load while saving on cost when demand goes down. 

You can start using predictive autoscaling without worry as it’s fully compatible with the current autoscaler. Autoscaler will calculate enough VMs to cover both forecasted as well as real-time CPU load—whichever is higher. This works with other autoscaling features as well: you can scale based on schedule, your Load Balancer request target or Cloud Monitoring metrics. Autoscaler provides enough capacity to all of your configurations by taking the highest number of VMs needed to meet all your targets.

Getting started

You can enable predictive autoscaling in the Google Cloud Console. Select an autoscaled MIG from the instance groups page and click Edit group. Change predictive autoscaling configuration from Off to Optimize for availability.

compute google console.jpg

To better understand whether predictive autoscaling is good for your application, click the link See if predictive autoscaling can optimize your availability. This will show you a comparison of the last seven days with your current autoscaling configuration vs. with predictive autoscaling enabled.

instance group autoscaling.jpg

In the above chart, 

  • Average VM minutes overloaded per day shows how often your VMs exceed your CPU utilization target. This happens when demand is higher than available capacity. Predictive autoscaling can reduce this by starting VMs ahead of anticipated load. 
  • Average VMs per day is a proxy for cost. This shows how much additional VM capacity you need to keep your CPU utilization within the target you have set. You can optimize your cost by adjusting Minimum instances andCPU utilization as explained below. 

Optimizing your configuration

Make sure your Cool down period reflects how long it takes for your application to initialize from VM boot time until it’s ready to serve the load. Predictive autoscaling will use this value to start VMs ahead of forecasted load. If you set it to 10 minutes (600 seconds) your VMs will start 10 minutes before the load is expected to increase.

Review your autoscaling CPU utilization target and Minimum number of instances. With predictive autoscaling you no longer need a buffer to compensate for the time it takes for a VM to start. If your application works best at 70% CPU utilization you don’t need to set target to a much lower value as predictive autoscaling will start VMs ahead of usual load. A higher CPU utilization and lower Minimum number of instances allows you to reduce the cost as you don’t need to pay for additional capacity to prepare for growing demand.

Try predictive autoscaling today

Predictive autoscaling is generally available across all Google Cloud regions. For more information on how to configure, simulate and monitor predictive autoscaling, consult the documentation.

2857

Of your peers have already watched this video.

16:00 Minutes

The most insightful time you'll spend today!

How-to

Conversational AI in Search, Maps and Online Shopping!

Did you know about 77 percent of customers are likely to make a purchase from a brand they can message with? Direct interaction with the brand to gather product information shortens buyers’ journey and personalizes it with appropriate messages. To helps businesses add speed, simplicity and convenience in brand-customers interaction, Google’s Business Messages helps add chat feature in Search, Maps or other mediums.

Watch the video to learn to integrate conversational AI based on Google’s superior AI and ML features and build interactive and connected chat automation using Google Cloud’s suite of tools and products!

Blog

You Can Now ‘Listen’ to Over 50 Tech Blogs on Google Cloud Reader

4876

Of your peers have already read this article.

1:30 Minutes

The most insightful time you'll spend today!

You can now give your eyes some rest and yet catch up on the Google's latest tech blogs in audio format with Google Cloud Reader. Listen to your favorite from the 50 blogs or episodes on Google Podcasts, Apple Podcasts and Spotify.

🎧 Prefer to listen? Check out this episode on the Google Cloud Reader podcast

If you’re anything like me, you love reading, but also appreciate that sometimes your eyes need to be doing other things; whether it’s finding your exit off the highway, or keeping your puppy from destroying the couch.

And sometimes the thought of sitting down to read something just feels like it’s going to take valuable multi-tasking time away from my day. I know, I know, multitasking can be frowned upon, but it’s the way I live a good chunk of my life, and it’s working out so far. And while I’m not alone in my multitasking, I’m also not alone in my desire for a non-visual way to get this content, or any content.

*Google Cloud Reader enters the chat*

Google Cloud Reader is a podcast that lets you listen to the Google Cloud Blog posts that aren’t as dependent on visuals. This means they’re articles that are, or are adapted to be, less focused on graphs, or code samples, and instead describe the meaning behind those visual aids. 

It’s an easy, audible way to absorb content around all things new in Cloud, while still being able to make sure Ruthie doesn’t eat my work from home equipment. 

ruthie
Ruthie, a French Shepherd puppy, with her giant ears and feet dangerously close to filming equipment

So by now you’re probably thinking “OK, so you started a podcast during the pandemic, even though you definitely seemed like the type to start making sourdough”—and you’re right. My 53 plants agree with you. But rest assured, one can listen to an episode of this podcast *while* creating a macramé plant hanger, or waiting for bread to rise—multitasking, am I right?

We’re a little over 50 episodes/macrame plant hangers in, so you should check it out (Ruth and I would appreciate it).

Some of my personal favorites 

  • Beginners Guide to Painless Machine Learning – Learn how to get started with Google Cloud AI tools
  • Introducing GKE Autopilot: A Revolution in Managed Kubernetes – Learn more about GKE Autopilot, a revolutionary mode of operations for managed Kubernetes that lets you focus on your software, while GKE Autopilot manages the infrastructure.
  • Cook up your own ML recipes with AI Platform – ​​Learn about Mars Wrigley’s new ML-inspired recipe experiment on Google Cloud and how you can get started with your own.
  • Recovering Global Wildlife Populations using ML – Review Google’s Wildlife Insight’s ML project and help users create an image classification model for motion-sensor cameras (called camera traps) used to help protect wildlife in an non-invasive way by collecting and tagging species via pictures.

Let me know your favorite episodes, and what other articles you’d like to hear on Twitter @jbrojbrojbro!

No matter why you prefer an audio format, we’ve got you covered; Google Cloud Reader, where we read the tech blog for you, and to you.

Get all the Google Cloud Reader on your favorite podcast platform, including Google PodcastsApple Podcasts, and Spotify.

Blog

Startup Success Blueprint: Insights on Cloud Provider Selection from One AI

908

Of your peers have already read this article.

4:30 Minutes

The most insightful time you'll spend today!

Explore the world of generative AI through the lens of One AI, a rising startup, as they share insights on the critical aspects of selecting a cloud provider. Learn how MongoDB on Google Cloud empowers startups to accelerate their AI journey.

From the newsroom to the boardroom, everywhere we turn these days the topic of conversation is artificial intelligence (AI). From the smallest startups to the largest enterprises, every business is looking for ways to incorporate generative AI technology into their products or services. 

Generative AI is a new breed, able to not only discern patterns in data but generalize from them and create new data—and it’s moving fast. With new technologies rapidly emerging, there’s a lot to consider when choosing a tech stack that’s right for your startup. 

Our mission is to empower startups to build generative AI applications quickly, efficiently, and responsibly. In addition to our technology, we offer a variety of fresh educational and consulting initiatives, as well as comprehensive plans tailored to specific industry applications. 

At the recent Google Cloud Startup Summit, leaders from One AI and MongoDB weighed in on various decisions startups face when choosing a development platform to integrate AI into their products. (You can watch the full conversation here.) They shared key insights into the challenges startups face when applying AI technology to real-world business and product use cases. 

In this blog, we’ll explore why One AI – a platform that empowers businesses to deploy tailored AI solutions – relies on MongoDB on Google Cloud for performance, scale, functionality, and TCO. We will also cover the things that startups should keep in mind when building their generative AI stack.

Fine-tuning generative AI for startups

First, let’s start with a little background. 

AI-based capabilities have been around for decades now, but generative AI is distinct. It’s a more mature version of AI, powered by models pre-trained on very large datasets composed of massive quantities of images, text, and data. These models are known as foundation models, and they include large language models (LLMs), text-to-image models, multimodal models, and more. Foundation models let intelligent applications  generate new images, text, and data based on queries or prompts. This in turn allows companies to deploy those capabilities in their products and services faster since they don’t have to retrain the whole AI from scratch. 

Over the course of their diverse startup careers, Amit Ben, CEO at One AI and his team have built AI-based capabilities from the ground up for various products in various fields.

“And each time, we had to rebuild the tech stack over again,” Amit explains. “With the advent of generative AI, startups now have the ability to deploy much faster — with a lower TCO and higher confidence — and deliver the capabilities they need into their products and services. It finally makes sense for every company to have AI in its product portfolio.”

“For us to be able to focus on that,” Amit adds, “we need to make sure we have a rock-solid foundation that we can build on.”

That foundation is MongoDB on Google Cloud. 

Helping startups build fast and with flexibility

Startups can scale from ideation to growth with Google Cloud’s global availability, market-leading sustainability, and the same zero-trust security model that Google itself depends on.

With MongoDB, startups can take advantage of iteration cycles that are 3-5X faster, reduce sprawl and complexity, and benefit from the scalable infrastructure and advanced analytics tools on Google Cloud

With MongoDB on Google Cloud, Amit and his team are confident they can adapt to new schemas, to new data, and to the scale they need for both writing and reading, all while operating on a scalable platform and infrastructure they can rely on for the long haul. 

A common mistake for startups is turning to niche, single-point solutions. But this can backfire when they realize their solution doesn’t provide the security, scalability, and performance they need to grow. 

Additionally, startups tend to have tight iteration cycles as they find their ideal product market fit. The ability to build fast and with improved flexibility is a key differentiator in a startup environment. 

From cutting-edge automation to rock-solid redundancy and performance, there are many reasons why startups choose MongoDB on Google Cloud. And now, a dedicated partnership helps startups like One AI scale more quickly, more securely, and more successfully. 

Google Cloud and MongoDB for startups

Choosing the right technology to accelerate time to market is critical to a startup’s success. Not only is it easy to get started, but Google Cloud and MongoDB also provide the foundation for users to scale without limits — so startups can focus on innovating and growing their businesses. 

Learn more about the powerful startup programs available from Google Cloud and MongoDB.

Case Study

Cohere uses Google Cloud’s new TPU v4 Pods on its quest to create larger and more powerful language models

2814

Of your peers have already read this article.

3:30 Minutes

The most insightful time you'll spend today!

Cohere has entered into a multi-year tech partnership with Google Cloud. With this liaison, Cohere will leverage Google Cloud’s advanced AI and ML infrastructure and custom-designed machine learning chips optimized for large-scale ML.

Over the past few years, advances in training large language models (LLMs) have moved natural language processing (NLP) from a bleeding-edge technology that few companies could access, to a powerful component of many common applications. From chatbots to content moderation to categorization, a general rule for NLP is that the larger the model, the greater the accuracy it’s able to achieve in understanding and generating language.

But in the quest to create larger and more powerful language models, scale has become a major challenge. Once a model becomes too large to fit on a single device, it requires distributed training strategies, which in turn require extensive compute resources with vast memory capacity and fast interconnects. You also need specialized algorithms to optimize the hardware and time resources.

Cohere engineers are working on solutions to this scaling challenge that have already yielded results. Cohere provides developers a platform for working with powerful LLMs without the infrastructure or deep ML expertise that such projects typically require. In a new technical paper, Scalable Training of Language Models using JAX pjit and TPUv4, engineers at Cohere demonstrate how their new FAX framework deployed on Google Cloud’s recently announced Cloud TPU v4 Pods addresses the challenges of scaling LLMs to hundreds of billions of parameters. Specifically, the report reveals breakthroughs in training efficiency that Cohere was able to achieve through tensor and data parallelism.

This framework aims to accelerate the research, development, and production of large language models with two significant improvements: scalability and rapid prototyping. Cohere will be able to improve its models by training larger ones more quickly, delivering better models to its customers faster. The framework also supports rapid prototyping of models that address specific objectives — for example, creating a generative model that powers customer-service chatbot — by experimenting and testing new ideas. The ability to switch back and forth among model types and optimize for different objectives will ultimately allow Cohere to offer models optimized for particular use cases.

The FAX framework relies heavily on the partitioned just-in-time compilation (pjit) feature of JAX, which abstracts the relationship between device and workload. This allows Cohere engineers to optimize efficiency, and performance by aligning devices and processes in the ideal configuration for the task at hand. Pjit works by compiling an arbitrary function into a single program (an XLA computation), that runs on multiple devices — even those residing on different hosts.

Cohere’s new solution also takes advantage of Google Cloud’s new TPU v4 Pods to perform tensor parallelism. which is more efficient than the earlier pipeline parallelism implementation. As the name suggests, the pipeline parallel approach uses accelerators in a linear fashion to scale a workload, like a single long assembly line. Accelerators must process each micro-batch of data before passing it along to the next one, and then run the backward pass in reverse order.

Tensor parallelism eliminates the accelerator idle time of pipeline parallelism, also known as the pipeline bubble. Tensor parallelism involves partitioning large tensors (mathematical arrays that define the relationship among multiple objects such as the words in a paragraph) across accelerators to perform computations at the same time on multiple devices. If pipeline parallelism is an ever-lengthening assembly line, tensor parallelism is a series of parallel assembly lines — one making the engine, the other the body, etc. — that simultaneously come together to form a complete car in a fraction of the time.

These computations are then collated, a process made practical thanks to Google Cloud TPU v4 VMs, which more than double the computational power. The superior performance of v4 chips has enabled Cohere to iterate on ideas and validate them 1.7X faster in computation than before.

At Cohere, we build cutting-edge natural language processing (NLP) services, including APIs for language generation, classification, and search. These tools are built on top of a set of language models that Cohere trains from scratch on Cloud TPUs using JAX. We saw a 70% improvement in training time for our largest model when moving from Cloud TPU v3 Pods to Cloud TPU v4 Pods, allowing faster iterations for our researchers and higher quality results for our customers. The exceptionally low carbon footprint of Cloud TPU v4 Pods was another key factor for us.


Aidan Gomez
CEO and co-founder, Cohere

Why Google Cloud for LLM training?

As part of a multiyear technology partnership, Cohere leverages Google Cloud’s advanced AI and ML infrastructure to power its platform. Cohere develops and deploys its products on Cloud TPUs, Google Cloud’s custom-designed machine learning chips that are optimized for large-scale ML. Cohere’s recently announced their new model improvements and scalability by training an LLM using FAX on Google Cloud TPUs, and this model has demonstrated that transitioning from TPU v3 to TPU v4 has so far enabled them to achieve a total speedup of 1.7x. In addition to a significant performance boost, TPUs provide an excellent user experience with the new TPU VM architecture. Importantly, Google Cloud ensures that Cohere’s state-of-the-art ML training is achieved with the highest standards of sustainability, powered by 90% carbon-free energy in the world’s largest publicly available ML hub.

By adopting Cloud TPUs, Cohere is making LLM training faster, more economical, and more agile. This helps them provide larger and more accurate LLMs to developers, and put NLP technology in the hands of developers and businesses of all sizes.

To learn more about these LLM training advances, you can read the full paper, Scalable Training of Language Models using JAX pjit and TPUv4. To learn more about Cohere’s best practices and AI principles, you can check this article co-authored with Open AI and AI 21 Labs.

More Relevant Stories for Your Company

Blog

Enabling Real-time AI with Streaming Ingestion in Vertex AI

Many machine learning (ML) use cases, like fraud detection, ad targeting, and recommendation engines, require near real-time predictions. The performance of these predictions is heavily dependent on access to the most up-to-date data, with delays of even a few seconds making all the difference. But it’s difficult to set up

How-to

AutoML Vision: Among the Fastest and Easiest Way to Adopt AI for Your Enterprise

What’s among the largest impediment to the adoption of AI within enterprises? Not enough access to skills. According to 80 percent of business respondents to an EY survey, the top challenge to an enterprise AI program is the lack of requisite talent. What companies need today is a way to

Case Study

Beany’s Cloud-Based Accounting Solutions Transform Small Business Finance

Many people start a small business that aligns with their passions, but soon discover the day-to-day running of a business is very different than anticipated. Dealing with accounting, finance, and other daily activities can quickly overwhelm even the most promising of new businesses. Recognizing the unique challenges facing small businesses,

Blog

Notebook Executor Feature of Vertex AI Workbench to Schedule Notebooks Ad Hoc or on Recurring Basis

When solving a new ML problem, it’s common to start by experimenting with a subset of your data in a notebook environment. But if you want to execute a long-running job, add accelerators, or run multiple training trials with different input parameters, you’ll likely find yourself copying code over to

SHOW MORE STORIES