Manipal Group: Delivering High-Quality Patient Care with Google Cloud - Build What's Next
Case Study

Manipal Group: Delivering High-Quality Patient Care with Google Cloud

6488

Of your peers have already read this article.

4:30 Minutes

The most insightful time you'll spend today!

Manipal Group of Hospitals deployed a mobile application that automates nurse rostering, reducing personnel requirements, costs, and stress on nurses and freeing up senior nurses for more valuable tasks.

One of India’s best-known healthcare brands, Manipal Group prides itself on clinical excellence and a patient-centric approach. From humble beginnings in 1953 as a single teaching hospital—the Kasturba Medical College—in a university town in Karnataka, India, Manipal Group has grown to a presence in seven cities in India and operations in Malaysia and Nigeria. “Our founder, Dr T. M. A. Pai, established Kasturba Medical College just six years after India gained independence,” says C. G. Muthana, Chief Operating Officer, Teaching Hospitals, Manipal Health Enterprises Pvt. Ltd, part of Manipal Group.

Manipal Health Enterprises Pvt. Ltd operates 11 corporate, or for-profit, hospitals and five teaching hospitals. “We have about 7,000 beds India-wide, and on that measure, we are the third-largest private-sector healthcare provider in India today,” says Muthana. The organization now employs about 6,000 people in its corporate hospitals and 5,000 people in its teaching hospitals.

Manipal Health Enterprises Pvt. Ltd acquires and operates world-class medical technologies in its teaching and corporate hospitals. However, with 75% of corporate hospital wards allocated to fee-paying private patients—compared to just 25% of the wards in teaching hospitals—the differences in information technology budgets are significant. “In the general wards that comprise most of the wards in teaching hospitals, patients are typically treated for free or at heavily subsidized rates,” explains Muthana. “So revenues and costs of delivery vary between hospitals, while employee costs remain similar.”

Rostering nurses a critical task

Rostering nurses to work shifts is one of the most important tasks at Manipal Group corporate and teaching hospitals. As at all hospitals, nurses administer medications, monitor patients, maintain records, manage intravenous lines, and work with doctors to heal patients. They can also provide advice and support to patients and loved ones, including teaching them how to administer medication outside a hospital setting. If too few nurses are rostered for a particular shift, the quality of patient care may suffer.

Google Cloud results

  • Enabled the hospital to correctly size its nursing workforce
  • Lowers stress on nurses by delivering more equitable rostering
  • Presents opportunity to better manage nurses’ leave and other administration tasks

However, rostering at the group’s hospitals was a time-consuming exercise. Senior nurses on each ward would have to spend up to 45 minutes per day manually amending paper-based rosters to accommodate requested changes to duty shifts for personal or other circumstances. The organization began receiving complaints from patients, doctors, and other hospital staff that, on occasion, too few nurses were rostered on for certain shifts—particularly at its flagship hospital in Bangalore. “We had based our rostering calculations on the number of beds occupied by patients and, when we investigated, we found at a macro level, we were rostering on the correct number of nurses,” says Muthana. “However, on some days, on wards optimally staffed by, say, 10 nurses per shift, we might have 14 nurses rostered on for one shift and seven rostered on for another shift. Those times we had seven, we had a clear shortage.”

In 2012, Muthana asked a consultant to develop algorithms to help automate the rostering. “Unfortunately, there were so many variables, the consultant failed to solve the problem,” he says. The Chief Operating Officer’s next step was to ask the founder and Chief Executive Officer of predictive analytics business Retigence Technologies—already working with the business on a materials management project—to develop an application to manage rostering.

Removing the daily drudgery

“Our plan with the automation project was to relieve our senior nurses of the daily drudgery of amending the rosters and to deliver rosters that were as fair as possible to all our staff,” says Muthana. “We also saw an opportunity to reduce our costs by reducing the overall number of nurses needed to look after our patients.”

Retigence Technologies’ team members then worked with the group’s nurses to capture the variables and requirements for the project. For example, a minimum number of nurses with one year experience or more needs to be rostered on for each shift. Retigence Technologies then started building an application using compute resources available through Compute Engine, a Google Cloud product. “We selected Google primarily because of its pioneering work in artificial intelligence (AI) and machine learning,” says Srinibas Behera, founder and Chief Executive Officer of Retigence Technologies.

With the nurses rostering application developed, Manipal Health Enterprises Pvt. Ltd undertook several pilots to build user acceptance. “Some of our senior nurses took time to accept the fact automation removed their control over rostering assignments, but were finally convinced by the better transparency the product offered,” says Muthana. The organization deployed the application to a smaller hospital and secured user support before rolling out the application to its Bangalore flagship.

Eliminating stress

Deploying the application has enabled Manipal Health Enterprises Pvt. Ltd to remove a buffer of about 100 nurses retained to accommodate the variations in number of nurses rostered for individual shifts. “The savings on those salaries more than paid for the cost of developing the application,” says Muthana.

The application also enabled the organization to reduce the stress on nurses—both the nurses in charge of the rosters and the nurses subject to the rosters. “Night shifts were more equitably distributed among the nurses, while we have been able to reduce the 45 minutes per day required to amend rosters to just 10 minutes,” says Muthana. “In Bangalore alone, we have 51 nurses in charge of rostering—so the combined saving there equates to nearly 30 hours per day.” This is freeing up these senior nurses to complete more important tasks.

The application was subsequently implemented at Kasturba Hospital, Manipal, again reducing the time needed to generate complete nursing rosters to less than 10 minutes.

Managing leave and training

The organization now plans to extend the application to manage leave and training for its nurses and other employees. “I would like to see every employee given an annual leave plan that is added to the roster at the start of the year,” says Muthana. “The flexibility and control afforded by the application would enable us to address challenges such as managing leave across a workforce with high attrition rates.” The organization would also be able to create a calendar to ensure nurses receive all their required training.

“We also plan to keep fine-tuning the application to deploy nurses more efficiently and continue to reduce their stress levels,” adds Muthana. “We also want to create a nursing load indicator tailored to patients’ specific circumstances. For example, a sedated patient may not require much nursing care, whereas a patient who comes in with a broken leg and may be on a ventilator may require assistance from three nurses at once.” The business plans to use Google’s AI and machine learning APIs in the future to improve the value and user experience of the product.

Blog

Harnessing the Power of Data and AI to Transform Life Science Supply Chains

3973

Of your peers have already read this article.

3:30 Minutes

The most insightful time you'll spend today!

Maximize efficiency and make data-driven decisions in your life science supply chain by harnessing the power of data and AI. Discover how advanced analytics and machine learning can enhance visibility and control.

Global life science supply chains are lengthy and complex with many moving parts. One small disruption can create serious delays and affect your ability to deliver therapeutics for patients.

Supply chain disruptors

Over the last few years, healthcare organizations have encountered a range of obstacles, from both internal and external factors, that have resulted in supply networks failing to get drugs and medical devices to where they need to be on time. These obstacles include:

  • Labor and supply shortages
  • Rising material costs
  • Raw material constraints
  • Geo-political events
  • Unpredictable weather

How do you overcome supply chain disruptors that are out of your control?

The intelligent healthcare supply chain

While many organizations have already implemented data-driven supply chains, organizations are still faced with the challenges of static, siloed, and different functional supply chain applications; limited data exchange with key trading partners across upstream and downstream operations; and the inability to effectively leverage relevant external data.

At Google Cloud, we believe the key to meaningful and effective change is a data-driven supply chain that allows you to achieve visibility, flexibility, and innovation.

Our solutions help you prepare for the unpredictable and enhance the value of your data. By unlocking AI-driven insights, you can strengthen distribution networks and optimize your workflows and supply chains to become more reliable, intelligent, and sustainable. Some of the business challenges we address include:

Make sure you’re prepared for the unpredictable with real-time visibility over your distribution networks. Learn how you can harness the power of AI and analytics and gain actionable insights that enhance your supply chain.

How-to

How to Build A Basic Image Search Utility for Natural Language Queries

4700

Of your peers have already read this article.

6:00 Minutes

The most insightful time you'll spend today!

Read the post to walk through the components required for building a basic image search utility for natural language queries and how learn how the components are connected to each other!

This post shows how to build an image search utility using natural language queries. Our aim is to use different GCP services to demonstrate this. At the core of our project is OpenAI’s CLIP model. It makes use of two encoders – one for images and one for texts. Each encoder is trained to learn representations such that similar images and text embeddings are projected as close as possible.

We will first create a Flask-based REST API capable of handling natural language queries and matching them against relevant images. We will then demonstrate the use of the API through a Flutter-based web and mobile application. Figure 1 shows how our final application would look like:

Figure 1 real
Figure 1: Final application overview.

All the code shown in this post is available as a GitHub repository. Let’s dive in. 

Application at a high-level

Our application will take two queries from the user:

  • Tag or keyword query. This is needed in order to pull a set of images of interest from Pixabay. You can use any other image repositories for this purpose. But we found Pixabay’s API to be easier to work with. We will cache these images to optimize the user experience. Suppose we wanted to find images that are similar to this query: “horses amidst flowers”. For this, we’d first pull in a few “horse” images and then run another utility to find out the images that best match our query.  
  • Longer or semantic query that we will use to retrieve the images from the pool created in the step above. These images should be semantically similar to this query. 

Note: Instead of two queries, we could have only taken a single long query and run named-entity extraction to determine the most likely important keywords to run the initial search with. For this post, we won’t be using this approach. 

Figure 2 below depicts the architecture design of our application and the technical stack used for each of the components.

Figure 2
Figure 2: Architecture design and flow.

Figure 2 also presents the core logic of the API we will develop in bits and pieces in this post. We will deploy this API on a Kubernetes cluster using the Google Kubernetes Engine (GKE). The following presents a brief directory structure of our application code-base:

after figure 2

Next, we will walk through the code and other related components for building our image search API. For various machine learning-related utilities, we will be using PyTorch

Building the backend API with Flask

First, we’d need to fetch a set of images with respect to user-provided tags/keywords before performing the natural language image search. The utility below from the pixabay_utils.py script can do this for us:

  def fetch_images_tag(pixabay_search_keyword, num_images):
    """
    Fetches images from Pixabay w.r.t a keyword.
    :param pixabay_search_keyword: Keyword to perform the search on Pixabay.
    :param num_images: Number of images to retrieve.
    :return: List of PIL images.
    :return: List of image URLs.
    """
    query = (
        PIXABAY_API
        + "&q="
        + pixabay_search_keyword.lower()
        + "&image_type=photo&safesearch=true&per_page="
        + str(num_images)
    )
    
    response = requests.get(query)
    output = response.json()

    all_images = []
    all_image_urls = []

    for each in output["hits"]:
        imageurl = each["webformatURL"]
        response = requests.get(imageurl)
        image = Image.open(BytesIO(response.content)).convert("RGB")
        all_images.append(image)
        all_image_urls.append(imageurl)

    return (all_images, all_image_urls)

Note that all the API utilities are logging relevant information. But for brevity, we have omitted the lines of code responsible for that. Next, we will see how to invoke the CLIP model and select the images that would best match a given query semantically.  For this, we’ll be using Hugging Face, an easy-to-use Python library offering state-of-the-art NLP capabilities. We’ll collate all the logic related to this search inside a SimilarityUtil class:

  class SimilarityUtil:
    def __init__(self):
        self.model = CLIPModel.from_pretrained(CLIP_MODEL)
        self.processor = CLIPProcessor.from_pretrained(CLIP_PREPROCESSOR)
        self.device = "cuda" if torch.cuda.is_available() else "cpu"

    def perform_sim_search(self, images, query_phrase, top_k=3):
        """
        Performs similarity search between the images and query.
        :param images: A list of PIL images initially retrieved with
        respect to some entity e.g. Tiger.
        :param query_phrase: A list containing a single text query,
        e.g. "Tiger drinking water".
        :param top_k: Number of top images to return from `images`.
        :return: Top-k indices matching the query semantically and
        their similarity scores.
        """
        model = self.model.to(self.device)
        # Obtain the text-image similarity scores
        with torch.no_grad():
            inputs = self.processor(
                text=[query_phrase], images=images, return_tensors="pt", padding=True
            )
            inputs = inputs.to(self.device)
            outputs = model(**inputs)

        # Image-text similarity scores
        logits_per_image = outputs.logits_per_image.cpu()
        (top_indices, top_scores) = self.sort_scores(logits_per_image, top_k)

        return (top_indices, top_scores)

    def sort_scores(self, scores, top_k):
        """
        Sorts the scores in a descending manner.
        :param scores: Scores to sort through.
        :param top_k: Number of top scores to return.
        :return: Top-k scores and their indices.
        """
        values, indices = scores.squeeze().topk(top_k)
        top_indices, top_scores = [], []

        for score, index in zip(values, indices):
            top_indices.append(int(index.numpy()))
            score = score.numpy().tolist()
            top_scores.append(round(score, 3))

        return (top_indices, top_scores)

CLIP_MODEL uses a ViT-base model to encode the images for generating meaningful embeddings with respect to the provided query. The text-based query is also encoded using A Transformers-based model for generating the embeddings. These two embeddings are matched with one another during inference. To know more about the particular methods we are using for the CLIP model please refer to this documentation from Hugging Face. 

In the code above, we are first invoking the CLIP model with images and the natural language query. This gives us a vector (logits_per_image) that contains the similarity scores between each of the images and the query. We then sort the vector in a descending manner. Note that we are initializing the CLIP model while instantiating the SimilarityUtil to save us the model loading time. This is the meat of our application and we have tackled it already. If you want to interact with this utility in a live manner you can check out this Colab Notebook

Now, we need to collate our utilities for fetching images from Pixabay and for performing the natural language image search inside a single script – perform_search.py. Following is the main class of that script:

  class Searcher:
    def __init__(self):
        self.similarity_model = SimilarityUtil()

    def get_similar_images(self, keyword, semantic_query, pixabay_max, top_k):
        """
        Finds semantically similar images.
        :param keyword: Keyword to search with on Pixabay.
        :param semantic_query: Query to find semantically similar images retrieved from Pixabay.
        :param pixabay_max: Number of maximum images to retrieve from Pixabay.
        :param top_k: Top-k images to return.
        :return: Tuple of top_k URLs and the similarity scores of the images present inside the URLs.
        """
        images_redis_key = keyword + "_images"
        urls_redis_key = keyword + "_urls"

        if redis_client.exists(images_redis_key) and redis_client.exists(
            urls_redis_key
        ):
            keyword_images = redis_client.get(images_redis_key)
            keyword_image_urls = redis_client.get(urls_redis_key)
        else:
            (keyword_images, keyword_image_urls) = fetch_images_tag(
                keyword, pixabay_max
            )
            redis_client.set(images_redis_key, keyword_images)
            redis_client.set(urls_redis_key, keyword_image_urls)

        (top_indices, top_scores) = self.similarity_model.perform_sim_search(
            keyword_images, semantic_query, top_k
        )

        top_urls = [keyword_image_urls[index] for index in top_indices]

        return (top_urls, top_scores)

Here, we are just calling the utilities we had previously developed to return the URLs of the most similar images and their scores. What is even more important here is the caching capability. For that, we combined GCP’s MemoryStore and a Python library called direct-redis. More on setting up MemoryStore later. 

MemoryStore provides a fully managed and low-cost platform for hosting Redis instances. Redis databases are in memory and light-weight making them an ideal candidate for caching. In the code above, we are caching the images fetched from Pixabay and their URLs. So, in the event of a cache hit, we won’t need to call the CLIP model and this will tremendously improve the response time of our API. 

Other options for caching

We can cache other elements of our application. For example, the natural language query. When searching through the cached entries to determine if it’s a cache hit, we can compare two queries for semantic similarity and return results accordingly. 

Consider that a user had entered the following natural language query: “mountains with dark skies”. After performing the search, we’d cache the embeddings of this query. Now, consider that another user entered another query: “mountains with gloomy ambiance”. We’d compute its embeddings and run a similarity search with the cached embeddings. We’d then compare the similarity scores with respect to a threshold and parse the most similar queries and their corresponding results. In case of a cache miss, we’d just call the image search utilities we developed above. 

When working on real-time applications we often need to consider these different aspects and decide what enhances the user experience and maximizes business at the same time. 

All that’s left now for the backend is our Flask application – main.py:

  @app.route("/search", methods=["GET"])
def get_images():
    tag = request.args.get("t").lower()
    query = request.args.get("s_query").lower()
    top_k = request.args.get("k")

    (top_urls, top_scores) = searcher.get_similar_images(
        tag, query, MAX_PIXABAY_SEARCH, int(top_k)
    )

    return jsonify({"top_urls": top_urls, "top_scores": top_scores})

Here we are first parsing the query parameters from the request payload of our search API.  We are then just calling the appropriate function from perform_search.py to handle the request. This Flask application is also capable of handling CORS. We do this via the flask_cors library:

  cors = CORS(
    app,
    resources={
        r"/search/*": {"origin": "*"},
        r"/test/*": {"origin": "*"},
    },
)

And this is it! Our API is now ready for deployment. 

Deployment with Compute Engine and GKE

The reason why we wanted to deploy our API on Kubernetes is because of the flexibility Kubernetes offers for managing deployments. When operating at scale, auto scalability and load balancing are very important. With the comes the requirement of security — we’d not want to expose the utilities for interacting with any internal services such as databases. With Kubernetes, we can achieve all these easily and efficiently. 

GKE provides secured and fully managed functionalities for operationalizing Kubernetes clusters. Here are the steps to deploy the API on GKE at a glance:

  • We first build a Docker image for our API and then push it to the Google Container Registry (GCR).
  • We then create a Kubernetes cluster on GKE and initialize a deployment.
  • We then add scalability options.
  • If any public exposure is needed for the API, we then tackle it. 

We can assimilate all the above into a shell script – k8s_deploy.sh:

  ## Docker build and push ## 
# We are inside the `server` directory
docker build -t gcr.io/${PROJECT_ID}/search_service .
docker push gcr.io/${PROJECT_ID}/search_service

## Deploy on a GKE cluster ##
# Configure the Docker command-line tool to authenticate to Container Registry
gcloud auth configure-docker

# Create a Kubernetes cluster
gcloud container clusters create image-search-nlp 

# Create a deployment
kubectl create deployment clip-search --image=gcr.io/${PROJECT_ID}/search_service

# Number of worker replicas
kubectl scale deployment clip-search --replicas=3

# HorizontalPodAutoscaler resource
kubectl autoscale deployment clip-search --cpu-percent=80 --min=1 --max=5

# Expose deployment
kubectl  expose deployment clip-search --name=clip-search-service --type=LoadBalancer --port 80 --target-port 8080

These steps are well explained in this tutorial that you might want to refer to for more details. We can configure all the dependencies on our local machine and execute the shell script above. We can also use the GCP Console to execute it since a terminal on the GCP Console is pre-configured with the system-level dependencies we’d need. In reality, the Kubernetes cluster should only be created once and different deployment versions should be created under it. 

After the above shell script is run successfully, we can run kubectl get service to know the external IP address of the service we just deployed:

  NAME                       CLUSTER-IP    EXTERNAL-IP     PORT(S)          AGE
clip-search-service      10.3.251.122    203.0.113.0     80:30877/TCP     10s

We can now consume this API with the following base URI: http://203.0.113.0/. If we wanted to deal with only http-based API requests, then we are done here. But secured communication is often a requirement in order for applications to operate reliably. In the following section, we are to discuss how to configure the additional items to allow our Kubernetes cluster to allow https requests.  

Configurations for handling https requests with GKE

A secure connection is almost often a  must-have requirement in modern client/server applications. The front-end Flutter application would be hosted on GitHub Pages for this project, and it requires https-based connection as well. Even if configuring https connection particularly for a GKE-based cluster can be considered a chore, its setup might seem daunting at first.

There are six steps to configure https connection in the GKE environment: 

  1. You need to have a domain name, and there are a lot of inexpensive options that you can buy. For instance, mlgde.com domain for this project is acquired via Gabia which is a Korean service provider.
  2. A reserved (static) external IP address has to be acquired via gcloud command or GCP console
  3. You need to bind the domain name with the acquired external IP address. This is a platform-specific configuration that issued the domain name to you. 
  4. There is a special ManagedCertificate resource which is specific to the GKE environment. ManagedCertificate resource specifies the domain that the SSL certificate will be created for, so you need this. 
  5. An Ingress resource should be created by listing the static external IP address, ManagedCertificate resource, and the service name and port which the incoming traffic will be routed to. The Service resource could remain the same as in the above section with only changes from LoadBalancer to ClusterIP. 
  6. Last but not least, you need to modify the existing Flask application and Deployment resource to support liveness and readiness probes which are used to check the health status of the Deployment. The Flask application side can be simply modified with the flask-healthz Python package, and you only need to add livenessProbe and readinessProbe sections in the Deployment resource. In the code example below, the livenessProbe and readinessProbe are checked via /alive and /ready endpoints respectively.
  from flask_healthz import healthz
from flask_healthz import HealthError


app = Flask(__name__)
app.register_blueprint(healthz, url_prefix="/")


def printok():
   print("Everything is fine")


def liveness():
   try:
       printok()
   except Exception:
       raise HealthError("Can't connect to the file")


def readiness():
   try:
       printok()
   except Exception:
       raise HealthError("Can't connect to the file")


app.config.update(
   HEALTHZ = {
       "alive": "main.liveness",
       "ready": "main.readiness",
   }
)

One thing to be careful of is the initialDelaySeconds attribute of the probes. It is uncommon to configure this attribute with a big number, but it could be bigger than 90 – 120 seconds depending on the size of the model to be used. For this project, it is configured in 90 seconds in order to wait until the CLIP model is fully loaded into memory (full YAML script here).

  livenessProbe:
 httpGet:
   path: /alive
   port: 8080
 initialDelaySeconds: 90
 periodSeconds: 10

readinessProbe:
 httpGet:
   path: /ready
   port: 8080
 initialDelaySeconds: 90
 periodSeconds: 10

Again, these steps may seem daunting at first, but it will become clear when you have done it once. Here is the official document for Using Google-managed SSL certificates You can find all the GKE-related resources used in this project here

Once every step is completed you should be able to see your server application running on the GKE environment. Please make sure to run kubectl apply command whenever you create Kubernetes resources such as Deployment, Service, Ingress, and ManagedCertificate, and it is important to wait for more than 10 minutes until the ManagedCertifcate provisioning is done. 

You can run gcloud compute addresses list command to find out the static external IP address that you have configured.

  NAME  ADDRESS/RANGE  TYPE      PURPOSE  NETWORK  REGION  SUBNET  STATUS
gde      34.149.231.34       EXTERNAL                                                                    IN_USE

Then, the IP address has to be mapped to the domain. Figure 3 is a screenshot of a dashboard from where we got the mlgde.com domain. It clearly shows mlgde.com is mapped to the static external IP address configured in GCP.

Figure 3
Figure 3: API endpoints mapped to our custom domain.

In case you’re wondering why we didn’t deploy this application on App Engine, well that is because of the compute needed to execute the CLIP model. App Engine instance won’t fit in that regime. We could have also incorporated compute-heavy capabilities via a VPC Connector. That is a design choice that you and your team would need to consider. In our experiments, we found the GKE deployment to be easier and suitable for our needs. 

Infrastructure for the CLIP model

As mentioned earlier, at the core of our application is the CLIP model. It is computationally a bit more expensive than the regular deep learning models. This is why it makes sense to have the hardware infrastructure set up accordingly to execute it. We ran a small benchmark in order to see how a GPU-based environment could be beneficial here. 

We ran the CLIP on a Tesla P100-based machine and also on a standard CPU-only machine 1000 times. The code snippet below is the meat of what we executed:

  DEVICE = "cuda" if torch.cuda.is_available() else "cpu"

start_time = time.time()
for _ in range(1000):
    with torch.no_grad():
        model = model.to(DEVICE)
        inputs = processor(text=[semantic_search_phrase], 
                        images=all_images, return_tensors="pt", padding=True)
        inputs = inputs.to(DEVICE)
        outputs = model(**inputs)
end_time = time.time() - start_time
print(f"Total time: {end_time:.3f} seconds.")

As somewhat expected, with the GPU, the code took 13 minutes to complete execution. With no GPU, it took about 157 minutes.

It is uncommon to leverage GPUs for model prediction because of cost restrictions, but sometimes we have to access GPUs for deploying a big model like CLIP. We configured a GPU-based cluster on GKE and compared the performance differences with and without it. It took about 1 second to handle a request with GPU and MemoryStore cache while it took more than 4 seconds with MemoryStore only (without the GPUs). 

For the purposes of this post, we used a CPU-based cluster on Kubernetes. But It is easy to configure GPU usage in a GKE cluster. This document shows you how to do so. For a short summary, there are two steps. First, a node should be configured with GPUs when creating a GKE cluster. Second, GPU drivers should be installed in GKE nodes. You don’t need to visit and manually install GPU drivers for each node by yourself. Rather you can simply apply the DaemonSet resource to GKE as described here.

Setting up MemoryStore

In this project, we first query the general concept of images to Pixabay, then we filter the images with a semantic query using CLIP. It means we can cache the initially retrieved images from Pixabay for the next specific semantic query. For instance, you may want to search with “gentleman wearing tie” at first, then you may want to retry searching for “gentleman wearing glass”. In this case, the base images remain all the same, so they could be stored in a cache server like Redis. 

MemoryStore is a GCP service wrapping the Redis which is an in-memory data store, so you can simply use a standard Redis Python package for accessing it. The only thing to be careful about when provisioning a MemoryStore Redis instance is to make sure it is in the same region where your GKE cluster or Compute Engine instance is.

Figure 4
Figure 4: MemoryStore setup.

The code snippet below shows how to make a connection to the Redis instance in Python. Nothing specific to GCP, but you only need to be aware of the usage of the standard redis-py package

  # REDISHOST is the IP address to the MemoryStore instance
redis_host = os.environ.get("REDISHOST", "localhost")
redis_port = int(os.environ.get("REDISPORT", 6379))
redis_client = redis.StrictRedis(host=redis_host, port=redis_port)

After creating a connection, you can store and retrieve data from MemoryStore. There are more advanced use cases of Redis, but we only used exists, get, and set methods for the demonstration purpose. These methods should be very familiar if you know maps, dictionaries, or other similar data structures. For the code portion that uses Redis-related utilities, please refer to the Searcher Python class we discussed in an earlier section. 

In the URLs below, you can find side-by-side comparisons of using MemoryStore:

Putting everything together

All that’s left now is to collate the different components we developed in the sections above and deploy our application with a frontend. All the frontend-related code is present here

  • The front-end application is written in the Flutter development kit. The main screen contains two text fields for queries to Pixabay and CLIP model respectively. When you click the “Send Query” button, it will send out a RestAPI request to the server. After receiving the result back from the server, the retrieved images from the semantic query will be displayed at the bottom section of the screen. 
    • Please note that a Flutter application can be deployed to various environments including desktop, web, iOS, and Android. In order to keep as simple as possible, we chose to deploy the application to the GitHub Pages. Whenever there is any change to a client-side source directory, the GitHub Action will be triggered to build a web page and deploy the latest version to the GitHub Pages. 

Our final application is deployed here and it looks like so:

Figure 1
Figure 5: Live application screen.

Note that due to constraints, the above-mentioned URL will only be live for one or two months.  

It is also possible to redeploy the back-end application with a GitHub Action. 

  • The very first step is to craft a Dockerfile like below. Since Python is a scripting language, and there are lots of heavy packages that the application is dependent on, it is important to cache the steps. For instance, installing the dependencies should be separated from other commands.
  FROM pytorch/pytorch:latest
WORKDIR /app

# install the dependencies
COPY requirements.txt requirements.txt
RUN pip install -r requirements.txt

COPY . .  

# set up environment variables
ENV PIXABAY_API_KEY="..."
ENV REDIS_IN_USE="true"
ENV REDISHOST="..."

# expose port that Flask app is listening on
EXPOSE 8080

# run the Flask app
CMD [ "python3", "main.py" ]
  • With the Dockerfile defined, we can use a GitHub Action like this for automatic deployment. 

Edge cases

Since the CLIP model is pre-trained on a large corpus of image and text pairs it’s likely that it may not generalize well to every natural language query we throw at it. Also, because we are limiting the number of images on which the CLIP model can operate, this somehow restricts the expressivity of the model.

We may be able to improve the performance for the second situation by increasing the number of images to be pre-fetched and by indexing them into a low-cost and high-performance database like Datastore

Costs

In this section, we wanted to provide the readers a breakdown of the costs they might incur in order to consume the various services used throughout the application. 

  • Frontend hosting
    • The front-end application is hosted on GitHub Pages, so there is no expenditure for this.
  • Compute Engine
    • With an e2-standard-2 instance type without GPUs, the cost is around $48.92 per month. In case you want to add a GPU (NVIDIA K80), the cost goes up to $229.95 per month.
  • MemoryStore
    • The cost for MemoryStore depends on the size. With 1GB of space, the cost is around $35.77 per month, and whenever you add more GBs the cost will be doubled.
  • Google Kubernetes Engine
    • The monthly cost for a 3 node GKE cluster with n2-standard-2 (vCPUs: 2, RAM: 8GB without GPUs) is about $170.19. If you add one GPU (NVIDIA K80) to the cluster, the cost goes up to $835.48.

While you may think that is a lot cost-wise, it is good to know that Google gives away free $300 credits when you create a new GCP account. It is still not enough for leveraging GPUs, but it is enough to learn and experiment with GKE and MemoryStore usage.

Conclusion

In this post, we walked through the components needed to build a basic image search utility for natural language queries. We discussed how these different components are connected to each other. Our image search API is able to utilize caching and was deployed on a Kubernetes cluster using GKE. These elements are essential when building a similar service to cater to a much bigger workload. We hope this post will serve as a good starting point for that purpose. Below are some references on similar areas of work that you can explore:

Acknowledgments: We are grateful to the Google Developers Experts program for supporting us with GCP credits. Thanks to Karl Weinmeister and Soonson Kwon of Google for reviewing the initial draft of this post. 

How-to

An AI-Powered Cost Cutting Guide: 8 Strategies for Maximizing Profits

1714

Of your peers have already read this article.

3:30 Minutes

The most insightful time you'll spend today!

Want to stay ahead of the curve and keep your business thriving? It's time to start harnessing the power of data and AI. In this post, we'll share 8 actionable tips for cutting costs and driving profits, all while staying ahead of the competition.

We are increasingly seeing one question arise in virtually every customer conversation: How can the organization save costs and drive new revenue streams? 

Everyone would love a crystal ball, but what you may not realize is that you already have one. It’s in your data. By leveraging Data Cloud and AI solutions, you can put your data to work to achieve your financial objectives. Combining your data and AI reveals opportunities for your business to reduce expenses and increase profitability, which is especially valuable in an uncertain economy. 

Google Cloud customers globally are succeeding in this effort, across industries and geographies. They are improving ROI by saving money and creating new revenue streams. We have distilled the strategies and actions they are implementing—along with customer examples and tips—in our eBook, “Make Data Work for You.” In it, you’ll find ways you can pare costs, increase profitability, and monetize your data.  

Find money in your data 

Our Google Cloud teams have identified eight strategies that successful organizations are pursuing to trim expenses and uncover new sources of revenue through intelligent use of data and AI. These use cases range from scaling small efficiencies in logistics to accelerating document-based workflows, monetizing data, and optimizing marketing spend.

https://storage.googleapis.com/gweb-cloudblog-publish/images/1_Cost_Optimization.max-900x900.jpg

The results are impressive. They include massive cost savings and additional revenue. On-time deliveries have increased sharply at one company, and procure-to-pay processing costs have fallen by more than half at another. Other organizations have reaped big gains in ecommerce upselling and customer satisfaction.

We’ve found that businesses across every industry and around the globe are able to take action on at least one of these eight strategies. Contrary to common misperceptions, implementation does not require massive technology changes, crippling disruption to your business, or burdensome new investments. 

What success looks like 

If you worry your business is not ready or you need to gain buy-in from leadership, the success stories of the 15 companies in this report are helpful examples. Learning how organizations big and small, in different industries and parts of the world, have implemented these data and AI strategies makes the opportunities more tangible.

Carrefour 
Among the world’s largest retailers, Carrefour operates supermarkets, ecommerce, and other store formats in more than 30 countries. To retain leadership in its markets, the company wanted to strengthen its omnichannel experience.

Carrefour moved to Google Data Cloud and developed a platform that gives its data scientists secure, structured access to a massive volume of data in minutes. This paved the way for smarter models of customer behavior and enabled a personalized recommendation engine for ecommerce services. 

The company saw a 60% increase in ecommerce revenue during the pandemic, which it partly attributes to this personalization. 

ATB Financial 
ATB Financial, a bank in the Canadian province of Alberta, uses its data and AI to provide real-time personalized customer service, generating more than 20,000 AI-assisted conversations monthly. Machine learning models enable agents to offer clients real-time tailored advice and product suggestions. 

Moreover, marketing campaigns and month-end processes that used to take five to eight hours now run in seconds, saving over CA$2.24 million a year. 

Bank BRI
Bank BRI, which is owned by the Indonesian government, has 75.5 million clients. Through its use of digital technologies, the institution amasses a lot of valuable data about this large customer base. 

Using Google Cloud, the bank packages this data through more than 50 monetized open APIs for more than 70 ecosystem partners who use it for credit scoring, risk management, and other applications. Fintechs, insurance companies, and financial institutions don’t have the talent or the financial resources to do quality credit scoring and fraud detection on their own, so they are turning to Bank BRI. 

Early in the effort, the project generated an additional $50 million in revenue, showing how data can drive new sources of income. 

How to get going now

Make Data Work for You” will help you launch your financial resiliency initiatives by outlining the steps to get going. The process lays the groundwork for realizing your own cost savings and new revenue streams by leveraging data and AI.

Among these steps include building frameworks to operate cost efficiently, make informed decisions related to spending and optimize your data and AI budgets.

https://storage.googleapis.com/gweb-cloudblog-publish/images/2_Cost_Optimization.max-900x900.jpg

Operate: Billing that’s specific to your use-case
Control your costs by choosing data and analytics vendors who offer industry-leading data storage solutions and flexible pricing options. For example, multiple pricing options such as flat rate and pay-as-you-go allow you to optimize your spend for best price-performance.

Inform: make informed decisions based on usage
Use your cloud vendor’s dashboards or build a billing data report to gain insights on your spending over time. Make use of cost recommendations and other forecasting tools to predict what your future expenses are going to be.

Optimize: Never pay more than you use 
While planning data analytics capacity, organizations often overprovision and overpay than what they actually use. Consider migrating your workloads that have unpredictable demand to a data warehousing solution that offers granular level autoscaling features so that you never have to pay for more than what you use.

There are other key moves that will set your initiative up for success including how to shorten time to value in building AI models and measuring impact. You can find details in the report.

A brighter future

The teams at Google Cloud helped the companies in “Make Data Work for You,” along with many more organizations, use their data and AI to achieve meaningful results. Download the full report to see how you can too.

Blog

DocAI Lowers Customer’s Document Processing Cost by 60 Percent. Learn How

5621

Of your peers have already read this article.

2:00 Minutes

The most insightful time you'll spend today!

Document processing in most organizations require human intervention for review, data extraction and unlocking valuable insights. With Google Cloud Doc AI platform, document-intensive industries and businesses can end their guess work through automation, document processing accuracy, ML-based predictions and better UX at 60 percent lesser document processing cost.

Some of the most important data at your company isn’t living in databases, but in documents, and most business processes begin, involve or end with a document. 

Yet most companies are still manually entering data and reliant on guesswork to make sense of it all as the volume and variety of data explodes. Organizations are also leaving heaps of value on the table in the form of new and better customer experiences that can be unlocked with artificial intelligence (AI) applied to documents. 

The latest releases of Document (Doc) AI platformLending DocAI and Procurement DocAI, built on decades of AI innovation at Google, bring powerful and useful solutions to these challenges. Under the hood are Google’s industry-leading technologies:

  • Computer vision (including OCR) and Natural Language Processing (NLP) that creates pre-trained models for high-value, high-volume documents. 
  • Google Knowledge Graph to validate and enhance the fields in your documents.
  • Training and creation of your own custom document models. 
  • Human interaction with AI to ensure accuracy where needed.

Google Cloud DocAI platform, Lending DocAI and Procurement DocAI are now generally available. Thousands of customers have tried these products in the preview phase—and DocAI has already processed tens of billions of pages of documents across lending, insurance, government and other industries.

“The vast majority of enterprise content still resides in unstructured sources like documents. Google Cloud’s Document AI brings a fresh new perspective to the problem informed by the company’s decades of experience making sense of the largest unstructured corpus in the world—the world wide web.” Ritu Jyoti, VP of AI Research, IDC

Cut document processing costs by up to 60%

Lending DocAI helps banks, mortgage brokers and other lending institutions fast track the loan application process from weeks to days, dramatically reducing the cost of issuing a loan. And Procurement DocAI enables companies to automate procurement data capture at scale, lowering processing costs by up to 60%.

These solutions are built on DocAI platform, a unified console for document processing that lets you quickly access all parsers and tools. From the platform, you can automate and validate documents to streamline workflows, reduce guesswork, and keep data accurate and compliant. 

Get more value from AI with DocAI’s industry-specific solutions

According to Accenture’s AI: Built to Scale report: “Companies that scale successfully see 3x the return on their AI investments compared to those who have not fully rolled out AI capabilities.” 

Core to our strategy at Google Cloud is the creation of industry-specific solutions that help companies get maximum value out of their investments in AI. We announced Lending DocAI, our first solution designed specifically for the financial services industry, at the Mortgage Bankers Association convention last year. It processes borrowers’ income and asset documents using a set of specialized machine learning (ML) models, and automates routine document reviews so that mortgage providers can focus on more important work. 

Lending DocAI is now generally available and includes more specialized parsers for critical loan documents including paystubs, bank statements, and more. Our goal is to provide the right tools to help borrowers and lenders have a better experience and close home loans faster. For more, watch this video.

Procurement DocAI is also now generally available. This solution helps companies accelerate document processing for invoices, receipts, and other valuable documents in the procurement cycle. 

Automating data capture is helping our customers increase accuracy and also lower their procure-to-pay processing costs. We are continually expanding the types of documents Procurement DocAI can process—the latest is a utility parser for electric, water and other bills. In addition, Procurement DocAI leverages Google Knowledge Graph to validate and enrich parsed information to make the data even more useful. Check out this overview video for more details. 

One company that lives and breathes AI-enabled document management is AODocs. It uses Procurement DocAI to simplify invoice processing for enterprise customers and launched a new Gmail add-on, Invoice to Sheet, for SMB customers who just want to track their invoices in Google Sheets.

“Google Cloud’s Procurement DocAI service allows our document management platform to better automate the processing of invoices; AODocs customers who have tested our new account payables workflow estimate that the productivity of their A/P team has more than doubled, thanks to the reduction of manual data input brought by the Procurement DocAI.”Stéphan Donzé, Founder and CEO, AODocs

The new specialized parsers for Lending and Procurement DocAI can be used alongside our existing AutoML Text & Document Classification and AutoML Document Extraction services. These technologies provide a state-of-the-art toolset for creating new document models and have been widely deployed by customers in financial services and other industries. 

Partner to accelerate your AI deployment and results

Having the right partner to ease the complexity of rolling out your AI-strategy in mortgage document processing is critical to transforming your customers’ experience. We’re excited to announce a partnership with Mr. Cooper, a leader in mortgage servicing, to provide customers with more automation and workflow tools throughout their entire mortgage life cycle. As part of this agreement, both companies will collaborate on digitizing Mr. Cooper’s core mortgage platform, creating a more personal customer experience utilizing AI, and driving a broader culture of innovation to imagine and develop services and solutions that will transform the mortgage experience for American homeowners.

“Over the last few years, we have made substantial investments in our servicing technology and core mortgage platform that have revolutionized the customer experience, while providing dramatic efficiencies in operating cost. Our partnership with Google Cloud AI will build on those advances and help make these technologies available for the mortgage industry.” —Jay Bray, Chairman and CEO, Mr. Cooper Group

This builds upon the robust partner ecosystem we’re creating to help customers revolutionize the home loan experience, which includes last year’s partnership announcement with Roostify.

Integrate human review into ML predictions 

Next up is the general availability of Human-in-the-Loop AI, a new DocAI feature that will help companies achieve higher document processing accuracy with the assurance of human review. Adding human review can increase accuracy and help businesses interpret predictions using purpose-built tools to enable those reviews. 

Processing documents quickly and cost-effectively is important. But it’s often necessary to have a high level of assurance on data accuracy for compliance. CIOs and IT decision-makers need highly accurate ML predictions to fulfill compliance requirements, improve employee experience (e.g. less rework), and raise customer satisfaction (e.g. fewer data errors). Including human participation in ML processes allows AI and humans to work together for the best possible results.

gcp human-in-the-loop ai.gif

Human-in-the-Loop AI provides the workflow to manage human review tasks and produces a percentage confidence score of how “sure” it is that the AI ingested the document correctly. Document AI extracts data from documents with ML, and when paired with Human-in-the-Loop AI, human reviewers are able to verify the data captured. This system is customizable, providing the flexibility to set different thresholds and assign individual groups of reviewers to various stages of the workflow. With Human-in-the-Loop AI, developers can choose trusted reviewers to assign to the task; these reviewers can be from within their own or partner organizations.

More Document AI resources

To learn more, check out the Document AI webpage and watch a demo of how to process sample forms in AI Platform notebooks to inspect data extraction and confidence scores. For more on how customers and partners like Workday, AODocs, and Mr. Cooper are using Document AI, listen to our fireside chat. And stay tuned for the exciting evolution of these technologies in future releases of DocAI.

Blog

Google is a Leader in the 2023 Gartner® Magic Quadrant™ for Enterprise Conversational AI Platforms

1293

Of your peers have already read this article.

1:30 Minutes

The most insightful time you'll spend today!

We’re excited to share that Gartner has recognized Google as a Leader in the 2023 Gartner® Magic Quadrant™ for Enterprise Conversational AI Platforms, authored by Bern Elliot and Gabriele Rigon.

We believe this recognition is a testament to Google Cloud’s robust investments and commitment to innovation in AI, coupled with a deep understanding of enterprise customer needs. Enterprises are increasingly investing in AI-driven solutions that balance addressing customer expectations with operational efficiency. At a time when the demand for quality, performant, and trustworthy conversational AI has never been higher, we’re thrilled to continue to deliver best-in-class technologies, purpose-built to solve our customers’ most critical use cases.  

In 2022, Google Cloud delivered cutting-edge conversational AI technologies with many launches to our Conversational AI API portfolio. Together, these enabled developers to leverage Google’s technologies to power their applications with our end-to-end Contact Center AI (CCAI) suite, designed to solve the needs of Customer Experience (CX) and contact center leaders. 

Google Cloud’s Conversational AI APIs include pre-trained models for Speech to TextText to Speech and Natural Language Understanding. This conversational AI core leverages Google Research’s technology for speech, understanding and interaction, enabling and orchestrating high-quality conversational experiences at scale.

With Contact Center AI, organizations see improved customer satisfaction, higher agent productivity and reduced costs through increased agent efficiency. By focusing on user needs, Google Cloud provides comprehensive and integrated solutions that are ready for the enterprise. CCAI encompasses a comprehensive set of offerings to address the needs of the contact center.

In 2022, we launched Contact Center AI Platform, our AI-first, mobile-first, user-first contact center as a service (CCaaS), providing AI-powered experiences, CRM-centered design and deployment flexibility in a single platform without the need for multiple providers. Contact Center AI Platform auto-scales on the backend, with capacity for up to 100k concurrent users on a single tenant. It also offers multi-provider voice resiliency with global low-latency routing for best possible call quality and is optimized for enhanced customer/business data security, reduced downtime, and increased agent productivity. Segra, one of the largest independent fiber infrastructure bandwidth companies in the Eastern U.S., is leveraging CCAI Platform to reimagine their customer experience through predictive flows for common experiences and greatly expanding their channels for customer interaction.

CCAI also includes Dialogflow for building virtual agents, enabling businesses to meet their customers across multiple channels. It offers robust, flexible self-service voice and chat interactions that are just as natural as a live agent. Dialogflow enables both a great customer experience and a cost-effective way to scale services. 

Our Agent Assist service gives businesses the ability to transition a call from a virtual agent to a human agent while maintaining context. It efficiently guides the agent to an accurate response, while providing real-time suggestions, more accurate responses and informed recommendations.

To improve contact center operations, CCAI Insights analyzes all customer conversations to provide leaders with real-time, actionable data points on customer queries, agent performance, and sentiment trends. Its topic modeling capabilities enable deeper understanding of key investment areas and greater classification accuracy.

To ensure our enterprise customers deploying CCAI realize value faster, Google Cloud offers CCAI through three defined transformation stages with out-of-the-box packages. The first stage starts with efficiency basics in the first week that include transcription and summarization. The second stage covers automation basics within six months using Agent Assist and Insights. The final stage is full automation within a year with industry use cases and pre-built components. All of this leads to higher agent efficiency, improved customer satisfaction and increased containment.

As we look forward to the rest of 2023 and beyond, elevating the customer experience through user-first design, AI-first capabilities and accelerating time-to-value will be our north star. We plan to announce exciting new capabilities over the next few months to enable that vision to become a reality for many more organizations. 

We are honored to be a Leader in the 2023 Gartner® Magic Quadrant™ for Enterprise Conversational AI Platforms, and look forward to continuing to innovate and partner with customers on their digital transformation journeys. 

Download the complimentary copy of the report: 2023 Gartner® Magic Quadrant™ for Enterprise Conversational AI Platforms.

Learn more about how organizations are transforming their business with Google Cloud solutions with Contact Center AI.


GARTNER is a registered trademark and service mark of Gartner and Magic Quadrant is a registered trademark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and are used herein with permission. All rights reserved.

This graphic was published by Gartner, Inc. as part of a larger research document and should be evaluated in the context of the entire document. The Gartner document is available upon request from Google.

Gartner does not endorse any vendor, product or service depicted in its research publications and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner research publications consist of the opinions of Gartner’s research organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this research, including any warranties of merchantability or fitness for a particular purpose.

More Relevant Stories for Your Company

Blog

Elevating Voice Solutions: Google Cloud Introduces Speech-to-Text v2 API and Chirp for Businesses

As one of the most innate and ubiquitous forms of expression, speech is a fundamental pillar of human interaction. It comes as no surprise, then, that Google Cloud’s Speech API  has become a crucial tool for enterprise customers, launched to general availability (GA) over six years ago and, now, processing over 1

How-to

Experts’ Guideline for Personalizing Platforms with the Right Recommendation System on Google Cloud

Over the past two decades, consumers have become accustomed to receiving personalized recommendations in all facets of their online life. Whether that be recommended products while shopping on Amazon, a curated list of apps in the Google Play store, or relevant videos to watch next on YouTube. In fact, in

Blog

Data-first Digitization Helps Leverage the Cloud for Your Mainframe Assets

For many enterprises, the venerable mainframe is home to decades’ worth of data about the company’s customers, processes and operations. And it goes without saying that the business would like access to that mainframe data — to report on it, to analyze it with big data analysis tools, or to

Blog

Volkswagen + Google Cloud: Using Machine Learning to Drive Smarter with Energy Efficient Cars

Volkswagen strives to design beautiful, performant, and energy efficient vehicles. This entails an iterative process where designers go through many design drafts, evaluating each, integrating the feedback, and refining. For example, a vehicle’s drag coefficient—its resistance to air—is one of the most important factors of energy efficiency. Thus, getting estimates

SHOW MORE STORIES