What’s New and What’s Next in Smart Analytics - Build What's Next

3080

Of your peers have already watched this video.

15:30 Minutes

The most insightful time you'll spend today!

Explainer

What’s New and What’s Next in Smart Analytics

Data platform architectures that were designed 20 years ago struggle to solve the business problems of 2020 and beyond.

Learn how the latest innovations in Google Cloud’s smart analytics platform can help strip out layers of complexity and analyze data seamlessly in a multi-cloud world.

In this keynote, Debanjan Saha, GM & VP, Engineering, Google Cloud and Vittorio Cretella, Chief Information Officer, The Procter & Gamble Company, share a compelling vision of what’s next in our smart analytics platform and demonstrate our newest solution innovations.

Plus, Luminary customers share how they are using Google Cloud to accelerate their data-driven digital transformations.

Case Study

Zeotap Uses Big Data to Enable Personalized, Precision Marketing at Scale

3363

Of your peers have already read this article.

5:30 Minutes

The most insightful time you'll spend today!

Zeotap migrated from its previous provider to Google Cloud because it was attracted to the benefits of BigQuery as a serverless cloud data warehouse. BigQuery allowed it to “operate more efficiently, removing 80% of the operation load”.

Customers today expect brands to know what they want. In fact, Forbes cited that 71% of customers feel frustrated when the shopping experience is impersonal. From interests to purchase habits, companies need to target customers effectively in order to stand out from the plethora of competitors in the online marketplace.

“Our mission is to help brands engage customers with the most appropriate marketing messages at the right place and at the right time while respecting consumer privacy and ensuring regulatory compliance. The experience of Google Cloud in working with large-scale datasets makes the platform an excellent fit for us.”

Projjol Banerjea, co-founder and Chief Product Officer, Zeotap

Leveraging first-, second-, and third-party data, Zeotap’s Customer Intelligence Platform (CIP)—a CDP with additional identity resolution and third-party data enrichment capabilities—allows brands to engage with their known and unknown users with tailored marketing messages and next best offers or actions. It also offers out-of-box algorithms for standard marketing use cases such as promoting a specific product that will resonate with audiences and custom algorithms for very niche and specific use cases.

“Our mission is to help brands engage customers with the most appropriate marketing messages at the right place and at the right time while respecting consumer privacy and ensuring regulatory compliance. The experience of Google Cloud in working with large-scale datasets makes the platform an excellent fit for us,” says Projjol Banerjea, co-founder and Chief Product Officer at Zeotap.

“With BigQuery, we are able to operate more efficiently, removing 80% of the operation load that we would otherwise have to manage on a different cloud provider. As a result, we get to build and deliver products faster.”

Sathish K S, VP Engineering, Zeotap

Improving data processing to deliver products faster

The cloud-native company migrated from its previous provider to Google Cloud in October 2019 because it was attracted to the benefits of BigQuery as a serverless cloud data warehouse. This was a feature that wasn’t available with its previous provider.

Sathish K S, Engineering Vice-President at Zeotap says, “With BigQuery, we are able to operate more efficiently, removing 80% of the operation load that we would otherwise have to manage on a different cloud provider. As a result, we get to build and deliver products faster.” Since coming onboard BigQuery, Zeotap has been able to increase the speed of its product delivery by 40%.

BigQuery also offers a much higher query limit, which is helpful because Zeotap can converge all of its previous operation load across different platforms into one data warehouse without worrying about it being overloaded.

One of the key services Zeotap offers is customer relationship management (CRM) platform integration. Using Dataflow, Zeotap can easily stream first-party data into BigQuery at scale.

The data integration capabilities of Cloud Data Fusion also complements Dataflow to help Zeotap enhance its API offerings. To illustrate, an ecommerce company may use Zeotap’s internal API to understand their customer’s cart and check-out habits and add Zeotap’s extra layer of data intelligence to it, so they can improve their marketing efforts. Cloud Data Fusion enables a seamless flow of real-time data so that the company can deliver the relevant ads or suggest products just before they make a purchase, which increases the likelihood of them buying more.

Zeotap uses Dataproc and Pub/Sub to manage the large flow of data that goes in and out of its database. Because many of its customers are pushing data in real time, Pub/Sub acts as a queuing engine that feeds them into Dataflow. Likewise, the data from Dataflow is also being channeled back smoothly to the customer endpoints using Pub/Sub.

Being able to automate all of its processes benefitted the DevOps team, as they now have time to focus on developing and improving Zeotap’s product offerings instead of spending time and energy on operations.

Zeotap founders

Managing a full migration seamlessly

Today, all of Zeotap’s data is hosted on Google Cloud. The company began its lift and shift migration in October 2019 and was complete by April 2020. “We have close to 405 data pipelines, and even after the migration, we had to verify the right datasets. We had terabytes of data to move, and despite this, our operations were not affected,” says Aditya Chandra, Vice-President of Infrastructure and Security at Zeotap.

Apart from the BigQuery capabilities, there were other major things Zeotap looked out for in its decision to move to a new cloud provider. The functional parity of Google Cloud was important so that it did not have to rebuild many of its existing structures. Zeotap also didn’t want any drop in performance during and after the move to a new cloud provider. Finally, as a multi-region company, it used to have to spot instance problems and run them on demand, which had cost implications. With the move, it no longer had to face such issues.

“Whatever Spark code that was already running on our existing system had to be lifted and shifted into Dataproc without any library changes because 60% of our codebase is out of there. We’re glad that this was possible,” says Ameya Agnihotri, Chief Technology Officer.

Managing business continuity with a trusted partner

Zeotap now works with Rackspace Technology (Rackspace) as it continues to optimize its infrastructure operations. As a marketing and technology company, Zeotap needs niche expertise to manage big data processing issues.

Ameya Agnihotri shares, “Given its global expertise in dealing with on-premises as well as cloud solutions, Rackspace can really add value because they are better equipped to troubleshoot and handle complex data operations.” Zeotap has plans to explore more of Google Cloud’s artificial intelligence (AI) and machine learning (ML) capabilities, so having a partner in this journey will be valuable to the team.

With so much data on hand, security is paramount to Zeotap. It does a lot of internal security assessments daily. This is one of the areas where the team works with Rackspace to ensure that it adopts best practices and stays on top of the latest updates. With its ability to handle big data operations, Rackspace also acts as an extension of the Zeotap team that supports the infrastructure side of the business.

“The level of engagement from Google Cloud has been excellent, particularly the ability to bring in expertise from different parts of the organization. The personal touch meant we could reach product owners and managers of specific products. This kind of support is what we really value and appreciate.”

Ameya Agnihotri , Chief Technology Officer, Zeotap

Future plans and collaboration with Google Cloud

Zeotap is currently exploring Cloud Bigtable and API Gateway as it continues to grow the business and explore new types of services to offer to its customers. Although it has a software-as-a-service (SaaS) offering for customers to self-serve, some may prefer to do a server-to-server integration so that they do not have to worry about the back end at all. For this reason, Zeotap is exploring the possibility of integrating these customers’ APIs in the back end of its API, using Google Cloud’s API Gateway features that are available out of the box. The company is also looking forward to the new features being released on BigQuery, as well as exploring more of the alerting and monitoring capabilities of Google Cloud.

“The level of engagement from Google Cloud has been excellent, particularly the ability to bring in expertise from different parts of the organization,” says Ameya. He shares that the team has had many fortnightly calls when there were issues to address, and the Google Cloud team would help to fix the problems. “The personal touch meant we could reach product owners and managers of specific products. This kind of support is what we really value and appreciate.”

With big plans ahead for Zeotap that include becoming a Google Cloud Technology Partner in the near future, Swapnasarit Sahu, Chief Data and Analytics Officer, is confident that the company has chosen the right cloud solution. “We see Google as one of the frontrunners in the space of AI and ML solutions, which can really help data scientists in developing new and more innovative products and solutions, and we look forward to building great products with them,” he says.

Zeotap team

4546

Of your peers have already watched this video.

1:30 Minutes

The most insightful time you'll spend today!

Explainer

AI-powered Cameras Help You Serve Your Pets Better. Learn How

Watch the video to learn how you can leverage Google Cloud platform and AI functions to build pet-detection cameras that can capture and send image-based alert messages on your phone to notify when your pets arrive or leave.

Case Study

Quantum Metric Increases Business 10-fold

9278

Of your peers have already read this article.

5:30 Minutes

The most insightful time you'll spend today!

Would your company like access to a system which you can ask: "Show me high-loyalty customers, located in specific geographic areas, who visited the web site at least five times, based on specific campaigns, and never booked a seat on a flight.” Read this.

At Quantum Metric, we’re in the business of bringing our customers business insights that are based on customer experience data and analytics for mid-market and Fortune 500 companies.

Our software, powered by big data, machine intelligence, and Google Cloud, helps our customers identify, quantify, prioritize and measure opportunities to improve digital experiences.

As companies move to a more agile product lifecycle, including continuous deployment and continuous integration, they’re finding that it’s critical to receive perpetual quantified feedback and insights from their data in real time to understand where the largest opportunities exist. 

Each year, billions of customer interactions are captured through browsers or mobile apps on PCs, tablets, and mobile devices. This data, fed into the Quantum Metric platform, can show if a customer had a password problem they couldn’t solve or struggled when trying to purchase something and abandoned their cart.

It also can show if the customer tried in vain to complete an online change to their service provider’s subscription, to reach tech support, or couldn’t find the size or color they were looking for while shopping online.

Most importantly, the Quantum Metric platform quantifies the business value of the issue, helping organizations prioritize where they can make the largest impact to their business.

Success overwhelms our initial architecture
Initially, the Quantum Metric experience analytics software ran on a MySQL open source relational database management system (RDBMS). The MySQL RDBMS worked great for simple queries, when there was a specific question to ask of the data.

Soon, though, we knew we needed to offer more advanced data science capabilities. Our bigger customers wanted to ask questions across very large data sets—days, weeks, months, and years worth of data. They wanted to pose iterative questions using complex filters to answer their most challenging business questions. 

With more complex queries across more data, response times from our RDBMS went from 100 to 500 milliseconds to as long as 20 minutes.

That delay was slowing down our ability and time to insights, which also reduced the value we could provide to our customers, since iterative exploration and analysis requires real-time query responses. Because of the need for real-time responses, there were certain questions that we just weren’t able to ask of the data. It became clear that we needed a much more robust data warehouse solution. 

There were also operational challenges with MySQL and massive-scale data ingestion. We spent a lot of time into the wee hours of the night and morning handling errors and recovering databases.

We tried to address these challenges by sharding, partitioning, and indexing the data to optimize for the types of questions customers were asking. But the problems were escalating and happening more often, from once a month across the customer base to monthly for at least 20 different customers.

We could tune the platform for today and tomorrow’s workload, with good guesses at where indexes could be used, but we simply couldn’t continue to horizontally scale MySQL in a cost-efficient and operationally efficient manner. 

Speed breeds innovation
Once we started exploring options that could better scale with our business, we looked at NoSQL technologies like Cassandra (a partitioned row-store database), MySQL’s Column Store (a columnar store database), and Vertica (a columnar store database)—each with unique ways of handling data storage and accessibility.

But with high volumes of complex queries across large data stores, all of these solutions began to fail, bogged down with multiple, simultaneous users. We could have solved the problems with more raw compute and storage, but it would have been prohibitively expensive to run and require a large team to operate. 

We then decided to try BigQuery, and it was transformative.

We connected our front end to BigQuery via APIs. Once data is 15 minutes old, it is automatically extracted, loaded, and transformed (ETL) to BigQuery.

We continuously update the legacy MySQL RDBMS so its data is integrated with BigQuery data when queries require real-time data. Most query response times are within 100-200 milliseconds, matching what we initially experienced with MySQL.

When traffic from our customers scales up, we can now scale on-demand to accommodate it, thanks to BigQuery’s hundreds of thousands of CPUs. Our customers no longer run into slow response times, and we’ve gained confidence that we can offer them—and their users—advanced insights and better experiences without delay.

More importantly, with this scale of query power, we were able to build data science algorithms into the platform, which iteratively query BigQuery based on the results, and help quantify the impact of a specific issue to a specific segment of users. Adding these capabilities was possible because of the massive scale of BigQuery. 

In addition to new insights and fast response times, we wanted our customers to be able to ask complex questions using very simple language.

For example: “Show me high-loyalty customers, located in specific geographic areas, who visited the web site at least five times, based on specific campaigns, and never booked a seat on a flight.”

This was exactly the kind of query that was used by a major U.S. airline to understand the multi-million dollar impact of a failure affecting their most valuable customers: their high-loyalty members.

And this was all done while maintaining the highest standard of care of customer data and privacy by default, using multiple layers of encryption of data in transit, at rest, and a unique military-grade encryption approach. This approach encrypts PII, including even session cookies, with a RSA-2048 key available only to a select few and used for use cases such as fraud analysis.   

It’s no exaggeration to say that BigQuery has totally transformed our business. It provides the petabyte scale and speed we were missing, in addition to taking care of operational maintenance, a task that was burying our team with MySQL.

We’re now able to support some of the largest companies in the world that require real-time, petabyte-scale analytics. That lets them serve more customers faster with higher quality, and take advantage of BigQuery’s power and scale to innovate.

There are other cloud solutions that can address petabyte analytics, but the most unique value proposition of BigQuery was its on-demand scaling and operational management, with extremely cost-effective pay-as-you-go billing. While today we are at a scale where we have round-the-clock querying needs, our early days had very sporadic query loads where we needed instant scale, then a long lull of nothing. The unique business model of BigQuery’s pay-by-bytes-scanned allowed us to have access to a massive-scale querying platform without breaking the bank. 

Using BigQuery powers better customer experience and reduces purchasing friction
Among the many features of Quantum Metric is the ability to replay online customer sessions. In the example below from a mobile e-commerce site, each action is displayed chronologically. Why did this customer’s transaction fail?

Diving deeper, the session replay shows that the user tried to change the item quantity in the checkout cart, which resulted in a failed API call. Powered by BigQuery, Quantum Metric can then show how many other end users had this issue, with a simple click of “Show More Errors Like This.”

With BigQuery’s massive scale, Quantum Metric will then quantify the impact of that issue, so companies can prioritize which issues need attention immediately. If this is the issue that’s impacting the business the most, our customer can use a single click to open a Jira ticket, forwarding the discovery to their product and engineering teams. Those teams can then re-engineer the experience in near-real time, addressing the failed API call and cutting out the frustrating time it takes for engineers to reproduce the issue.

Quantum Metric platform.png

Once we had a powerful back end in BigQuery, we realized that Quantum Metric’s platform could be used to ask complex questions from vast datasets. We built some of the processes that data scientists use to formulate those queries right into our product.  

For example, we added one series of processes to our platform to help customers understand whether a suspect issue is really impacting end-user experience. Is it something that should be prioritized and fixed? Does this really affect the user experience? Does it have financial impact? These and other questions can be pre-defined as a complex query in Quantum Metric to let our customer quickly gain insight on how an issue is impacting the business. Customers were blown away when they heard this was possible. It was the holy grail for what they were looking for in data science. It really sets us apart from our competition.

Today, with every company heavily dependent on data, those companies that can uncover and act on insights fastest are the ones that will succeed. BigQuery gives us the data warehouse platform we need to provide our customers with fast, reliable technology tools. It frees us from having to deal with the minutiae of technology infrastructure operations, so we can focus on finding and extracting the magic in customer data. With the power and scale of BigQuery, combined with the real-time capture of every user experience with 100% fidelity, we’re able to offer a self-service analytics platform that provides insights into digital journey friction points and acts as the indisputable arbitrator of truth. 

Whitepaper

Google Cloud named a leader in the Forrester Wave: Streaming Analytics

DOWNLOAD WHITEPAPER

4195

Of your peers have already downloaded this article

15:30 Minutes

The most insightful time you'll spend today!

There’s a lot of pressure on companies today to become data-driven, but they need high-performing technology to make that a reality.

As part of Google’s larger cloud data analytics platform, Cloud Pub/Sub and Cloud Dataflow are designed for ease of use, scalability, and performance.

Google Cloud unifies streaming analytics and batch processing the way it should be. No compromises.

Forrester

Forrester has named Google Cloud as a leader in The Forrester Wave™: Streaming Analytics, Q3 2019. This findings reflect Google Cloud’s market momentum, and what Google hears from enterprise customers who are using Cloud Pub/Sub and Cloud Dataflow in production for streaming analytics in conjunction with our broader platform.  

According to Forrester, many leading enterprises realize that real-time analytics—the analytics of what’s happening with data in the present—is an incredible competitive advantage. With it, they can act right away to serve customers, fix operational problems, power internet of things (IoT) apps, and respond decisively to competitors.

The report evaluates the top 11 vendors against 26 rigorous criteria for streaming analytics to help enterprise IT teams understand their options and make informed choices for their organizations. Google scored 5 out of 5 in Forrester’s report evaluation criteria of scalability, availability, aggregates, management, security, extensibility, ability to execute, solution roadmap, partners, community, and customer adoption. 

We are very excited about the productivity benefits offered by Cloud Dataflow and Cloud Pub/Sub. It took half a day to rewrite something that had previously taken over six months to build using Apache Spark.
Paul Clarke, Director of Technology, Ocado

While stream analytics is a leading business priority, real-life use cases frequently require batch data as an input as well. Google Cloud Platform (GCP) customers can simplify their pipeline development by reusing code across both batch and stream processing using Apache Beam, which also provides pipeline portability to OSS projects. Further, Google Cloud’s data analytics portfolio autoscales across ingestion, processing, and analysis, which eliminates the need for provisioning and makes handling varying volumes of streaming data automatic. We hear from users that they’re able to ingest and process data much more quickly than in the past, helping to get new business insights faster and allowing more users to do self-serve analytics.

2861

Of your peers have already watched this video.

17:30 Minutes

The most insightful time you'll spend today!

Explainer

GCP for Bioinformatics

Join Cloud GDE and developer Lynn Langit in this fast-paced session to get resources you can use to learn how to use the Google Cloud Platform for bioinformatics.

Lynn has created an open source course (on GitHub) to introduce researchers to using GCP to scale their analysis jobs. In this short talk, she’ll guide you through her course materials, so that you can get started learning using examples from genomics.

She talks is divided into four parts:

  • What is needed?
  • Why do we need to have patterns?
  • How can you use pattern information,
  • How can you learn more.

More Relevant Stories for Your Company

Podcast

Setting up Data Pipelines Easily for Streaming and Non-Streaming Data

We live in a world of incessant data. Our phones, factory equipment, smart watches, health devices, and other IoT devices are constantly streaming data. Some of the most popular uses of machine learning and AI are also based on streaming data. Think real-time credit card fraud detection, or software that

Blog

Haaretz on Google Cloud Guarantees Faster, Reliable & Responsive Services to its Audience

Israeli centenarian newspaper, Haaretz relied on on-prem infrastructure to serve readers digitally. As the need for scalability, security and business intelligence grew alongside their readership, Haaretz was looking for more than just a cloud-based solution to replace their infrastructure. Inon Gershovitz, CTO, Haaretz takes us through the journey of recreating

Case Study

Apache and Dataflow Help with Real-time Indices Processing for Financial Institutions

Financial institutions across the globe rely on real-time indices to inform real-time portfolio valuations, to provide benchmarks for other investments, and as a basis for passive investment instruments including exchange-traded products (ETPs). This reliance is growing—the index industry dramatically expanded in 2020, reaching revenues of $4.08 billion. Today, indices are calculated and distributed by

Explainer

Data Leaders in 2021 and Beyond: How to Prepare

As technologies and workplace environments are quickly evolving, how people and businesses leverage data in their day-to-day workflow is also changing. Data leaders are an emerging force navigating and charting pathways forward towards new horizons for how people and organizations experience data. In this presentation, Pedro Arellano, Product Marketing Director,

SHOW MORE STORIES