Your DW Need Scaling Up? Try What This Company Did: It Can Run 25,000 Events a Second

5057
Of your peers have already read this article.
7:30 Minutes
The most insightful time you'll spend today!
With access to more data than ever before, companies have never been better positioned to adopt precision marketing methods and target the right customers at the right time. Emarsys, a digital marketing platform, enables its clients to collect, analyze, and act on a wide variety of data. From websites to mobile apps to emails, Emarsys’ customers can handle data from all its digital channels on a single, easy-to-use platform. Emarsys also makes sure that customers receive the highest quality data possible, making for smarter decisions and better business practices.
“We were close to the limits of our internal data warehouse, scalability-wise. We didn’t want to get to the point where we’d have to delete data or cancel projects. In Google Cloud we saw a platform that could scale with our ambitions and be optimized for AI and real-time solutions.”
—Levente Otti, Head of Data, Emarsys
Since launching as an email solutions provider in 2000, Emarsys has grown into the world’s largest independent digital marketing platform, with more than 2,500 clients worldwide and reaching more than 1.4 billion people. By 2016, the company felt that its existing data warehouse platform was close to its limits, affecting not just day-to-day operations but also important strategic goals.
“We were close to the limits of our internal data warehouse, scalability-wise. We didn’t want to get to the point where we’d have to delete data or cancel projects,” says Levente Otti, Head of Data at Emarsys. “In Google Cloud, we saw a platform that could scale with our ambitions and be optimized for AI and real-time solutions.”
Minimal maintenance, unlimited scale with Google Cloud
Digital marketing is a highly competitive environment. Emarsys works alongside big players with a huge market share on the one hand and smaller, specialist companies on the other. It has thrived by successfully combining the all-inclusive offerings of the former with the agility of the latter, constantly looking for ways to innovate and improve. In recent years, the company had started to feel that the ability to handle large quantities of data was no longer enough. The next challenge was speed. “We truly believe that in the future, everything will be done in real time, including data processing, analytics, and AI predictive models,” Levente says.
At the start of 2016, Emarsys’ existing data warehouse was a software-as-a-service solution running on-premises, which required hardware and software maintenance in order to keep up with the company’s growing appetite for data-heavy use cases such as prediction and analytics. The existing platform had proven its worth processing large amounts of data in batches, but its real-time capabilities were limited. Moreover, Emarsys had begun to experiment with AI technology, but found that its data warehouse couldn’t scale to accommodate some of the more resource-intensive processes, such as training the predictive models. The company decided that it needed a new, cloud-based data platform.
After evaluating some of the leading cloud providers, Emarsys chose Google Cloud for its mature AI capabilities and its ease of use. “With the other solutions, we still had to rent virtual machines and hardware and be responsible for maintenance. At the time, Google Cloud was the only provider that could take that management overhead away from us, while keeping customers accounted for on every query level,” Levente says.
To implement its new data platform, Emarsys teamed up with Google Cloud Partner Aliz. Over a series of meetings, workshops, and architecture reviews, Aliz helped Emarsys navigate the Google Cloud ecosystem to find the right products for the solution it was looking for. “Aliz really helped us set off in the right direction,” explains Levente.
With Google BigQuery, we can run queries which process terabytes of data, in seconds. We can also develop our own user-defined functions, incorporating Bayesian statistics into our predictive algorithms. That means we can take into account historical data, resulting in much more accurate predictions in a scalable way within seconds.”
—Levente Otti, Head of Data, Emarsys
Emarsys’ new data platform would actually be two: one platform for batch processing data and one for real-time analysis and interactions. Firstly, a proprietary publishing component gathered all the data points from Emarsys’ various channels including the website, mobile, emails, and custom events. With Cloud Pub/Sub and Cloud Dataflow, Emarsys transported and processed the data into BigQuery, which allows for further work and reviews that take into account errors or delayed events. After this, the data was exported to the main batch processing platform, which ran on BigQuery. For the real-time analytics, Emarsys used Cloud Bigtable to access data and Cloud Dataflow to pipeline it into the real-time platform, which could communicate with AI components or interaction components via an API to deliver real-time interactions with customers.
On top of the overall data infrastructure, Emarsys built a new AI platform with Google Cloud components. Training the predictive models had been an issue in the past due to the large number of resources required, so Emarsys chose to use Google Kubernetes Engine clusters, which can scale up and down on demand, without the need for hardware configuration or management. The trained models were held securely in Cloud Storage. From here, they were integrated with BigQuery for power and flexibility, allowing Emarsys to improve not just the speed of its AI predictions but also the quality.
“With Google BigQuery, we can run queries which process terabytes of data, in seconds,” shares Levente. “We can also develop our own user-defined functions incorporating Bayesian statistics into our predictive algorithms. That means we can take into account historical data, resulting in much more accurate predictions in a scalable way within seconds.”
Real-time insight, long-term satisfaction
Google Cloud enabled Emarsys to build a scalable data and AI platform that delivers powerful, actionable insights in real time. According to Levente, the company wanted to spend less time managing overload and more time considering how it should handle data. An immediate result of the new platform has been that data is now available in a scalable way, without hardware additions and management.
“With Google Cloud, we’ve been able to build a truly real-time data platform. The norm used to be daily batch processing of data. Now, if an event happens, marketing actions can be executed within seconds, and customers can react immediately. That makes us very competitive in our market.”
—Levente Otti, Head of Data, Emarsys
The clear and innovative pricing schemes of Google Cloud have also brought a new level of accountability to Emarsys’ costs in a way that wasn’t possible with its on-premises infrastructure. “Now that we only pay for what we use, we can assign costs to specific customers or queries, which has a huge impact on our pricing and product development strategies,” says Levente.
Thanks to the power of BigQuery and the scale at which it can handle data, Emarsys can now apply its analytics and AI tools to their full potential. “It’s very important to enable our clients to create the best possible experience for customers,” says Levente. At the same time, the company has cut its AI platform costs by 70% with Kubernetes while increasing scalability compared to the previous solution. The whole data platform was built to be scalable, and its first big test came during the retail peak of Black Friday, when it comfortably handled 250,000 events per second. “Perhaps the biggest impact on the business came with the real-time nature of the new platform,” says Levente.
“With Google Cloud, we’ve been able to build a truly real-time data platform,” he explains. “The norm used to be daily batch processing of data. Now, if an event happens, marketing actions can be executed within seconds, and customers can react immediately. That makes us very competitive in our market.”
Since implementing the new platform, Emarsys has continued to innovate with it and is about to release a new Real-Time Decision Framework, which will provide customers with even more real-time products and tools. The company continues to work with Aliz and Google Cloud, exploring other products such as Google BigQuery ML and TensorFlow to improve its AI processes. “We had a problem that we wanted to tackle now, and for us, Google Cloud was the best way of doing that,” says Levente. “But it was also about looking ahead. We felt that Google Cloud offered us the best way of future-proofing our platform.”
Notebook Executor Feature of Vertex AI Workbench to Schedule Notebooks Ad Hoc or on Recurring Basis

4442
Of your peers have already read this article.
3:00 Minutes
The most insightful time you'll spend today!
When solving a new ML problem, it’s common to start by experimenting with a subset of your data in a notebook environment. But if you want to execute a long-running job, add accelerators, or run multiple training trials with different input parameters, you’ll likely find yourself copying code over to a Python file to do the actual computation. That’s why we’re excited to announce the launch of the notebook executor, a new feature of Vertex AI Workbench that allows you to schedule notebooks ad hoc, or on a recurring basis. With the executor, your notebook is run cell by cell on Vertex AI Training. You can seamlessly scale your notebook workflows by configuring different hardware options, passing in parameters you’d like to experiment with, and setting an execution schedule, all via the Console UI or the notebooks API.
Built to Scale
Imagine you’re tasked with building a new image classifier. You start by loading a portion of the dataset into your notebook environment and running some analysis and experiments on a small machine. After a few trials, your model looks promising, so you want to train on the full image dataset. With the notebook executor, you can easily scale up model training by configuring a cluster with machine types and accelerators, such as NVIDIA GPUs, that are much more powerful than the current instance where your notebook is running.
Your model training gets a huge performance boost from adding a GPU, and you now want to run a few extra experiments with different model architectures from TensorFlow Hub. For example, you can train a new model using feature vectors from various architectures, such as Inception, ResNet, or MobileNet, all pretrained on the ImageNet dataset. Using these feature vectors with the Keras Sequential API is simple; all you need to do is pass the TF Hub URL for the particular model to hub.KerasLayer.

Instead of running these trials one by one in the notebook, or making multiples copies of your notebook (inceptionv3.ipynb, resnet50.ipynb, etc) for each of the different TF Hub URLs, you can experiment with different architectures by using a parameter tag. To use this feature, first select the cell you want to parameterize. Then click on the gear icon in the top right corner of your notebook.

Type “parameters” in the Add Tag box and hit Enter. Later when configuring your execution, you’ll pass in the different values you want to test.

In this example, we create a parameter called feature_extractor_model, and we’ll pass in the name of the TF hub model we want to use when launching the execution. That model name will be substituted into the tf_hub_uri variable, which is then passed to the hub.KerasLayer, as shown in the screenshot above.
After you’ve discovered the optimal model architecture for your use case, you’ll want to track the performance of your model in production. You can create a notebook that pulls the most recent batch of serving data that you have labels for, gets predictions, and computes the relevant metrics. By scheduling these jobs to execute on a recurring basis, you’ve created a lightweight monitoring system that tracks the quality of your model predictions over time. The executor supports your end-to-end ML workflow, making it easy to scale up or scale out notebook experiments written with Vertex AI Workbench.
Configuring Executions
Executions can be configured through the Cloud Console UI or the Notebooks API.
In your notebook, click on the Executor icon.

In the side panel on the right specify the configuration for your job, such as the machine type and the environment. You can select an existing image, or provide your own custom docker container image.

If you’ve added parameter tags to any of your notebook cells, you can pass in your parameter values to the executor.

Finally, you can choose to run your notebook as a one time execution, or schedule recurring executions.

Then click SUBMIT to launch your job.

In the EXECUTIONS tab, you’ll be able to track the status of your notebook execution.

When your execution completes, you’ll be able to see the output of your notebook by clicking VIEW RESULT.

You can see that an additional cell was added with the comment # Parameters, that overrides the default value for feature_extractor_model, with the value we passed in at execution time. As a result, the feature vectors used for this execution came from a ResNet50 model instead of an Inception model.
What’s Next?
You now know the basics of how to use the notebook executor to train with a more performant hardware profile, test out different parameters, and track model performance over time. If you’d like to try out an end-to-end example, check out this tutorial. It’s time to run some experiments of your own!
A Breakdown of Cloud-based Data Ingestion Practices

8807
Of your peers have already read this article.
3:00 Minutes
The most insightful time you'll spend today!
Businesses around the globe are realizing the benefits of replacing legacy data silos with cloud-based enterprise data warehouses, including easier collaboration across business units and access to insights within their data that were previously unseen. However, bringing data from numerous disparate data sources into a single data warehouse requires you to develop pipelines that ingest data from these various sources into your enterprise data warehouse. Historically, this has meant that data engineering teams across the organization procure and implement various tools to do so. But this adds significant complexity to managing and maintaining all these pipelines and makes it much harder to effectively scale these efforts across the organization. Developing enterprise-grade, cloud-native pipelines to bring data into your data warehouse can alleviate many of these challenges. But, if done incorrectly, these pipelines can present new challenges that your teams will have to spend their time and energy addressing.
Developing cloud-based data ingestion pipelines that replicate data from various sources into your cloud data warehouse can be a massive undertaking that requires significant investment of staffing resources. Such a large project can seem overwhelming and it can be difficult to identify where to begin planning such a project. We have defined the following principles for data pipeline planning to begin the process. These principles are intended to help you answer key business questions about your effort and begin to build data pipelines that address your business and technical needs. Each section below details a principle of data pipelines and certain factors your teams should consider as they begin developing their pipelines.
Principle 1: Clarify your objectives
The first principle to consider for pipeline development is clarify your objectives. This can be broadly defined as taking a holistic approach to pipeline development that encompasses requirements from several perspectives: technical teams, regulatory or policy requirements, desired outcomes, business goals, key timelines, available teams and their skill sets, and downstream data users. Clarifying your objectives clearly identifies and defines requirements from each key stakeholder at the beginning of the process and continually checks development against these requirements to ensure the pipelines built will meet these requirements.This is done by first clearly defining the desired end state for each project in a way that addresses a demonstrated business need of downstream data users. Remember that data pipelines are almost always the means to accomplish your end state, rather than the end state itself. An example of an effectively defined end-state is “enabling teams to gain a better understanding of our customers by providing access to our CRM data within our cloud data warehouse” rather than “move data from our CRM to our cloud data warehouse”. This may seem like a merely semantic difference, but framing the problem in terms of business needs helps your teams make technical decisions that will best meet these needs.
After clearly defining the business problem you are trying to solve, you should facilitate requirement gathering from each stakeholder and use these requirements to guide the technical development and implementation of your ingestion pipelines. We recommend gathering stakeholders from each team, including downstream data users, prior to development to gather requirements for the technical implementation of the data pipeline. These will include critical timelines, uptime requirements, data update frequency, data transformation, DevOps needs, and security, policy, or regulatory requirements by which a data pipeline must meet.
Principle 2: Build your team
The second principle to consider for pipeline development is build your team. This means ensuring you have the right people with the right skills available in the right places to develop, deploy, and maintain your data pipelines. After you have gathered your pipeline requirements, you can begin to develop a summary architecture that will be used to build and deploy your data pipelines. This will help you identify the human talent you will need to successfully build, deploy, and manage these data pipelines and identify any potential shortfalls that would require additional support from either third-party partners or new team members.
Not only do you need to ensure you have the right people and skill sets available in aggregate, but these individuals need to be effectively structured to empower them to maximize their abilities. This means developing team structures that are optimized for each team’s responsibilities and their ability to support adjacent teams as needed.
This also means developing processes that prevent blockers to technical development whenever possible, such as ensuring that teams have all of the appropriate permissions they need to move data from the original source to your cloud data warehouse without violating the concept of least privilege. Developers need access to the original data source (depending on your requirements and architecture) in addition to the destination data warehouse. Examples of this are ensuring that developers have access to develop and/or connect to a Salesforce Connected App or read access to specific Search Ads 360 data fields.
Principle 3: Minimize time to value
The third principle to consider for pipeline development is minimize time to value. This means considering the long-term maintenance burden of a data pipeline prior to developing and deploying it in addition to being able to deploy a minimum viable pipeline as quickly as possible. Generally speaking, we recommend the following approach to building data pipelines to minimize their maintenance burden: Write as little code as possible. Functionally, this can be implemented by:
1. Leveraging interface-based data ingestion products whenever possible. These products minimize the amount of code that requires ongoing maintenance and empower users who aren’t software developers to build data pipelines. They can also reduce development time for data pipelines, allowing them to be deployed and updated more quickly.
- Products like Google Data Transfer Service and Fivetran allow for managed data ingestion pipelines by any user to centralize data from SaaS applications, databases, file systems, and other tooling. With little to no code required, these managed services enable you to connect your data warehouse to your sources quickly and easily.
- For workloads managed by ETL developers and data engineers, tools like Google Cloud’s Data Fusion provide an easy-to-use visual interface for designing, managing and monitoring advanced pipelines with complex transformations.
2. Whenever interface-based products or data connectors are insufficient, use pre-existing code templates. Examples of this include templates available for Dataflow that allow users to define variables and run pipelines for common data ingestion use cases, and the Public Datasets pipeline architecture that our Datasets team uses for onboarding.
3. If neither of these options are sufficient, utilize managed services to deploy code for your pipelines. Managed services, such as Dataflow or Dataproc, eliminate the operational overhead of managing pipeline configuration by automatically scaling pipeline instances within predefined parameters.
Principle 4: Increase data trust and transparency
The fourth principle to consider for pipeline development is increase data trust and transparency. For the purposes of this document, we define this as the process of overseeing and managing data pipelines across all tools. Numerous data ingestion pipelines that each leverage different tools or are not developed under a coordinated management plan can result in “tech sprawl”, which significantly increases the management overhead of data ingestion pipelines as the quantity of data pipelines increases. This becomes especially cumbersome if you are subject to service-level agreements, or legal, regulatory, or policy requirements for overseeing data pipelines. Preventing tech sprawl is, by far, the best strategy for dealing with it by developing streamlined pipeline management processes that automate reporting. Although this can theoretically be achieved by building all of your data pipelines using a single cloud-based product, we do not recommend doing so because it prevents you from taking advantage of features and cost optimizations that come with choosing the best product for your use case.
A monitoring service such as Google Cloud Monitoring Service or Splunk that automates metrics, events, and metadata collection from various products, including those hosted in on-premise and hybrid computing environments, can help you centralize reporting and monitoring of your data pipelines. A metadata management tool such as Google Cloud’s Data Catalog or Informatica’s Enterprise Data Catalog can help you better communicate the nuances of your data so users better understand which data resources are best fit for a given use case. This significantly reduces your pipeline’s governance burden by eliminating manual reporting processes that often result in inaccuracies or lagging updates.
Principle 5: Manage costs
The fifth principle to consider for pipeline development is manage costs. This encompasses both the cost of cloud resources and the staffing costs necessary to design, develop, deploy, and maintain your cloud resources. We believe that your goal should not necessarily be to minimize cost, but rather maximizing the value of your investment. This means maximizing the impact of every dollar spent by minimizing waste in cloud resource utilization and human time. There are several factors to consider when it comes to managing costs:
- Use the right tool for the job – Different data ingestion pipelines will have different requirements for latency, uptime, transformations, etc. Similarly, different data pipeline tools have different strengths and weaknesses. Choosing the right tool for each data pipeline can help your pipelines operate significantly more efficiently. This can reduce your overall cost, free up staffing time to focus on the most impactful projects, and make your pipelines much more efficient.
- Standardize resource labeling – Implement and utilize a consistent labeling schema across all tools and platforms to have the most comprehensive view of your organization’s spending. One example is requiring all resources to be labeled by the cost center or team at time of creation. Consistent labeling allows you to monitor your spend across different teams and calculate the overall value of your cloud spending.
- Implement cost controls – If available, leverage cost controls to prevent errors that result in unexpectedly large bills.
- Capture cloud spend – Capture your spend on all cloud resource utilization for internal analysis using a cloud data warehouse and a data visualization tool. Without it, you won’t understand the context of changes in cloud spend and how they correlate with changes in business.
- Make cost management everyone’s job – Managing costs should be part of the responsibilities of everyone who can create or utilize cloud resources. To do this well, we recommend making cloud spend reporting more transparent internally and/or implementing chargebacks to internal cost centers based on utilization.
Long-term, the increased granularity in cost reporting available within Google Cloud can help you better measure your key performance indicators. You can shift from cost-based reporting (i.e. – “We spent $X on BigQuery storage last month”) to value-based reporting (i.e. – “It costs $X to serve customers who bring in $Y revenue”).
To learn more about managing costs, check out Google Cloud’s “Understanding the principles of cost optimization” white paper.
Principle 6: Leverage continually improving services
The sixth principle is leverage continually improving services. Cloud services are consistently improving their performance and stability, even if some of these improvements are not obvious to users. These improvements can help your pipelines run faster, cheaper, and more consistently over time. You can take advantage of the benefits of these improvements by:
- Automating both your pipelines and pipeline management: Not only should data pipelines be automated, but almost all aspects of managing your pipelines can also be automated. This includes pipeline/data lineage tracking, monitoring, cost management, scheduling, access management and more. This helps reduce long-term operational costs of each data pipeline that can significantly alter your value proposition and prevent any manual configurations from negating the benefits of later product improvements.
- Minimizing pipeline complexity whenever possible: While ingestion pipelines are relatively easy to develop using UI-based or managed services, they also require continued maintenance as long as they are in use. The most easily maintained data ingestion pipelines are typically the ones that minimize complexity and leverage automatic optimization capabilities. Any transformation in a data ingestion pipeline is a manual optimization of the pipeline that may struggle to adapt or scale as the underlying services improve. You can minimize the need for such transformations by building ELT (extract, load, transform) pipelines rather than ETL (extract, transform, load) pipelines. This pushes transformations down to the data warehouse that is use a specifically optimized query engine to transform your data rather than manually configured pipelines.
Next steps
If you’re looking for more information about developing your cloud-based data platform, check out our Build a modern, unified analytics data platform whitepaper. You can also visit our data integration site to learn more and find ways to get started with your data integration journey.
Once you’re ready to begin building your data ingestion pipelines, learn more about how Cloud Data Fusion and Fivetran can help you make sure your pipelines address these principles.
3032
Of your peers have already watched this video.
24:00 Minutes
The most insightful time you'll spend today!
How to Modernize Your Data infrastructure and Analytics: Case Study
Data modernization is key to digital transformation. Legacy systems are slow, expensive and cannot keep up with changing requirements. Data teams often spend more time on data pipelines than they do on data analysis or data science.
in this video, Harish Ramachandraiah, Director Eng. & Analytics, Sunrun, speaks with Joel Mckelvey, Product Marketing Manager, Google Cloud. They discuss how Sunrun leveraged Looker and BigQuery to reduce the complexity of legacy extract-transform-load processes, improve database performance, and adapt quickly to changes in data.
Now Sunrun makes data accessible where it’s needed, in near-real time.
Mckelvey shows how a modern data stack helps companies simplify the data pipeline, consolidate vital data, and accelerate analytics performance while improving agility, efficiency, and data governance.
5116
Of your peers have already watched this video.
1:47 Minutes
The most insightful time you'll spend today!
How adidas’s Marketing Team Benefits from Google Cloud Marketing Platform
We live in a world where the consumer journey is no longer linear, making it difficult to know when to deliver a brand message or a product message.
“We talk a lot about the consumer journey no longer being linear. It’s very difficult to know when the consumer is going to come in contact with your message. And because of that it makes it much harder to deliver the right message at the right time,” says Chris Murphy, Head of Digital Experience, adidas.
To solve the problem, adidas turned to the Google Cloud Marketing Platform to use audience insights to inform the art of storytelling and help teams, from brand marketing to retail to e-commerce, deliver relevant messages across channels.
“In a singular environment like Google Cloud Marketing Platform, when you’re able to see audience insights and information…we know when someone should be receiving a brand message versus a product message — when they’re thinking about buying or when they’re actually ready to purchase,” says Murphy.
It wasn’t easy at the start. “Being in brand marketing for so long, there was this initial pushback on data and the science behind it, because it was an art. But what we’re seeing now is the ability for data and insights to really inform the art of storytelling. It allows us to understand the impact that we’re having in real time,” says Kelly Olmstead, VP, Brand Activation for North America, adidas.
Today adidas’ marketing team depends more heavily on data and analysis to guide its strategy. “Starting every meeting with data and insights, that’s something that we want to take outside of just the marketing organization, and really infuse it across our culture here at adidas,” says Olmstead.
Watch how adidas benefits from using the Google Cloud Marketing Platform.
Zeotap Uses Big Data to Enable Personalized, Precision Marketing at Scale

3352
Of your peers have already read this article.
5:30 Minutes
The most insightful time you'll spend today!
Customers today expect brands to know what they want. In fact, Forbes cited that 71% of customers feel frustrated when the shopping experience is impersonal. From interests to purchase habits, companies need to target customers effectively in order to stand out from the plethora of competitors in the online marketplace.
“Our mission is to help brands engage customers with the most appropriate marketing messages at the right place and at the right time while respecting consumer privacy and ensuring regulatory compliance. The experience of Google Cloud in working with large-scale datasets makes the platform an excellent fit for us.”
—Projjol Banerjea, co-founder and Chief Product Officer, Zeotap
Leveraging first-, second-, and third-party data, Zeotap’s Customer Intelligence Platform (CIP)—a CDP with additional identity resolution and third-party data enrichment capabilities—allows brands to engage with their known and unknown users with tailored marketing messages and next best offers or actions. It also offers out-of-box algorithms for standard marketing use cases such as promoting a specific product that will resonate with audiences and custom algorithms for very niche and specific use cases.
“Our mission is to help brands engage customers with the most appropriate marketing messages at the right place and at the right time while respecting consumer privacy and ensuring regulatory compliance. The experience of Google Cloud in working with large-scale datasets makes the platform an excellent fit for us,” says Projjol Banerjea, co-founder and Chief Product Officer at Zeotap.
“With BigQuery, we are able to operate more efficiently, removing 80% of the operation load that we would otherwise have to manage on a different cloud provider. As a result, we get to build and deliver products faster.”
—Sathish K S, VP Engineering, Zeotap
Improving data processing to deliver products faster
The cloud-native company migrated from its previous provider to Google Cloud in October 2019 because it was attracted to the benefits of BigQuery as a serverless cloud data warehouse. This was a feature that wasn’t available with its previous provider.
Sathish K S, Engineering Vice-President at Zeotap says, “With BigQuery, we are able to operate more efficiently, removing 80% of the operation load that we would otherwise have to manage on a different cloud provider. As a result, we get to build and deliver products faster.” Since coming onboard BigQuery, Zeotap has been able to increase the speed of its product delivery by 40%.
BigQuery also offers a much higher query limit, which is helpful because Zeotap can converge all of its previous operation load across different platforms into one data warehouse without worrying about it being overloaded.
One of the key services Zeotap offers is customer relationship management (CRM) platform integration. Using Dataflow, Zeotap can easily stream first-party data into BigQuery at scale.
The data integration capabilities of Cloud Data Fusion also complements Dataflow to help Zeotap enhance its API offerings. To illustrate, an ecommerce company may use Zeotap’s internal API to understand their customer’s cart and check-out habits and add Zeotap’s extra layer of data intelligence to it, so they can improve their marketing efforts. Cloud Data Fusion enables a seamless flow of real-time data so that the company can deliver the relevant ads or suggest products just before they make a purchase, which increases the likelihood of them buying more.
Zeotap uses Dataproc and Pub/Sub to manage the large flow of data that goes in and out of its database. Because many of its customers are pushing data in real time, Pub/Sub acts as a queuing engine that feeds them into Dataflow. Likewise, the data from Dataflow is also being channeled back smoothly to the customer endpoints using Pub/Sub.
Being able to automate all of its processes benefitted the DevOps team, as they now have time to focus on developing and improving Zeotap’s product offerings instead of spending time and energy on operations.
Managing a full migration seamlessly
Today, all of Zeotap’s data is hosted on Google Cloud. The company began its lift and shift migration in October 2019 and was complete by April 2020. “We have close to 405 data pipelines, and even after the migration, we had to verify the right datasets. We had terabytes of data to move, and despite this, our operations were not affected,” says Aditya Chandra, Vice-President of Infrastructure and Security at Zeotap.
Apart from the BigQuery capabilities, there were other major things Zeotap looked out for in its decision to move to a new cloud provider. The functional parity of Google Cloud was important so that it did not have to rebuild many of its existing structures. Zeotap also didn’t want any drop in performance during and after the move to a new cloud provider. Finally, as a multi-region company, it used to have to spot instance problems and run them on demand, which had cost implications. With the move, it no longer had to face such issues.
“Whatever Spark code that was already running on our existing system had to be lifted and shifted into Dataproc without any library changes because 60% of our codebase is out of there. We’re glad that this was possible,” says Ameya Agnihotri, Chief Technology Officer.
Managing business continuity with a trusted partner
Zeotap now works with Rackspace Technology (Rackspace) as it continues to optimize its infrastructure operations. As a marketing and technology company, Zeotap needs niche expertise to manage big data processing issues.
Ameya Agnihotri shares, “Given its global expertise in dealing with on-premises as well as cloud solutions, Rackspace can really add value because they are better equipped to troubleshoot and handle complex data operations.” Zeotap has plans to explore more of Google Cloud’s artificial intelligence (AI) and machine learning (ML) capabilities, so having a partner in this journey will be valuable to the team.
With so much data on hand, security is paramount to Zeotap. It does a lot of internal security assessments daily. This is one of the areas where the team works with Rackspace to ensure that it adopts best practices and stays on top of the latest updates. With its ability to handle big data operations, Rackspace also acts as an extension of the Zeotap team that supports the infrastructure side of the business.
“The level of engagement from Google Cloud has been excellent, particularly the ability to bring in expertise from different parts of the organization. The personal touch meant we could reach product owners and managers of specific products. This kind of support is what we really value and appreciate.”
—Ameya Agnihotri , Chief Technology Officer, Zeotap
Future plans and collaboration with Google Cloud
Zeotap is currently exploring Cloud Bigtable and API Gateway as it continues to grow the business and explore new types of services to offer to its customers. Although it has a software-as-a-service (SaaS) offering for customers to self-serve, some may prefer to do a server-to-server integration so that they do not have to worry about the back end at all. For this reason, Zeotap is exploring the possibility of integrating these customers’ APIs in the back end of its API, using Google Cloud’s API Gateway features that are available out of the box. The company is also looking forward to the new features being released on BigQuery, as well as exploring more of the alerting and monitoring capabilities of Google Cloud.
“The level of engagement from Google Cloud has been excellent, particularly the ability to bring in expertise from different parts of the organization,” says Ameya. He shares that the team has had many fortnightly calls when there were issues to address, and the Google Cloud team would help to fix the problems. “The personal touch meant we could reach product owners and managers of specific products. This kind of support is what we really value and appreciate.”
With big plans ahead for Zeotap that include becoming a Google Cloud Technology Partner in the near future, Swapnasarit Sahu, Chief Data and Analytics Officer, is confident that the company has chosen the right cloud solution. “We see Google as one of the frontrunners in the space of AI and ML solutions, which can really help data scientists in developing new and more innovative products and solutions, and we look forward to building great products with them,” he says.
More Relevant Stories for Your Company

How to Predict the Cost of a Managed Streaming and Batch Analytics Service
The value of streaming analytics comes from the insights a business draws from instantaneous data processing, and the timely responses it can implement to adapt its product or service for a better customer experience. “Instantaneous data insights,” however, is a concept that varies with each use case. Some businesses optimize

Making Streaming Analytics Simpler and More Cost-Effective With Cloud Dataflow
Streaming analytics helps businesses to understand their customers in real time and adjust their offerings and actions to better serve customer needs. It’s an important part of modern data analytics, and can open up possibilities for faster decisions. For streaming analytics projects to be successful, the tools have to be

Post-COVID: Times Driven by Data Analytics and Intelligence
As we think about economic recovery from COVID-19—both inside Google and outside through working with Google Cloud customers—we’ve made many important observations. Among them is the recognition that the ways software developers and IT practitioners work together will shift in the post COVID-19 world. Our economic recovery today will look

Speech Analytics and AI: Leverage Every Customer Interaction as a Data Point
Contact centers can transform customer experience using AI-driven speech analytics to evaluate every customer interaction and use it as a 'data point' to identify key patterns, enquiries, pain points, reviews and feedback to enhance customer experience with real-time, personalized recommendations. Knowlarity's AI-based cloud telephony solutions for businesses built with Google






