6084
Of your peers have already watched this video.
1:30 Minutes
The most insightful time you'll spend today!
Explore the Innovations and Architecture Powering Spanner and BigQuery
Previously, databases had architectures with tightly coupled storage and compute. This resulted in higher latency, and with faster networks these constraints no longer surface. With Google Cloud’s BigQuery and CloudSpanner, the storage and compute architecture have been separated, allowing for better scalability and availability to address businesses’ high throughput data needs.
Watch the video to understand how these database and analysis products leverage Google’s distributed storage system, in-house custom network hardware and software, internal cluster management system and more!
Home Depot Leverages Google Cloud’s BigQuery and DataFlow to Break Data Silos and Craft Personalized CX

5013
Of your peers have already read this article.
4:00 Minutes
The most insightful time you'll spend today!
The Home Depot, Inc., is the world’s largest home improvement retailer with annual revenue of over $151B. Delighting our customers—whether do-it-yourselfers or professionals—by providing the home improvement products, services, and equipment rentals they need, when they need them, is key to our success.
We operate more than 2,300 stores throughout the United States, Canada, and Mexico. We also have a substantial online presence through HomeDepot.com, which is one of the largest e-commerce platforms in the world in terms of revenue. The site has experienced significant growth both in traffic and revenue since the onset of Covid-19.
Because many of our customers shop at both our brick-and-mortar stores and online, we’ve embarked on a multi-year strategy to offer a shopping experience that seamlessly bridges the physical and digital worlds. To maximize value for the increasing number of online shoppers, we’ve shifted our focus from event marketing to personalized marketing, as we found it to be far more effective in improving the customer experience throughout the sales journey. This led to changing our approach to marketing content, email communications, product recommendations, and the overall website experience.
Challenge: launching a modern marketing strategy using legacy IT
For personalized marketing to be successful, we had to improve our ability to recognize a customer at the point of transaction so we could—among other things—suspend irrelevant and unnecessary advertising. Most of us have experienced the annoyance of receiving ads for something we’ve already purchased, which can degrade our perception of the brand itself. While many online retailers can identify 100% of their customer transactions due to the rich information captured during checkout, most of our transactions flow through physical stores, making this a more difficult problem to solve.
Our old legacy IT system, which ran in an on-premises data center and leveraged Hadoop, also challenged us since maintaining both the hardware and software stack required significant resources. When that system was built, personalized marketing was not a priority, so it took several days to process customer transaction data and several weeks to roll out any system changes. Further, managing and maintaining the large Hadoop cluster base presented its own set of issues in terms of quality control and reliability, as did keeping up with open-source community updates for each data processing layer.
Adopting a hybrid approach
As we worked through the challenges of our legacy system, we started thinking about what we wanted our future system to look like. Like many companies, we began with a “build vs. buy” analysis. We looked at several products on the market and determined that while each of them had their strengths, none was able to offer the complete set of features we needed.
Our project team didn’t think it made sense to build a solution from scratch, nor did we have access to the third-party data we needed. After much consideration, we decided to adopt a solution that combined a complete rewrite of the legacy system with the support of a partner to help with the customer transaction matching process.
Building the foundation on Google Cloud
We chose Google Cloud’s data platform, specifically BigQuery, Dataflow, DataProc, Cloud Storage, and Cloud Composer. Google Cloud platform empowered us to break down data silos and unify each stage of the data lifecycle from ingestion, storage, and processing to analysis and insights. Google Cloud offered best-in-class integration with open-source standards and provided the portability and extensibility we needed to make our hybrid solution work well. The open standards of BigQuery’s BQ Storage API allowed us to leverage fast BQ storage layers to be utilized with other compute platforms, e.g., DataProc.
We used BigQuery combined with Dataflow to integrate our first- and third-party data into an enterprise data and analytics data lake architecture. The system then combined previously siloed data and used BigQuery ML to create complete customer profiles spanning the entire shopping experience, both in-store and online.
Understanding the customer journey with the help of Dataflow and BigQuery
The process of developing customer profiles involves aggregating a number of first- and third-party data sources to create a 360-degree view of the customer based on both their history and intent. It starts with creating a single historical customer profile through data aggregation, deduplication, and enrichment. We used several vendors to help with customer resolution and NCOA (Change of Address) updates, which allows the profile to be house-holded and transactions to be properly reconciled to both the individual and the household. This output is then matched to different customer signals to help create an understanding of where the customer is in their journey—and how we can help.
The initial implementation used Google Dataflow, Google’s streaming analytics solution, to load data from Google Cloud Storage into BigQuery and perform all necessary transformations. The Dataflow process was converted into BQML (BigQuery Machine Learning) since this significantly reduced costs and increased visibility into data jobs. We used Google Cloud Composer, a fully managed workflow orchestration service, to help orchestrate all data operations and DataProc and Google Kubernetes Engine to enable special case data integration so we could quickly pivot and test new campaigns. The architecture diagram below shows the overall structure of our solution.

Taking full advantage of cloud-native technology
In our initial migration to Google Cloud, we moved most of our legacy processes in their original form. However, we quickly learned that this approach didn’t take full advantage of the cloud-native and more improved features Google Cloud offered such as auto scaling of resources, flexibility to decouple storage from the compute layer, and a wide variety of options to choose the best tool for the job. We refactored our Hadoop-based data pipelines written in Java-based Map Reduce and our Pig Latin jobs to Dataflow and BigQuery jobs. This dramatically reduced processing time and made our data pipeline code concise and efficient.
Previously, our legacy system processes ran longer than intended, and data was not used efficiently. Optimizing our code to be cloud-native and leveraging all the capabilities of Google Cloud services resulted in reduced run times. We decreased our data processing window from 3 days to 24 hours, improved resource usage by dramatically reducing the amount of compute we used to possess this data, and built a more streamlined system. This in turn reduced cloud costs and provided better insight. For example, DataFlow offers powerful native features to monitor data pipelines, enabling us to be more agile.
Leveraging the flexibility and speed of the cloud to improve outcomes
Today, using a continuous integration/continuous delivery (CI/CD) approach, we can deploy multiple system changes each week to further improve our ability to recognize in-store transactions. Leveraging the combined capabilities of various Google Cloud systems—BigQuery, DataFlow, Cloud Composer, Dataproc, and Cloud Storage–we drastically increased our ability to recognize transactions and can now connect over 75% of all transactions to an existing household. Further, the flexible Google Cloud environment coupled with our cloud-native application makes our team more nimble and better able to respond to emerging problems or new opportunities.
Increased speed has led to better outcomes in our ability to match transactions across all sales channels to a customer and thereby improve their experience. Before moving to Google Cloud, it took 48 to 72 hours to match customers to their transactions, but now we can do it in less than 24 hours.
Making marketing more personal—and more efficient
The ability to quickly match customers to transactions has huge implications for our downstream marketing efforts in terms of both cost and effectiveness. By knowing what a customer has purchased, we can turn off ads for products they’ve already bought or offer ads for things that support what they’ve bought recently. This helps us use our marketing dollars much more efficiently and offer an improved customer experience.
Additionally, we can now apply the analytical models developed using BQML and Vertex AI to sort customers into audiences. This allows us to more quickly identify a customer’s current project, such as remodeling a kitchen or finishing a basement, and then personalize their journey by offering them information on products and services that matter most at a given point through our various marketing channels. This provides customers with a more relevant and customized shopping journey that mirrors their individual needs.
Protecting a customer’s privacy
With this ability to better understand our customers, we also have the responsibility to ensure we have good oversight and maintain their data privacy. Google’s cloud solutions provide us the security needed to help protect our customers’ data, while also being flexible enough to allow us to support state and federal regulations, like the California Customer Privacy Act. This way we can provide a customer the personalized experience they desire without having to fear how their data is being used.
With flexible Google Cloud technology in place, The Home Depot is well positioned to compete in an industry where customers have many choices. By putting our customers’ needs first, we can stay top of mind whenever the next project comes up.
2876
Of your peers have already watched this video.
26:30 Minutes
The most insightful time you'll spend today!
How to Build a Data Pipeline Across Hybrid and Multi-region Infrastructures
Building a data pipeline on Google Cloud is one of the most common things enterprises do. Increasingly, organizations want to build these data pipelines across hybrid infrastructures.
Using Apache Kafka as a way to stream data across hybrid and multi-region infrastructures is a common pattern to create a consistent data fabric. Using Kafka and Confluent allows customers to integrate legacy systems and Google services like BigQuery and Dataflow in real time.
Learn how to build a robust, extensible data pipeline starting on-premises by streaming data from legacy systems into Kafka using the Kafka Connect framework.
This session highlights how to easily replicate streaming data from an on-premises Kafka cluster to Google Cloud Kafka cluster. Doing this integrates legacy applications and analytics in the cloud, using different Google services like AI Platform, AutoML, and BigQuery.
11 Google Cloud Analytics Tools, Each Explained Simply in Under 2 Minutes

3563
Of your peers have already read this article.
22:00 Minutes
The most insightful time you'll spend today!
Need a quick overview of Google Cloud analytics technologies? Quickly learn these 11 Google Cloud products—each explained in under two minutes.
BigQuery in a minute
Storing and querying massive datasets can be time consuming and expensive without the right infrastructure. This video gives you an overview of BigQuery, Google’s fully-managed data warehouse. Watch to learn how to ingest, store, analyze, and visualize big data with ease.
Firestore in a minute
Cloud Firestore is a NoSQL document database that lets you easily store, sync, and query data for your mobile and web apps, at global scale. In this video, learn how use Firestore and discover features that simplify app development without compromising security.
Cloud Spanner in a minute
Cloud Spanner is a fully managed relational database with unlimited scale, strong consistency, and up to 99.999% availability. In this video, you’ll learn how Cloud Spanner can help you create time-sensitive, mission critical applications at scale.
Cloud SQL in a minute
Cloud SQL is a fully-managed database service that helps you set up, maintain, manage, and administer your relational databases on Google Cloud. In this video you’ll learn how Cloud SQL can help you with time-consuming tasks such as patches updates, replicas, and backups so you can focus on designing your application.
Memorystore in a minute
Memorystore is a fully managed and highly available in-memory service for Google Cloud applications. This tool can automate complex tasks, while providing top-notch security by integrating IAM protocols without increasing latency. Watch to learn what Memorystore is and what it can do to help in your developer projects.
Bigtable in a minute
Cloud Bigtable is a fully managed, scalable NoSQL database service for large analytical and operational workloads. In this video, you’ll learn what Bigtable is and how this key-value store supports high read and write throughput, while maintaining low latency.
BigQuery ML in a minute
BigQuery ML lets you create and execute machine learning models in BigQuery by using standard SQL queries. In this video, learn how you can use BigQuery ML for your machine learning projects.
Dataflow in a minute
Dataflow is a fully managed streaming analytics service that minimizes latency, processing time, and cost through autoscaling and batch processing. In this video, learn how it can be used to deploy batch and streaming data processing pipelines.
Cloud Pub/Sub in a minute
Cloud Pub/Sub is an asynchronous messaging service that decouples services that produce events from services that process events. In this video, you’ll learn how you can use it for message storage, real-time message delivery, and much more, while still providing consistent performance at scale and high availability.
Dataproc in a minute
Dataproc is a managed service that lets you take advantage of open source data tools like Apache Spark, Flink and Presto for batch processing, SQL, streaming, and machine learning. In this video, you’ll learn what Dataproc is and how you can use it to simplify data and analytics processing.
Data Fusion in a minute
Cloud Data Fusion is a fully managed, cloud-native, enterprise data integration service for quickly building and managing data pipelines. In this video, you’ll learn how Cloud Data Fusion can help you build smarter data marts, data lakes, and data warehouses.
Getting Started with BigQuery: 10 Quick Steps

3801
Of your peers have already read this article.
12:30 Minutes
The most insightful time you'll spend today!
BigQuery is Google’s fully managed, petabyte scale, low cost analytics data warehouse. BigQuery is NoOps—there is no infrastructure to manage and you don’t need a database administrator—so you can focus on analyzing data to find meaningful insights, use familiar SQL, and take advantage of our pay-as-you-go model.
That’s great, right?
So let’s get started. At the end of this, you will be able to you will use Google Cloud Client Libraries for .NET to query BigQuery public datasets with C#.
This quick course will go over:
- Setup and Requirements
- Enabling the BigQuery API
- Authenticating API requests
- Setup Access Control
- Installing the BigQuery client library for C#
- Query the works of Shakespeare
- Query the GitHub dataset
- Caching and statistics
- Loading data into BigQuery
How the City of Memphis Uses Technology to Identify 75 Percent More Potholes

7144
Of your peers have already read this article.
4:30 Minutes
The most insightful time you'll spend today!
At 340 square miles, the City of Memphis is among the largest in the United States in terms of land area. Memphis has over 6,800 lane-miles of city streets, enough to drive back and forth to Los Angeles four times. Keeping these streets well maintained and safe for citizens and visitors is a major priority for the city.
Lots of traffic, lots of roads, and a four-season climate prone to wintertime freeze-thaw-refreeze cycles means the opportunity for potholes. Although the city aims to fill potholes within five business days of notification, it can take longer, especially during winter and early spring. Last year, the city’s Public Works crews repaired some 63,000 potholes, only 20% of which were reported by residents. Approximately 32,000-man-hours each year are spent repairing potholes, with seasonal fluctuations requiring ten to twelve Street Maintenance crews working steadily during the winter months. Still, many went unreported, leading the city to flag pothole request resolution under “needs improvement” on its open data portal website.
Like many large cities, Memphis also struggles with vacant and blighted properties. Nearly 15,000 properties in Memphis are likely vacant, and city officials contend that many are owned by out-of-town investors who live elsewhere and do not take necessary restoration or maintenance steps. These properties can decrease the value of surrounding real estate and discourage new businesses and other residents from moving to an area. Citizen frustration and concerns over the number of blighted properties has made blight eradication a major focus of the City of Memphis.
Historically, residents reported potholes and blighted properties by calling 311, or more recently by using the Memphis 311 app. However, these reports only covered about 20 percent of the problems — often the worst cases. And by the time residents took the initiative to submit a 311 report, they usually weren’t feeling good about the situation.
Recognizing that potholes and vacant properties are often the most visible indicators of whether a city government is doing its job efficiently, Memphis Mayor Jim Strickland and CIO Mike Rodriguez began looking for ways they could apply technology to fix the problems. Mike approached Google for ideas, and Google recommended conducting a machine learning proof-of-concept (POC) with SpringML, a Google Cloud Partner.
“Memphis is focused on easy living, and we want to do everything we can to keep our citizens happy,” says Mike Rodriguez. “Working with Google and SpringML to reduce potholes and urban blight using machine learning and artificial intelligence was an easy decision.”
Bringing machine learning to city operations and budgets
The city’s goal is to detect potholes and abandoned properties by analyzing video footage of roads and residential properties. It wanted to classify potholes by width and depth, and share the information with workers who can repair them. For abandoned properties, it wanted to enable more strategic deployment of resources for homeowners citywide and take action to hold neglectful property owners accountable.
The POC began by training TensorFlow models for ML object detection using preconfigured AI Platform Deep Learning VM Images on Compute Engine. SpringML helped set up cameras and developed a user interface to collect pothole data and automate the 311 ticketing process.
Together, the teams analyzed 30 days of video from a moving city bus and high-resolution video from 360-degree cameras mounted to a code enforcement vehicle, overlaid with data from 311 reports. As the models were refined, accuracy quickly climbed from 50 percent to over 90 percent as models were taught to differentiate a pothole from a manhole cover or other object.
The city also imported routes, potholes, and paving data along with geolocation data from ArcGIS and Google Maps into BigQuery to better understand street conditions and the proximity of potholes to one another. BigQuery also analyzes city property records, tax records, 311 reports, and third-party survey data on-demand to predict where homes are starting to become run down and where neighborhood decay is most likely to occur. The SpringML team created a pilot analysis to begin vacant property protections and developed a user interface tool to interact with the model’s results.
“Google Cloud Platform made it possible for us to experiment with machine learning and artificial intelligence to help solve our city’s problems while working within the budget constraints of a municipal IT organization,” says Mike. “Google turned a ‘nice to have’ into a ‘let’s do this!'”
Identifying 75 percent more potholes
Memphis expects to substantially reduce the number of potholes on its streets, creating a better driving experience for residents and visitors alike. Because drivers won’t be as likely to swerve to miss a pothole, streets will be safer and friendlier to bicycles and scooters. Fewer potholes will also save the city between $10,000 and $20,000 annually in city claims that it pays out in cases where vehicle damage results from a pothole that was not addressed in a timely manner.
“Historically, Public Works has relied primarily upon Street Maintenance crews to proactively locate and fill potholes. As Memphis has over 6,800 lane-miles of public streets, it is a daunting task to reliably survey the entire system in an efficient and systematic way,” says Robert Knecht, Public Works Director for the City of Memphis. “The outcome of the data collected will be invaluable to Public Works so that it can ensure it is managing the city’s street system in a more proactive manner.”
Memphis will be able to better prioritize road maintenance based on condition and impact, increasing the efficiency of its Public Works road crews. Analyzing video of streets also gave the city visibility into issues it wasn’t previously aware of, such as curbs, gutters, and manhole covers that had been mistakenly paved over and need to be excavated. The ML process is easily transferrable to other concerns as well, helping the city identify illegal signs or spools of cable hanging on light posts that could be potentially unsafe.
Helping communities recover and thrive
Memphis is also having success in analyzing predictive trends to combat high rates of abandoned and blighted properties, surpassing 97.5 percent accuracy. “In the past, Public Works experimented with comprehensive, city-wide blight identification by using approximately 200 volunteers to survey and photograph over 237,000 city parcels. This effort was costly, took a long time to complete, and resulted in inconsistent data collection,” says Robert. “Blighted property conditions can change quickly in a city the size of Memphis. Now, with this new technology, Memphis will be able to make a significant difference in the efforts to proactively and comprehensively identify and manage blighted and substandard properties.”
Code Enforcement with better data-driven detection mechanisms enables the city to also identify cases where homeowners are not physically or financially able to keep up with the challenges of homeownership and make them aware of resources that are available to assist them. Memphis Code Enforcement can do a better job of finding people living in derelict properties that pose hazards to inhabitants’ health and safety, and help them fix those problems or find a new place to live.
“Using SpringML and Google Cloud Platform to detect indicators of vacant or blighted properties will help Memphis create safer neighborhoods that will be more attractive to businesses and home buyers,” says Mike. “Property values and employment will go up, crime will go down, and social services can be more focused and effective.”
Revolutionizing service delivery for citizens
Memphis is proving the viability of a cost-effective, cloud-based machine learning model that other cities can follow. The city is already looking into new applications of AI and ML that will further improve city services and help it build a better future for its 652,000 residents.
As part of his commitment to a transparent government, Memphis Mayor Jim Strickland created an open data policy that commits to releasing raw data and sharing it with citizens in a variety of downloadable formats. Going forward, this transparency will help citizens understand how their needs are being served and uncover new, innovative use cases for AI and ML.
“Our goal is to become a smart city, and technologies such as Google Cloud Platform and SpringML put us ahead of the game,” says Mayor Strickland. “Google understands data, and there isn’t a better company to help us analyze our data resources for actionable insights.”
More Relevant Stories for Your Company

Easy Access to Stream Analytics with Google Cloud
By 2025, more than a quarter of the data created in the global datasphere will be real-time in nature. “This is important because in the real-time world, “the window of opportunity diminishes and goes away really fast. You want to be able to respond to your customer needs, their asks,
Apache and Dataflow Help with Real-time Indices Processing for Financial Institutions
Financial institutions across the globe rely on real-time indices to inform real-time portfolio valuations, to provide benchmarks for other investments, and as a basis for passive investment instruments including exchange-traded products (ETPs). This reliance is growing—the index industry dramatically expanded in 2020, reaching revenues of $4.08 billion. Today, indices are calculated and distributed by

Analytics Hub for Secure Data Sharing and Analytics Unlocks True Data Value and Insights
Customers tell us that sharing and exchanging data with other organizations is a critical element of their analytics strategy, but it’s hamstrung by unreliable data and processes, and only getting harder with security threats and privacy regulations on the rise. Furthermore, traditional data sharing techniques use batch data pipelines that are

WPP Innovates with the Cloud
Unlocking the Power of Data and Creativity withthe Cloud: The WPP Story






