Google Migration and BigQuery Brings PedidosYa Closer towards its Goal of Becoming Data-driven - Build What's Next
Case Study

Google Migration and BigQuery Brings PedidosYa Closer towards its Goal of Becoming Data-driven

7243

Of your peers have already read this article.

2:00 Minutes

The most insightful time you'll spend today!

PedidosYa, a Latin American leader in the online food ordering space with over 20 million app downloads was looking to democratize data by gaining a secured access and building a comprehensive information ecosystem. PedidosYa was also challenged with its legacy data warehouse that couldn't keep with the brand's increasing analytics demands and was time-consuming and costly. To modernize their data warehouse and transform the analytics environment, PedidosYa chose Google Cloud for its serverless, managed, and integrated data platform, coupled with its seamless integration across open-source solutions. Also, Google Cloud's advanced cost and workload management coupled with its transparent log analytic gave the brand full visibility into any query performance issues to make improvements as they wanted. BigQuery helped PedidosYa achieve cost reduction per query by 5x and leverage its AL/ML-based stack for productivity benefits. Read on to learn more about PedidosYa's BigQuery and Google Cloud migration journey as a step closer towards becoming a data-driven organization.

Editor’s note: PedidosYa is the market leader for online food ordering in Latin America, serving 15 markets and over 400 cities. It’s also one of the largest brands within the German multinational company Delivery Hero SE. With over 20 million app downloads, PedidosYa provides the best online delivery experience through 71,000+ online partners, including restaurants, shops, drugstores, and specialized markets. 

Having constant access to fresh customer data is a key requirement for PedidosYa to improve and innovate our customer’s experience. Our internal stakeholders also require faster insights to drive agile business decisions. Back in early 2020, PedidosYa’s leadership tasked the data team to make the impossible possible. Our team’s mission was to democratize data by providing universal and secure access while creating a comprehensive information ecosystem across PedidosYa. We also had to achieve this goal while keeping costs under control— even during the migration stage and removing operational bottlenecks. 


Challenges with legacy cloud infrastructure

PedidosYa first built its data platform on top of AWS. Our data warehouse ran on Redshift, and our data lake was in S3. We used Presto and Hue as the user interfaces for our data analysts. However, maintaining this infrastructure was a daunting task. Our legacy platform couldn’t keep up with the increasing analytics demands. For example, the data stored on S3 complemented by Presto/Hue required high operational overhead. This was because Presto and our IAM (identity access management) didn’t integrate well in our legacy ecosystem. Managing individual users and mapping IAM roles with groups and Kerberos was operationally time-consuming and costly. Further, sharding access on the S3 files was far too complicated to enable seamless ACLs (access control lists).  

There were also challenges with workload management. Our data warehouse had batch data loaded overnight. If one analyst scheduled a query to run during the overnight ETL (extract, transform, load) workload, it would disrupt the current ETL task. This could stop the entire data pipeline. We’d have to wait until data engineers intervened with a manual fix.

It was also difficult to understand whether a query error was due to performance issues or platform resource exhaustion. This lack of clarity affected our data analysts’ ability to autonomously improve querying efficiency. Data team members needed to manually inspect personal queries looking for performance issues. Also,  the current architecture was prone to a ‘tragedy of the commons’ situation; it was seen as an unlimited and free resource. As a result, it was impossible to disentangle the infrastructure from different stakeholder teams, as all had very different needs. 

The decision to modernize our data warehouse

Given the growing challenges from our legacy platform, our tech team decided to transform our analytics environment with a modern data warehouse. They required the following key criteria from their next data platform: 

  • Scalability – The ability to grow with elastic infrastructure.
  • Cost control – Cost management and transparency. These factors promote efficiency and ownership—both key aspects of data democratization.
  • Metadata management – Intuitive data platform focusing on users’ previous SQL knowledge. Plus, being able to enrich the informational ecosystem with metadata,  to diminish data gatekeepers.
  • Ease of management – The team needed to reduce operational costs with a serverless solution. Data engineers wanted to focus on their key roles rather than acting as database administrators and infrastructure engineers. The team also wanted much higher availability, and to reduce the impact of maintenance windows and vacuum/analysis.
  • Data governance and access rights – With a growing employee base with varying data access requirements, the team needed a simple yet comprehensive solution to understand and track user access to data.

Migrating to Google Cloud

After exploring other alternatives, we concluded Google Cloud had an answer to each of our decision drivers. Google Cloud’s serverless, managed, and integrated data platform, coupled with its seamless integration across open-source solutions, was the perfect answer for our organization. In particular, the natural integration with Airflow as a job orchestrator and Kubernetes for flexible on-demand infrastructure was key.  

We  used Dataflow together with Pub/Sub and Cloud Functions for our data ingestion requirements, which has made our deployment process with Terraform seamless. Because we set up everything in our environment programmatically, operation time has diminished. Google Cloud reduced the deployment process from about 16 hours in our legacy platform to 4 hours.  This is partly due to the friendliness of automating the deployment (such as schema check, load test, table creation, build.) process with Terraform, Cloud Functions, Pub/Sub, Dataflow, and BigQuery on GCP. Input messages processed with Dataflow allow us to abstract and plan the schema changes according to the needs of the functional team. For example, schema changes raise an alarm, and then we can modify the raw layer table schema. By doing this, we ensure that backend modifications that we don’t control do not affect upper layers.

A key reason why we picked Google Cloud was because of its advanced cost and workload management coupled with its transparent log analytics. This information gives us a complete view into any query performance issues to make improvements on the fly. Further, we achieved a significant amount of cost savings by consolidating multiple tools to BigQuery.With BigQuery, we’ve been able to reduce our total cost per query by 5x.

This was due to a number of reasons:

  • Automating pipeline deployment made it much simpler to maintain the data processing processes. 
  • Analysts are conscious about what queries they’re running, resulting in running better, more optimized queries. 
  • Analysts use a Data Studio dashboard to see their queries and all the associated costs. As a result, there’s a lot more transparency for each persona.

 With these changes, we can easily manage and assign costs associated with each workload with their own cost centers using specific Google Cloud projects.

Change management is always challenging. However, BigQuery is intuitive and doesn’t have a steep learning curve from Hue/Hive on SQL basics. BigQuery also allowed the team to expand its capabilities and enabled them to properly work with nested structures, avoiding unnecessary joins and improving query efficiency. Additionally, we now use Data Catalog as our unique point of truth for metadata management. This allows our team to break the data access barriers and enable federation of data across the organization. By using Airflow to orchestrate everything, we keep track of every data stream. With this information, each end user can see their regularly used data entities’ status via the dashboard. This also adds transparency to our everyday data processes.

Finally, with Google Cloud’s IAM rules applied across the different products, data sharing and access is close to a noOps experience. We have programmatically implemented access according to roles and level access within the company. This allows certain pre-validated roles to view more sensitive information. These solutions help drive a more automated data governance experience. 

Up next: Google Cloud AI/ML

The new stack based on BigQuery has created significant productivity gains. Freed from the burden of operational management, PedidosYa’s data team can now focus on adding value through data tools and products.  

  • Our data engineers are better equipped to integrate constantly changing transactional and operational data.
  • The dataOps team can automate the infrastructure and provide autonomy to the end user.
  • Our data quality team can focus on bringing added value to data stakeholders. 
  • Data scientists and data analytics can spend more time analyzing data and less time asking data gatekeepers for data access.

PedidosYa can now democratize data access with a well-governed architecture. We are still at the beginning of our journey, but we are closer to achieving our vision of building a data-driven organization. Up next: expanding our artificial intelligence and machine learning capabilities.

Tune in to Google Cloud’s Applied ML Summit on June 10th, 2021, or listen on-demand later, to learn how to apply groundbreaking machine learning technology in your projects.

Blog

Cloud IoT Core Helps Businesses Leverage their IoT Data to Build a Competitive Edge

7101

Of your peers have already read this article.

1:30 Minutes

The most insightful time you'll spend today!

IoT devices produce tons of data that require an efficient, scalable and affordable way to analyze the information. IoT Core is a fully managed service for managing IoT devices that can bring a competitive edge for businesses. Learn how!

The ability to gain real-time insights from IoT data can redefine competitiveness for businesses. Intelligence allows connected devices and assets to interact efficiently with applications and with human beings in an intuitive and non-disruptive way. After your IoT project is up and running, many devices will be producing lots of data. You need an efficient, scalable, affordable way to both manage those devices and handle all that information. 

IoT Core is a fully managed service for managing IoT devices. It supports registration, authentication, and authorization inside the Google Cloud resource hierarchy as well as device metadata stored in the cloud, and the ability to send device configuration from other GCP or third-party services to devices. 

Main components

The main components of Cloud IoT Core are the device manager and the protocol bridges:

  • The device manager  registers devices with the service, so you can then monitor and configure them. It provides:
    • Device identity management 
    • Support for configuring, updating, and controlling individual devices
    • Role-level access control
    • Console and APIs for device deployment and monitoring
  • Two protocol bridges (MQTT and HTTP) can be used by devices to connect to Google Cloud Platform for:
    • Bi-directional messaging
    • Automatic load balancing
    • Global data access with Pub/Sub

How does Cloud IoT Core work?

Device telemetry data is forwarded to a Cloud Pub/Sub topic, which can then be used to trigger Cloud Functions as well as other third-party apps to consume the data. You can also perform streaming analysis with Dataflow or custom analysis with your own subscribers.

Cloud IoT Core supports direct device connections as well as gateway-based architectures. In both cases the real time state of the device and the operational data is ingested into Cloud IoT Core and the key and certificates at the edge are also managed by Cloud IoT Core. From Pub/Sub the raw input is fed into Dataflow for transformation, and the cleaned output is populated in Cloud Bigtable for real-time monitoring or BigQuery for warehousing and machine learning. From BigQuery the data can be used for visualization in Looker or Data Studio and it can be used in Vertex AI for creating machine learning models. The models created can be deployed at the edge using Edge Manager (in experimental phase). Device configuration updates or device commands can be triggered by Cloud Functions or Dataflow to Cloud IoT Core, which then updates the device.  

Design principles of Cloud IoT Core

As a managed service to securely connect, manage, and ingest data from global device fleets, Cloud IoT COre is designed to be:

  • Flexible, providing easy provisioning of device identities and enabling devices to access most of Google Cloud
  • IThe industry leader in IoT scalability and performance
  •  Interoperable, with supports for the most common industry-standard IoT protocols

Use cases

IoT use cases range across numerous industries. Some typical examples include:

  • Asset tracking, visual inspection, and quality control in retail, automotive, industrial, supply chain and logistics
  • Remote monitoring and predictive maintenance in oil & gas, utilities, manufacturing, and transportation
  • Connected homes and consumer technologies.
  • Vision intelligence in retail, security, manufacturing, and industrial sectors
  • Smart living in commercial, residential, and smart spaces 
  • Smart factories with predictive maintenance and real-time plant floor analytics

 For a more in-depth look into Cloud IoT Core check out the documentation.  

https://youtube.com/watch?v=76v16P-Wqe4%3Fenablejsapi%3D1%26

For more #GCPSketchnote, follow the GitHub repo. For similar cloud content follow me on Twitter @pvergadia and keep an eye out on thecloudgirl.dev.

Whitepaper

Forrester Surveyed Indian Retailers About Digital Transformation. Here’s What They Found

DOWNLOAD WHITEPAPER

3996

Of your peers have already downloaded this article

12:30 Minutes

The most insightful time you'll spend today!

As today’s empowered consumers demand more of the retail experience than ever before, leading retailers and brands in India are investing to rethink and reinvent in their customers’ cross-touchpoint experiences.

Our survey results demonstrate that retail decision makers understand that better customer experience can yield financial benefits, including faster revenue growth, and elevate the reach of influence and brand in the market.

Forty percent or more of retail executives are prioritizing revenue growth, improvement of customer experience (CX), and simplification of operations as the top priorities in their business agendas over the next year. 

The survey also covers:

  • Key Drivers For Retail Organizations To Migrate Application To Public Cloud 
  • Cloud Investments In The Retail Industry 
  • The Three Dimensions That The Industry’s Cloud Challenges Are Taking
  • The Top Agendas Retailers Want to Accomplish with the Public Cloud
Forrester’s retail report dives deep into the challenges Indian retailers are facing and what they want to accomplish with the cloud

Download Forrester’s Retail Report Now.

Blog

You Can Now ‘Listen’ to Over 50 Tech Blogs on Google Cloud Reader

4880

Of your peers have already read this article.

1:30 Minutes

The most insightful time you'll spend today!

You can now give your eyes some rest and yet catch up on the Google's latest tech blogs in audio format with Google Cloud Reader. Listen to your favorite from the 50 blogs or episodes on Google Podcasts, Apple Podcasts and Spotify.

🎧 Prefer to listen? Check out this episode on the Google Cloud Reader podcast

If you’re anything like me, you love reading, but also appreciate that sometimes your eyes need to be doing other things; whether it’s finding your exit off the highway, or keeping your puppy from destroying the couch.

And sometimes the thought of sitting down to read something just feels like it’s going to take valuable multi-tasking time away from my day. I know, I know, multitasking can be frowned upon, but it’s the way I live a good chunk of my life, and it’s working out so far. And while I’m not alone in my multitasking, I’m also not alone in my desire for a non-visual way to get this content, or any content.

*Google Cloud Reader enters the chat*

Google Cloud Reader is a podcast that lets you listen to the Google Cloud Blog posts that aren’t as dependent on visuals. This means they’re articles that are, or are adapted to be, less focused on graphs, or code samples, and instead describe the meaning behind those visual aids. 

It’s an easy, audible way to absorb content around all things new in Cloud, while still being able to make sure Ruthie doesn’t eat my work from home equipment. 

ruthie
Ruthie, a French Shepherd puppy, with her giant ears and feet dangerously close to filming equipment

So by now you’re probably thinking “OK, so you started a podcast during the pandemic, even though you definitely seemed like the type to start making sourdough”—and you’re right. My 53 plants agree with you. But rest assured, one can listen to an episode of this podcast *while* creating a macramé plant hanger, or waiting for bread to rise—multitasking, am I right?

We’re a little over 50 episodes/macrame plant hangers in, so you should check it out (Ruth and I would appreciate it).

Some of my personal favorites 

  • Beginners Guide to Painless Machine Learning – Learn how to get started with Google Cloud AI tools
  • Introducing GKE Autopilot: A Revolution in Managed Kubernetes – Learn more about GKE Autopilot, a revolutionary mode of operations for managed Kubernetes that lets you focus on your software, while GKE Autopilot manages the infrastructure.
  • Cook up your own ML recipes with AI Platform – ​​Learn about Mars Wrigley’s new ML-inspired recipe experiment on Google Cloud and how you can get started with your own.
  • Recovering Global Wildlife Populations using ML – Review Google’s Wildlife Insight’s ML project and help users create an image classification model for motion-sensor cameras (called camera traps) used to help protect wildlife in an non-invasive way by collecting and tagging species via pictures.

Let me know your favorite episodes, and what other articles you’d like to hear on Twitter @jbrojbrojbro!

No matter why you prefer an audio format, we’ve got you covered; Google Cloud Reader, where we read the tech blog for you, and to you.

Get all the Google Cloud Reader on your favorite podcast platform, including Google PodcastsApple Podcasts, and Spotify.

Blog

Army Google Workspace is Now Operational and Live

3215

Of your peers have already read this article.

1:00 Minutes

The most insightful time you'll spend today!

The U.S. Army is rolling out Google Workspace for soldiers. It is live and operational months after the service quietly began testing the software suite as a potential solution to previous IT issues. Read now!

Soldiers can now use their CAC cards to log into the service’s newly established Google Workspace account. With it, soldiers who don’t need the full spectrum of information-sharing provided by Army365 can nevertheless find comrades, schedule meetings and coordinate other activities.

“All new entrants to the Army will be provisioned automatically with a Google Workspace account with a new email address … when they receive their CAC card,” Raj Iyer, the service’s chief information officer, said in a social-media post.

“Soldiers will retain their Google account till they fully transition to their first unit after completion of basic training and AIT [advanced infantry training], at which point their commanders will determine if they need an Army365 account.”

The directive applies to all active-duty, reserve, and National Guard soldiers.

Blog

Media CDN to Intelligently Deliver Streaming Experiences to Viewers around the World!

3304

Of your peers have already read this article.

3:00 Minutes

The most insightful time you'll spend today!

Video streaming constitute 50+% of internet bandwidth traffic! Imagine the merits for media companies with Media CDN built on the success of Cloud CDN portfolio for web & API acceleration and combine it with immersive media experiences.

The digital media and entertainment industry is experiencing dramatic growth, as audiences migrate to online experiences and content providers seek to deliver new and innovative content. According to The Global Internet Phenomena Report, streaming video accounted for 53.7% of internet bandwidth traffic, up by 4.8% from a year ago. This rapid growth of over-the-top content is straining existing infrastructure, fueling media companies’ shift to the public clouds with their global presence and greater distribution capacities. In addition, other use cases such as gaming, social networks, AR/VR experiences, and education continue to fuel the need for intelligent media services and operations.

Today, at the 2022 NAB Show Streaming Summit, we’re excited to announce the general availability of Media CDN — a modern, extensible platform for delivering immersive experiences with unparalleled scale and intelligence. Media CDN will enable media and entertainment customers to efficiently and intelligently deliver streaming experiences to viewers anywhere in the world. The same infrastructure that Google has built over the last decade to serve YouTube content to over 2 billion users is now being leveraged to deliver media at scale to Google Cloud customers with Media CDN.

Unparalleled planet-scale reach and scale


Media CDN’s foundational advantage is the Google network. We have invested decades of resources to build tremendous capacity and reach in over 200 countries and more than 1,300 cities around the globe. Modern video applications are sensitive to fluctuations in latency, so getting content closer to users enables higher bitrates and reduces rebuffers, resulting in a superior experience for the end user. Media CDN builds on the success of the existing Cloud CDN portfolio for web and API acceleration and complements it by enabling delivery of immersive media experiences.

In addition to running on planet-scale infrastructure, Media CDN tailors delivery protocols to individual users and network conditions. Media CDN includes out-of-the-box support for QUIC (HTTP/3), TLS 1.3, and BBR, optimizing for last-mile delivery . When the Chrome team rolled out widespread support for QUIC, video rebuffer time decreased by more than 9% and mobile throughput increased by over 7%.

Media CDN also achieves industry-leading offload rates. With multiple tiers of caching, we minimize calls to origin — even for infrequently accessed content. This alleviates performance or capacity stress in the content origin and saves costs. These features are built into the product and seamlessly support customer content hosted on Google Cloud, on-premises, or on a third-party cloud.

We are excited to leverage Media CDN to continue to deliver an exceptional streaming experience for Stan users across Australia. With Google’s massive network, and a deep reach into the ISPs, we are able to deliver the highest quality video for our users, no matter where they are”—John Hogan, Chief Technology Officer, Stan

“Our mission at U-NEXT is to deliver the highest quality and most entertaining content to our users. Google Cloud’s Media CDN helps us efficiently scale our infrastructure, which is challenging with a vast library of content. Media CDN offloaded 98.3% of requests from our origin server while delivering consistent great quality.”—Rutong Li, Chief Technology Officer, U-NEXT

Broader platform for monetization and immersive experiences


While global distribution is critical for a high-quality end-user experience, it’s only one piece of delivering a world-class platform for immersive experiences. Media CDN offers additional capabilities to enable this transformation — ad insertion, ecosystem integrations and platform extensibility, and powerful AI/ML analytics for interactive experiences.

Streaming providers can improve monetization through integrated ad serving via the Video Stitcher API, which allows manipulation of video content to dynamically insert ads.

Through extensible ecosystem integrations, Media CDN connects customers to key capabilities to simplify their operations. For example, the Transcoder API supports custom streaming formats, while the Live Stream API transcodes mezzanine live signals into direct-to-consumer streaming formats, for multiple device platforms.

Media CDN is built with AI/ML that will give viewers more control over how they see, experience, and even interact with content. For example, sports fans watching a game can obtain real-time stats and analytics, viewers can purchase items from virtual billboards, etc.

Cloud-native and developer-friendly operations


Media companies are under pressure to develop and deploy innovative experiences at a furious pace. Media CDN was built by developers, for developers, with automation and observability built in, giving media providers the speed and flexibility they need to integrate delivery provisioning and management into their content release processes.

Media CDN offers comprehensive APIs and automation tools such as Terraform. Detailed, pre-aggregated metrics and playback tracing make it easy to diagnose performance across the entire infrastructure stack. Real-time visibility is provided via Google Cloud’s operations suite, and integrates with tools that developers already use such as Grafana and ElasticSearch.

Leveraging the same infrastructure as YouTube, Google Cloud’s Media CDN combines geographic reach, API-first architecture and integration with the Cloud operations suite. This is a transformative move that is aligned with the future of the CDN industry.”— Ghassan Abdo, Research Vice President, WW Telecom, Virtualization and CDN, IDC

Viewers around the world are demanding best-in-class video quality and performance across modes of consumption. A video-first delivery network can be a game changer in this space. We’re excited to partner with Google Cloud and to leverage Media CDN to enable premium video experiences and customer engagements.”—Juan Martin, Founder and CTO, Firstlight Media

Planet-scale advanced security


Media CDN lets streaming media providers take advantage of Google’s decades-long experience delivering video safely, securely, and reliably. The platform includes deep integration with Google Cloud Armor for planet-scale DDoS protection and a rich set of capabilities to detect and mitigate attacks, prevent abuse, manage risk, and comply with regulatory or licensing requirements.

If you want to deliver rich, immersive experiences to global audiences with an extensible, modern delivery platform, we’d love to hear from you. For more information, including technical specifications and platform architecture, please visit cloud.google.com/media-cdn. To get started with Media CDN, contact your sales team.

More Relevant Stories for Your Company

Blog

Two Ways to Deploy SAP HANA System on Google Cloud

Many of the world’s leading companies run on SAP—and deploying it on Google Cloud extends the benefits of SAP even further. Migrating your current SAP S/4HANA deployment to Google Cloud—whether it resides on your company’s on-premises servers or another cloud service—provides your organization with a flexible virtualized architecture that lets

Blog

Public Cloud’s Zero Trust Architecture Keeps Enterprise Data Safe

Over the past decade, cybersecurity has posed an increasing risk for organizations. In fact, cyber incidents topped the recent Allianz Risk Barometer for only the second time in the survey’s history. The challenges in combating these risks only continue to grow. Adversaries tend to be agile and are consistently looking

Case Study

The Inside Story of How Home Depot Migrated to Google BigQuery From an On-prem DW Solution

In the media, you will often hear story of how born-in-the-cloud companies manage with massive infrastructure. But it is one thing is to be a startup, and build infrastructure with bespoke requirements. And quite another to have a complex, multinational organization with online, with mobile, with brick-and-mortar presence, and hundreds

Case Study

Google Cloud’s ML-based Image Classification App: A Key to Global Wildlife Conservation

Wildlife provides critical benefits to support nature and people. Unfortunately, wildlife is slowly but surely disappearing from our planet and we lack reliable and up-to-date information to understand and prevent this loss. By harnessing the power of technology and science, we can unite millions of photos from [motion sensored cameras]

SHOW MORE STORIES