WayFair and Google Cloud Get Together to Raise the Bar on World-class Experience!

3013
Of your peers have already read this article.
2:00 Minutes
The most insightful time you'll spend today!
On Dec 9th and 10th, Wayfair and Google Cloud came together for the inaugural Wayfair-Google Cloud Machine Learning Hackathon. Wayfair firmly believes that hackathons are a great way to fuel a culture of collaboration and experimentation. To fuel Wayfair’s incredible pace of innovation at scale, it’s team of more than 3,000 technologists is constantly experimenting and taking smart risks. It’s test-and-learn culture empowers everyone to think critically and creatively, and take big swings to raise the bar on its world-class experience.
The Wayfair-Google Cloud Machine Learning Hackathon was all about getting Wayfairians excited about and enabled on new technology. More specifically, this event was a contained test environment for Wayfair innovators to validate AI and ML tools in order to understand how to get better insight from their data. The projects worked on during this hackathon will help Wayfair build new use cases that could impact the business and the end customer in new and productive ways. Google Cloud’s focus for this Hackathon was to enable Wayfairains to harness the power of Machine Learning and AI to enable their own goals around continual improvement and relentless customer focus.
Prior to the event, Wayfair Data Scientists and Machine Learning Engineers who signed up and submitted ideas for the Hackathon were invited to optional Google Cloud enablement and training sessions. Google Cloud set up classrooms in Qwiklabs on topics including, but not limited to, BigQuery, Vertex AI, and Natural Language AI, so that all Hackathon participants could try out new technologies and tools in a learning environment. Google Cloud subject matter experts were available to field any questions Wayfairians had about the use cases they were hacking on.
The Hackathon was a hybrid virtual & physical event that hosted 67+ innovators, 49 of which registered for Google Cloud supported ideas. 15 judges from both Google Cloud and Wayfair oversaw the event. The esteemed list of judges included Steven Conine, Wayfair’s Co-Founder and Co-Chairman who has helped pave the way for the development of practical applications of next-generation technologies like augmented reality. The judges measured and evaluated the success of a project based on the following criteria:
- “Wow Factor”: How innovative is the project?
- Impact: How impactful is the project to Wayfair?
- Polish: How complete is the project?
- Presentation Quality: How clear and consistent is the demo?
The theme of the Hackathon was Machine Learning and AI. Teams were able to collaborate with participants globally, either in-person or virtually, and work together on projects in five categories: Relentless Customer Focus, Always Improving, Google’s Choice, People’s Choice, and Hackers’ choice.
At the end of the two days there were winners in all 5 categories. The Google’s Choice award went to “Entity Extraction for Order Matching”. Team “Project Clippy” was named both Hackers’ Choice and the winners of the Relentless Customer Focus category. See below for the results of the 5 categories:
- Relentless Customer Focus and Hackers’ Choice
- Winning Team: Project Clippy
- Hackers: Misha Balyasin, Alex Saad, Leo Smerling, and Gabriele Lanaro
- This project provided a gamified experience to make the process of leaving a product review even more seamless.
- Always Improving
- Winning Team: Customer Causal MetaLeaners
- Hackers: Colin Gray, Irene Wang, Huy Vo Tran, Wenhao Xu, and Santiago Velez Ferro
- Team Customer Causal MetaLeaners worked to build lightweight procedure(s) for computationally distributed, multi-target, customized loss functions to make causal meta-learners more applicable to real-world Wayfair problems.
- Google’s Choice Award:
- Winning Team: Entity Extraction for Order Matching
- Hackers: Roger Bock, Bradley West, Sina Moeini, and Jonathan de Melker Worms
- This team built out a solution to use text models to extract and identify the products that customers purchased from user reviews.
- People’s Choice:
- Winning Team: KNN and ANN on Vertex AI
- Hackers: Santosh Jhingade, Ashrith Marpaka, Nikhil Bhaip, Adam Schulze, and Brandon Sanders
- This team used Vertex AI to expand the impact they can have on suppliers and customers by providing accurate and real-time information of products that match either description or image.
Wayfair leaders reflected on the two days and shared their input. Matt Ferrari, Head of Ad Tech, Customer Intelligence, and Machine Learning; Engineering and Product at Wayfair said, “Wayfair has a lot of vendors, but very few strategic partners, and Google is that. Our Partner.” “Thank you all for the participation! I’m grateful to Google, and the many others for helping lead a successful event.”
Wayfair partners with Google Cloud to optimize performance and resiliency, support scaling data-driven decisions, and Increase employee productivity. “Our category is ripe for innovation, and our partnership with Google Cloud helps us ensure that great ideas can come from anywhere by empowering our technologists with cutting-edge products and solutions,” said Ferrari. “We’re proud to partner on efforts like hackathons that align with our team’s eagerness to work on complex, rewarding problems that push the envelope and challenge us to always think big.”
At first Google Cloud won Wayfair over with the speed, reliability and performance of their technology. Hackathons like this one exemplify that what is equally as important is Google Cloud’s willingness to work side-by-side with Wayfair at every level to enable a culture of innovation. Learn more about the Wayfair-Google Cloud partnership here.
Google Migration and BigQuery Brings PedidosYa Closer towards its Goal of Becoming Data-driven

7245
Of your peers have already read this article.
2:00 Minutes
The most insightful time you'll spend today!
Editor’s note: PedidosYa is the market leader for online food ordering in Latin America, serving 15 markets and over 400 cities. It’s also one of the largest brands within the German multinational company Delivery Hero SE. With over 20 million app downloads, PedidosYa provides the best online delivery experience through 71,000+ online partners, including restaurants, shops, drugstores, and specialized markets.
Having constant access to fresh customer data is a key requirement for PedidosYa to improve and innovate our customer’s experience. Our internal stakeholders also require faster insights to drive agile business decisions. Back in early 2020, PedidosYa’s leadership tasked the data team to make the impossible possible. Our team’s mission was to democratize data by providing universal and secure access while creating a comprehensive information ecosystem across PedidosYa. We also had to achieve this goal while keeping costs under control— even during the migration stage and removing operational bottlenecks.
Challenges with legacy cloud infrastructure
PedidosYa first built its data platform on top of AWS. Our data warehouse ran on Redshift, and our data lake was in S3. We used Presto and Hue as the user interfaces for our data analysts. However, maintaining this infrastructure was a daunting task. Our legacy platform couldn’t keep up with the increasing analytics demands. For example, the data stored on S3 complemented by Presto/Hue required high operational overhead. This was because Presto and our IAM (identity access management) didn’t integrate well in our legacy ecosystem. Managing individual users and mapping IAM roles with groups and Kerberos was operationally time-consuming and costly. Further, sharding access on the S3 files was far too complicated to enable seamless ACLs (access control lists).
There were also challenges with workload management. Our data warehouse had batch data loaded overnight. If one analyst scheduled a query to run during the overnight ETL (extract, transform, load) workload, it would disrupt the current ETL task. This could stop the entire data pipeline. We’d have to wait until data engineers intervened with a manual fix.
It was also difficult to understand whether a query error was due to performance issues or platform resource exhaustion. This lack of clarity affected our data analysts’ ability to autonomously improve querying efficiency. Data team members needed to manually inspect personal queries looking for performance issues. Also, the current architecture was prone to a ‘tragedy of the commons’ situation; it was seen as an unlimited and free resource. As a result, it was impossible to disentangle the infrastructure from different stakeholder teams, as all had very different needs.
The decision to modernize our data warehouse
Given the growing challenges from our legacy platform, our tech team decided to transform our analytics environment with a modern data warehouse. They required the following key criteria from their next data platform:
- Scalability – The ability to grow with elastic infrastructure.
- Cost control – Cost management and transparency. These factors promote efficiency and ownership—both key aspects of data democratization.
- Metadata management – Intuitive data platform focusing on users’ previous SQL knowledge. Plus, being able to enrich the informational ecosystem with metadata, to diminish data gatekeepers.
- Ease of management – The team needed to reduce operational costs with a serverless solution. Data engineers wanted to focus on their key roles rather than acting as database administrators and infrastructure engineers. The team also wanted much higher availability, and to reduce the impact of maintenance windows and vacuum/analysis.
- Data governance and access rights – With a growing employee base with varying data access requirements, the team needed a simple yet comprehensive solution to understand and track user access to data.
Migrating to Google Cloud
After exploring other alternatives, we concluded Google Cloud had an answer to each of our decision drivers. Google Cloud’s serverless, managed, and integrated data platform, coupled with its seamless integration across open-source solutions, was the perfect answer for our organization. In particular, the natural integration with Airflow as a job orchestrator and Kubernetes for flexible on-demand infrastructure was key.
We used Dataflow together with Pub/Sub and Cloud Functions for our data ingestion requirements, which has made our deployment process with Terraform seamless. Because we set up everything in our environment programmatically, operation time has diminished. Google Cloud reduced the deployment process from about 16 hours in our legacy platform to 4 hours. This is partly due to the friendliness of automating the deployment (such as schema check, load test, table creation, build.) process with Terraform, Cloud Functions, Pub/Sub, Dataflow, and BigQuery on GCP. Input messages processed with Dataflow allow us to abstract and plan the schema changes according to the needs of the functional team. For example, schema changes raise an alarm, and then we can modify the raw layer table schema. By doing this, we ensure that backend modifications that we don’t control do not affect upper layers.
A key reason why we picked Google Cloud was because of its advanced cost and workload management coupled with its transparent log analytics. This information gives us a complete view into any query performance issues to make improvements on the fly. Further, we achieved a significant amount of cost savings by consolidating multiple tools to BigQuery.With BigQuery, we’ve been able to reduce our total cost per query by 5x.
This was due to a number of reasons:
- Automating pipeline deployment made it much simpler to maintain the data processing processes.
- Analysts are conscious about what queries they’re running, resulting in running better, more optimized queries.
- Analysts use a Data Studio dashboard to see their queries and all the associated costs. As a result, there’s a lot more transparency for each persona.
With these changes, we can easily manage and assign costs associated with each workload with their own cost centers using specific Google Cloud projects.
Change management is always challenging. However, BigQuery is intuitive and doesn’t have a steep learning curve from Hue/Hive on SQL basics. BigQuery also allowed the team to expand its capabilities and enabled them to properly work with nested structures, avoiding unnecessary joins and improving query efficiency. Additionally, we now use Data Catalog as our unique point of truth for metadata management. This allows our team to break the data access barriers and enable federation of data across the organization. By using Airflow to orchestrate everything, we keep track of every data stream. With this information, each end user can see their regularly used data entities’ status via the dashboard. This also adds transparency to our everyday data processes.
Finally, with Google Cloud’s IAM rules applied across the different products, data sharing and access is close to a noOps experience. We have programmatically implemented access according to roles and level access within the company. This allows certain pre-validated roles to view more sensitive information. These solutions help drive a more automated data governance experience.
Up next: Google Cloud AI/ML
The new stack based on BigQuery has created significant productivity gains. Freed from the burden of operational management, PedidosYa’s data team can now focus on adding value through data tools and products.
- Our data engineers are better equipped to integrate constantly changing transactional and operational data.
- The dataOps team can automate the infrastructure and provide autonomy to the end user.
- Our data quality team can focus on bringing added value to data stakeholders.
- Data scientists and data analytics can spend more time analyzing data and less time asking data gatekeepers for data access.
PedidosYa can now democratize data access with a well-governed architecture. We are still at the beginning of our journey, but we are closer to achieving our vision of building a data-driven organization. Up next: expanding our artificial intelligence and machine learning capabilities.
Tune in to Google Cloud’s Applied ML Summit on June 10th, 2021, or listen on-demand later, to learn how to apply groundbreaking machine learning technology in your projects.
Say Goodbye to Manual W2 & Payslip Processing with Document AI

2503
Of your peers have already read this article.
2:00 Minutes
The most insightful time you'll spend today!
Documents like payslips and W2s are crucial to processes such as employment and income verification for mortgage loans, personal loans, personal finance, and benefits processing. Unfortunately, efficiently extracting data from these documents at scale can be challenging and time-consuming, with many organizations relying on manual examination of documents or automated approaches that don’t adequately capture the document data needed for given tasks. Google Cloud built Document AI to remove these barriers, empowering customers to deploy powerful machine learning models to more quickly process documents, save money, and discover insights. We’re excited to expand Document AI’s capabilities with the recent release of improved pre-trained models for W2s and payslips, built on Document AI Workbench.
Pre-trained models let developers focus on core application logic and leave the complex task of information extraction from the documents to Google’s AI technology. In many cases, the primary driver for automated data extraction is operational efficiency and cost savings, but Document AI can also open new possibilities. For example, a financial services company might use Document AI to enable fully self-serve loan applications on mobile devices, helping the organization to differentiate itself with simple, fast customer experiences.
We’ve heard from customers that more granular entity extraction from W2 and payslip documents is particularly important, with organizations requiring support for a wider variety of layouts and formats. The recent launch of the stable release of these pretrained models addresses these requests.
Here is what is new with W2 parser:
- The parser improves accuracy and entity specificity thanks to the ability to break down long entities such as addresses into fine-grained sub-entities like StreetAddressOrPostalBox, AdditionalStreetAddressOrPostalBox, City, State, and ZIP code.
- It can handle a wider variation of W2 forms, including multi-copies (2,3,4-ups) issued by various payroll vendors. The model is not limited to specific tax years, which means it should be able to process W2 for 2022 or beyond provided there are not significant changes to the format.
- It introduces eight new entities for Box 12 that represent both codes and values, enriching understanding of the various taxable and non-taxable components of the W2 recipient’s income.
Here is what is new with Payslip parser:
- Bonus, commissions, holiday, overtime, regular pay, and vacation are now part of earning_item/earning_this_period and earning_item/earning_ytd. The parser captures types of earnings beyond those categories, and maps them to their respective earning rates, hours, and pay (both for the period and year-to-date). This helps in building a more detailed understanding of the components of the payslip recipient’s income
- The parser now returns year-to-date and current-period taxes and deductions.
- Direct deposits are linked to corresponding bank account numbers.
- The parser now returns page numbers, state and federal tax exemptions, and filing statuses.
While these parsers have become more useful out of the box, with this release, the ability to uptrain makes them easy to modify as new needs arise. Uptraining lets developers further improve the accuracy of these models and extract additional fields with minimal development work. It also lets developers customize existing parsers to support new document types that are similar. For example, the parser is trained on U.S. data and could be uptrained to create a payslip parser for the U.K.
We’re pleased that parsers are already making a difference for customers. Bryan Jackson, CTO at lending automation firm Gateless, said, “High accuracy data extraction is critical to the success of our Smart Underwrite solution, and Document AI provided better results than competitors. Using the latest W2 & Payslip pretrained parsers, we saw a 48% increase in performance on pay stubs and a 15% performance improvement in W2s. The ability to easily uptrain models as new document variations are introduced ensures we continue to deliver optimal outcomes for our customers.”
Additional pre-trained models available as release candidates include parsers for 1040, 1099R, 1120, and 1120S documents. Check for details here. To learn more, talk to a Google Cloud sales executive about how Document AI can help your business, and check out our Document AI breakout session from Google Cloud Next ’22.
WayFair and Google Cloud Get Together to Raise the Bar on World-class Experience!

3014
Of your peers have already read this article.
2:00 Minutes
The most insightful time you'll spend today!
On Dec 9th and 10th, Wayfair and Google Cloud came together for the inaugural Wayfair-Google Cloud Machine Learning Hackathon. Wayfair firmly believes that hackathons are a great way to fuel a culture of collaboration and experimentation. To fuel Wayfair’s incredible pace of innovation at scale, it’s team of more than 3,000 technologists is constantly experimenting and taking smart risks. It’s test-and-learn culture empowers everyone to think critically and creatively, and take big swings to raise the bar on its world-class experience.
The Wayfair-Google Cloud Machine Learning Hackathon was all about getting Wayfairians excited about and enabled on new technology. More specifically, this event was a contained test environment for Wayfair innovators to validate AI and ML tools in order to understand how to get better insight from their data. The projects worked on during this hackathon will help Wayfair build new use cases that could impact the business and the end customer in new and productive ways. Google Cloud’s focus for this Hackathon was to enable Wayfairains to harness the power of Machine Learning and AI to enable their own goals around continual improvement and relentless customer focus.
Prior to the event, Wayfair Data Scientists and Machine Learning Engineers who signed up and submitted ideas for the Hackathon were invited to optional Google Cloud enablement and training sessions. Google Cloud set up classrooms in Qwiklabs on topics including, but not limited to, BigQuery, Vertex AI, and Natural Language AI, so that all Hackathon participants could try out new technologies and tools in a learning environment. Google Cloud subject matter experts were available to field any questions Wayfairians had about the use cases they were hacking on.
The Hackathon was a hybrid virtual & physical event that hosted 67+ innovators, 49 of which registered for Google Cloud supported ideas. 15 judges from both Google Cloud and Wayfair oversaw the event. The esteemed list of judges included Steven Conine, Wayfair’s Co-Founder and Co-Chairman who has helped pave the way for the development of practical applications of next-generation technologies like augmented reality. The judges measured and evaluated the success of a project based on the following criteria:
- “Wow Factor”: How innovative is the project?
- Impact: How impactful is the project to Wayfair?
- Polish: How complete is the project?
- Presentation Quality: How clear and consistent is the demo?
The theme of the Hackathon was Machine Learning and AI. Teams were able to collaborate with participants globally, either in-person or virtually, and work together on projects in five categories: Relentless Customer Focus, Always Improving, Google’s Choice, People’s Choice, and Hackers’ choice.
At the end of the two days there were winners in all 5 categories. The Google’s Choice award went to “Entity Extraction for Order Matching”. Team “Project Clippy” was named both Hackers’ Choice and the winners of the Relentless Customer Focus category. See below for the results of the 5 categories:
- Relentless Customer Focus and Hackers’ Choice
- Winning Team: Project Clippy
- Hackers: Misha Balyasin, Alex Saad, Leo Smerling, and Gabriele Lanaro
- This project provided a gamified experience to make the process of leaving a product review even more seamless.
- Always Improving
- Winning Team: Customer Causal MetaLeaners
- Hackers: Colin Gray, Irene Wang, Huy Vo Tran, Wenhao Xu, and Santiago Velez Ferro
- Team Customer Causal MetaLeaners worked to build lightweight procedure(s) for computationally distributed, multi-target, customized loss functions to make causal meta-learners more applicable to real-world Wayfair problems.
- Google’s Choice Award:
- Winning Team: Entity Extraction for Order Matching
- Hackers: Roger Bock, Bradley West, Sina Moeini, and Jonathan de Melker Worms
- This team built out a solution to use text models to extract and identify the products that customers purchased from user reviews.
- People’s Choice:
- Winning Team: KNN and ANN on Vertex AI
- Hackers: Santosh Jhingade, Ashrith Marpaka, Nikhil Bhaip, Adam Schulze, and Brandon Sanders
- This team used Vertex AI to expand the impact they can have on suppliers and customers by providing accurate and real-time information of products that match either description or image.
Wayfair leaders reflected on the two days and shared their input. Matt Ferrari, Head of Ad Tech, Customer Intelligence, and Machine Learning; Engineering and Product at Wayfair said, “Wayfair has a lot of vendors, but very few strategic partners, and Google is that. Our Partner.” “Thank you all for the participation! I’m grateful to Google, and the many others for helping lead a successful event.”
Wayfair partners with Google Cloud to optimize performance and resiliency, support scaling data-driven decisions, and Increase employee productivity. “Our category is ripe for innovation, and our partnership with Google Cloud helps us ensure that great ideas can come from anywhere by empowering our technologists with cutting-edge products and solutions,” said Ferrari. “We’re proud to partner on efforts like hackathons that align with our team’s eagerness to work on complex, rewarding problems that push the envelope and challenge us to always think big.”
At first Google Cloud won Wayfair over with the speed, reliability and performance of their technology. Hackathons like this one exemplify that what is equally as important is Google Cloud’s willingness to work side-by-side with Wayfair at every level to enable a culture of innovation. Learn more about the Wayfair-Google Cloud partnership here.
3245
Of your peers have already watched this video.
38:09 Minutes
The most insightful time you'll spend today!
TensorFlow: The Show-and-Tell Data Scientists Have Been Asking For
TensorFlow is among the most popular, if not, the most popular deep learning libraries today. According to one ranking, “TensorFlow is at least two standard deviations above the mean on all calculated metrics.”
Watch as Lak Lakshmanan, Technical Lead, Machine Learning and Big Data, Google Cloud, walks through a development workflow that will make operationalization easier to execute, including the process of building a complete machine learning pipeline covering ingest, exploration, training, evaluation, deployment, and prediction.
He also talks about the need for distributed training. But what’s the benefit of distributed TensorFlow? Many machine learning frameworks can only handle “toy problems”, or problems that can be solved by input data that fits into memory. These are small data sets.
But to build effective machine learning you need big data, feature engineering, and model architectures. With large amounts of data batching and distribution are very important. That’s where distributed training comes in.

5565
Of your peers have already downloaded this article
6:30 Minutes
The most insightful time you'll spend today!
Across industries and use cases, organizations that have implemented ML report demonstrable return on investment and substantial business benefits ranging from better, faster data analysis to improved efficiency and cost savings.
The vast majority of early adopters — nearly 90 percent, according to one study — believe that ML provides a competitive advantage, and more than half of business leaders who participated in another survey expect that ML will determine their companies’ future success. It’s also worth noting that most early adopters say that ML enhances their cybersecurity efforts.
We’ve experienced this effect firsthand here at Google Cloud, where
we use AI-powered methods to identify vulnerabilities and
thwart attacks.
What Are the Top Uses of ML?

And more importantly what are the top use cases in your industry?

Find out more, including the return-on-investment enterprises are witnessing, and the tricks many early adopters are employing to fast-track their adoption of machine learning. Download this survey report now!
More Relevant Stories for Your Company

Increasing Production Efficiency: How AI Can Improve Asset Utilization and Minimize Downtime
Today, manufacturers are advancing on their factory digitalization journey, betting on innovative technologies to strengthen competitiveness, deliver sustainable growth, and offer new services. Macroeconomic factors - such as high energy costs, increasing labor, and raw material shortages - drive the need for urgent operational optimizations and automation. Cloud capabilities have

Explore the Innovations and Architecture Powering Spanner and BigQuery
Previously, databases had architectures with tightly coupled storage and compute. This resulted in higher latency, and with faster networks these constraints no longer surface. With Google Cloud's BigQuery and CloudSpanner, the storage and compute architecture have been separated, allowing for better scalability and availability to address businesses' high throughput data

Scaling Machine Learning Operations with Vertex AI AutoML and Pipeline
When you build a Machine Learning (ML) product, consider at least two MLOps scenarios. First, the model is replaceable, as breakthrough algorithms are introduced in academia or industry. Second, the model itself has to evolve with the data in the changing world. We can handle both scenarios with the services

Google and AI Researchers Work towards Building Data-centric AI
AI researchers and engineers need better data to enable better AI solutions. The quality of an AI solution is determined by both the learning algorithm (such as a deep-neural network model) and the datasets used to train and evaluate that algorithm. Historically, AI research has focused much more on algorithms






