How AI and ML Helps Interpret Baseball Fandom during this MLB Season

4560
Of your peers have already read this article.
1:30 Minutes
The most insightful time you'll spend today!
The game of baseball has no shortage of statistics — from batting average to exit velocity, strikeouts to wins above replacement. Among all sports, Major League Baseball (MLB) arguably contains the most analytical and data-driven participants and fan base. Subconsciously or viscerally, players and managers on the field and those following from anywhere are constantly assessing and making decisions based off of game play trends and expectations — whether a batter will come through with a hit in an important situation, when a pitcher should be pulled. Less analyzed, however, is what leads fans to become engaged with certain players or teams, and what factors drive their love of the game. This is the motivation behind the problem being posed by Major League Baseball in their Kaggle competition for Player Digital Engagement Forecasting. Can you use machine learning to deconstruct baseball fandom?
This competition asks you to predict measures of digital engagement for each active player on a daily basis during the MLB season. So, how large was the surge in fan interest after Joe Musgrove threw the first no-hitter in Padres history? Is Shohei Ohtani’s engagement higher when he pitches well, when he hits a monster home run…or when he does both? You’re provided a wealth of game, team and player information – detailed stats, awards, rosters, and transaction information – as well as social and digital engagement data as your inputs. Data scientists will recognize this as an exciting forecasting problem with both traditional regression and time series components, where having this input data just prior to the prediction date is critical to determining which players will receive the most engagement.
With so many variables in the game, there are an endless number of vectors which could possibly influence fan engagement. Eleven-time All-Star Miguel Cabrera delighted fans by hitting the first home run of the season – in the snow! Occasionally a lesser-known player like Musgrove or Carlos Rodón “wins the day” with an unlikely no-hitter. And sometimes just getting traded to an iconic franchise like the Yankees generates a ton of fan interest, like it did for Rougned Odor in early April.

As these examples show, a player’s digital engagement can be pretty dynamic during the season, with many different potential contributors to who is “trending” on a given day. How can you use data to uncover which factors are the most influential of engagement with each player’s digital content?
Ready to play ball? Check out the competition on Kaggle for all the details. $50,000 in prizes is up for grabs in two prize categories. The code competition puts your machine learning skills to the test, to see who can build the most accurate forecasting models to predict daily digital engagement for every active player. You’ll have until July 31st to build your models and then be evaluated on a future time frame, which will determine the winners. For data visualization and exploration experts out there, the explainability prizes give you an opportunity to analyze more broadly which factors, even those outside of what we’re providing directly, most influence digital engagement. You’ll be evaluated on how well you can use what the data is telling you to support your findings.
And if you’re looking to get started, we’ve provided an introductory video and some notebook tutorials, including a starting point for harnessing the power of Vertex AI through tools including Cloud Notebooks, Explainable AI, and Vizier.
With the second half of the season upon us, it’s an exciting time to be an MLB fan. With this Kaggle competition, it’s also a perfect opportunity to use data science to help understand baseball fandom and potentially earn some of your own accolades in the process. Step up to the plate!
Major League Baseball trademarks and copyrights are used with permission of Major League Baseball. Visit MLB.com.
3311
Of your peers have already watched this video.
21:26 Minutes
The most insightful time you'll spend today!
What are TPUs and Why Should Data Scientists Care?
When it comes to machine learning, more data means better results.
But processing more data also requires more computing power. CPUs are great for sequential arithmetic calculations. They offer low latency with this type of workload. But to speed up machine learning, models have to perform multiple calculations in parallel. And that’s where GPUs perform better.
TensorFlow Processing Units, or TPUs, represent the most advanced hardware accelerators that has been architected ground up for TensorFlow by Google Cloud. Cloud TPUs deliver accelerated performance to help businesses train deep learning models in a matter of hours instead of weeks.
This allows enterprises to be more productive with their scarcest resource, the ML scientist, and drive more innovations, thanks to the advanced computing capability.
Watch John Barrus, Senior Product Manager, Google Cloud; and Zak Stone, Product Manager, Google Brain, show you benefits of using TPUs and how to get started.
3026
Of your peers have already watched this video.
16:30 Minutes
The most insightful time you'll spend today!
Predicting Treasury Settlement Failures with ML
BNY Mellon’s Government Securities Services (GSS) business is the sole provider of treasury settlement services in the United States of America. Given its unique market position, GSS is exploring how to help clients improve their forecasting of $70+ billion in daily settlement fails leveraging Google Cloud.
Sarthak Pattanaik, Chief Information Officer, Clearance and Collateral Technology, The Bank of New York Mellon and Victor O’Laughlen, Digital Business Leader, Clearance and Collateral, The Bank of New York Mellon, share how they utilized Google Cloud AI solutions to predict treasury settlement failures.
They take us through the business process, the steps they took to set up their AI solution, and what they have learnt on their journey—not just from a technical standpoint but from a cultural one as well.
Maximize Efficiency in Document Extraction Models with Document AI Workbench GA Release

1310
Of your peers have already read this article.
3:30 Minutes
The most insightful time you'll spend today!
Each day, more documents are created and used across companies to make decisions. However, the value in these documents is primarily expressed as unstructured data, which makes the value difficult and manually intensive to extract and use for business processes.
As the number and variety of documents used by businesses proliferate, machine learning (ML) solutions need to be more flexible to handle the broader set of use cases. That’s why we introduced the first Document AI Workbench model, Custom Document Extractor (CDE), in Public Preview at Google Cloud Next ‘22. CDE makes it fast and easy to apply ML to virtually any document-based workflow to extract structured data from unstructured document types, to automate business processes.
CDE lets developers and analysts use their own data to train models and extract fields from documents needed for the business. CDE lets organizations build models faster and with less data — thus accelerating time-to-value for processing and analysis of data in documents.
Today, we announce that Document AI Workbench is Generally Available (GA), open to all customers, ready for production use through APIs and the Google Cloud Console. Document AI Workbench is covered by the Document AI SLA — online and batch document prediction is supported with >=99.9% uptime. Furthermore, Document AI Workbench is now covered by Google Cloud’s GA product terms. For example, Google will notify customers at least 12 months before significantly modifying a customer-facing Google API in a backwards-incompatible manner.
In this blog post, we’ll explore ways customers are already using CDE and Document AI Workbench’s updated capabilities.
What users are saying about Workbench
Deliver higher model accuracy with Workbench
Users leverage Workbench to ultimately save time and money. A third party evaluated Document AI Workbench and concluded that it extracts data more accurately1 than several competing products for document types with variable layouts (e.g. invoice, receipt, bank statements, paystubs). Better accuracy drives higher automation rates, helping Workbench users save time and money.
Chris Jangareddy, managing director for Artificial Intelligence & Data at Deloitte Consulting LLP said, “Google Cloud Document AI is a leading document processing solution packed with rich features like multi-step classify and text extraction to automate sorting, classification, extraction, and quality assurance. By combining Document AI with Workbench, Google Cloud has created a forward-thinking and powerful AI platform for intelligent document processing that will allow for process transformation at an enterprise scale with predictable outcomes that can benefit businesses.”
Mansoor Khan, CEO of OneClinic said, “We help medical professionals scale their clinics through automation. We used Google’s Document AI Workbench to create a model to automatically extract data from patients’ insurance cards as part of our patient check-in software. Workbench is easy to use and we are really happy with the model accuracy — it extracts data more accurately than what we would expect from human data entry.”
Rajnish Palande, VP, Google Business Unit for BFSI, TCS said, “The Google Cloud Document AI Workbench leverages artificial intelligence (AI) to manage and glean insights from unstructured data. The Workbench brings together the power of classification, auto-annotation, page-number identification and multi-language support to help organizations rapidly deliver enhanced accuracy, improved operational efficiency, higher confidence in the information extract, and increased return on investment.”
Build production ready models faster with Workbench
Document AI Workbench helps users create machine learning models faster. For example, a third-party evaluation shows Document AI Workbench trains machine learning models up to 3x faster than a leading competitor. This is an important improvement which lowers total cost of ownership and increases value.
Dallas Dolen, Partner, Google Alliance Leader at PwC said, “Google Document AI Workbench helps to accelerate our custom parser models training as well as improves accuracy and performance using a custom document extractor with human in the loop. It helps us solve complex business problems for our clients in the financial services and healthcare industries.”
Ziang Jia, Senior DocAl Development Lead at Resultant, said, “Document AI Workbench has unlocked a brand-new machine learning development experience for information extraction solutions. Its simplicity and robustness enabled us to build models and deliver a highly accurate outcome in an agile way for a large government agency. We couldn’t be more impressed by its simplicity and robustness and are excited to see how the product will evolve in the future.”
Sean Earley, VP of Delivery Services of Zencore said, “Document AI Workbench allows us to develop highly accurate document parsing models in a matter of days. Our customers have automated tasks that formerly required significant human labor. For example, using Document AI Workbench, a team of two trained a model to split, classify and extract data from 15 document types to automate Home Mortgage Disclosure Act reporting. The mean trained model accuracy was 94%, drastically reducing the operational cost of our customer’s compliance reporting procedures.”
What’s new with Document AI Workbench
The latest Workbench capabilities make it even easier to train and deploy an extraction model:
- With Workbench’s public APIs, you can programmatically create, delete, train, evaluate and deploy models.
- Our updated dataset management tools automatically detect and create existing schema labels from your pre-annotated documents. They also provide you more flexibility when creating and managing schema.
- Our new DocAI Toolkit includes a labeled document converter so that you can easily convert your labeled documents to DocAI’s format and start training faster.
- We’ve reduced the cognitive load for labelers with efficiency enhancements to our Labeling UI.
- The revamped Processor Gallery helps you quickly identify the best model for your use case.
What’s next for Document AI Workbench
We continue to invest in Document AI Workbench to help you automate document processing. Here are a few things we’re working on that we’re excited about:
- Classify document types with the Custom Document Classifier (CDC), coming soon in public preview
- Copy processor versions across projects and processors to streamline managing development and production environments
- Support larger documents (e.g., longer than 50 pages) so you can process a wider array of documents
- Broader (non-latin) language support–equivalent to Document AI OCR
- And many more investments, using state of the art technology, to help you build world class models faster to automate document processing
Document AI Workbench is in GA and ready for production workloads. Learn more via Document AI Workbench documentation or try it out in the Google Cloud Console.
Acknowledgements: Tomas Moreno, Outbound Product Manager, Lukas Rutishauser, Software Engineering Manager, Michael Kwong, Software Engineering Manager, Rajagopal Janani, Software Engineering Manager, Michael Lanning, UX Designer.
1. When trained with 200+ documents
4952
Of your peers have already watched this video.
21:10 Minutes
The most insightful time you'll spend today!
Learn Modern App Development Practices to Ship Software Faster
Cloud-native, Kubernetes, Serverless have been the hottest and most widely discussed topics given the velocity and agility benefits.
Learn more about how you can leverage these modern app development practices to ship software faster, while reducing costs and improving security and compliance.
Learn how Google Cloud lets you modernize existing applications at your own pace using these technologies. Regardless of where you are in your app modernization journey, watch this video to learn how to improve the developer experience and deliver software faster.
3101
Of your peers have already watched this video.
21:00 Minutes
The most insightful time you'll spend today!
Google Cloud’s AI Adoption Framework: Helping You Build a Transformative AI Capability
AI can help organizations improve the decision-making process across most business functions. However, building an effective AI capability encompasses more than just creating a technology platform.
To do this effectively requires alignment to business objectives, strong executive sponsorship, and collaboration between skilled employees and strategic partners.
Additionally, you need your initiatives to be powered by secure data management and cloud-native services to scale and automate ML workloads, and ensure all of this is underpinned by responsible AI principles.
Successfully adopting AI in your business is determined by your practices in these areas. Learn more about the AI journey and how you can gain value every step of the way.
More Relevant Stories for Your Company

Case Study: Twitter is Taking Their CX to The Next Level with AutoML
Editor’s note: Since launching its Spaces feature, Twitter has demonstrated that hearing people’s voices can bring conversations on Twitter to life in a completely new way. Next, it aimed to make it easier for customers to join and listen to live conversations they personally care about. In this blog, we

Ahead of the Curve: 5 Data and AI Trends Set to Shape 2023
How will your organization manage this year's data growth and business requirements? Your actions and strategies involving data and AI will improve or undermine your organization's competitiveness in the months and years to come. Our teams at Google Cloud have an eye on the future as we evolve our strategies
Cloud as an Innovation Platform in Capital Markets
Public cloud, big data, and AI technologies offer competitive advantages and cost savings for capital markets firms ready to make the transition. This paper discusses the three phases capital markets firms go through in transitioning to public cloud, and the workloads, benefits, and cultural changes that characterize the three phases:

Innovate Faster & More Flexibly: How Our Commitment to Open Source Unlocks AI and ML Innovation
At Google, we believe anyone should be able to quickly and easily turn their artificial intelligence (AI) idea into reality. Open source software (OSS) has become increasingly important to this goal, heavily influencing the pace of innovation in AI and machine learning (ML) ecosystems. Over the last two decades, ML







