How Real IT Leaders Create a Machine Learning Strategy - Build What's Next

Hi There, Thank you for downloading the whitepaper

Whitepaper

How Real IT Leaders Create a Machine Learning Strategy

READ FULL INTRODOWNLOAD AGAIN

5136

Of your peers have already downloaded this article

7:30 Minutes

The most insightful time you'll spend today!

Blog

Document AI: A Platform for Businesses to Simplify Document Automation

2366

Of your peers have already read this article.

2:30 Minutes

The most insightful time you'll spend today!

Finding the right document in a haystack of thousands across the organization can be challenging. To solve this problem, Team Google launched Document AI, an AI agent that lets organizations simplify and automate documents. Read more...

As I type this blog in a Google Doc, I can’t help but think about how much we rely on digital documents to communicate, collaborate, and do business. Yet most of the data in documents remains un-analyzed. Even when documents are part of customer-facing workflows like processing a mortgage application or a contract, businesses frequently struggle with the data those documents contain. Often, even finding the right document in a haystack of thousands across the organization can be challenging. To resolve these difficulties, most organizations rely on manual, time-consuming, and resource-intensive processes—all of which are incredibly frustrating for employees at the forefront of document workflows.

Because of challenges like these, in 2020, we launched Document AI, an AI agent that lets organizations apply machine learning (ML) to their hardest document automation problems. Since then, we have introduced specialized models to extract data for industry-specific use cases such as mortgage processing and procurement. With the launch of Document AI Workbench and Document AI Warehouse at Google Cloud Next ‘22, we’ve continued to take significant steps in our mission to help organizations simplify and automate document processing. Let’s double click on each of these announcements.

Custom document processing with Document AI Workbench

With Document AI Workbench, organizations can process documents by creating custom ML models that are specific to their business needs and extract unstructured data with a high degree of accuracy. Thanks to the user-friendly interface, even business users who do not have extensive ML skills can get started training or uptraining models.

Moreover, if an organization wants to transfer learning from pretrained models and enhance a model further to, say, include new fields, users can now do so by what we call “uptraining.” The uptraining feature is especially valuable for the most common yet complex use cases because it helps to save time and resources, so businesses don’t have to start from scratch. Uptraining for the invoice, purchase order (PO), contracts, W2, 1099-R, payslip, and 1040 pre-trained models unlocks new possibilities for improving accuracy, adding new language support, and schema customization.

We’re continuing to invest in these pretrained models. At Next’22, we announced an update to our invoice and expense pre-trained models with improvements to normalization and line item entities detection, as well as new ID proofing capabilities via a flexible API designed to spot fake, altered, or doctored ID documents. We’ve also added support for five new languages across invoice and expense models, in addition to the 12 previously-supported languages, and expanded availability in Canada and Australia regions, in addition to previously-supported US, EU, and Singapore regions.

According to Daan De Groodt, Managing Director, Deloitte Consulting LLP, Document AI Workbench “is poised to be a game changer, because we can now uptrain various text documents and forms utilizing powerful Google Machine Learning models to get the desired accuracy creating greater time and resource efficiencies for our clients.”

And customers are already seeing benefits. Libeo used Document AI to uptrain an invoice parser with 1,600 documents and increase its testing accuracy from 75.6% to 83.9%. “Thanks to uptraining, the Document AI results now beat the results of a competitor and will help Libeo save ~20% on the overall cost for model training over the long run,” said Libeo chief technology officer, Pierre-Antoine Glandier.

Google-powered document search with Document AI Warehouse

With Document AI Warehouse we are bringing the best of Google’s semantic search to documents. Document AI Warehouse lets enterprises search, store, govern and manage documents and their AI-extracted data and metadata in a single platform. With Document AI Warehouse’s simple and intuitive web accessible user interface, users can explore, view, bulk update and organize documents into folders. Document AI Warehouse offers robust enterprise control and governance so you can control who has access at the document and folder levels and assign users and groups permissions to view, edit, manage (share, delete) documents. You can migrate, sync, or federate documents from other repositories, such as Microsoft SharePoint, Amazon S3, and IBM FileNet. Or if that’s not an option we simply index the content and any extracted/tagged metadata).

We also will consolidate a number of next-generation product enhancements on Document AI OCR and Form Parser by the end of this year – including deeper insights into document quality & semantics, a unified document OCR experience, expanded language coverage for Form Parser, and advanced tooling for model lifecycle management. Google’s DeepMind team developed a new method that allows the creation of document parsing ML models for utility bills and purchase orders with 50%-70% less training data than what was previously needed for Document AI. We’re working on integrating this method into Document AI Workbench in the coming months.

Getting started

I’m very excited about what the future holds for Document AI as a platform for businesses to simplify document automation. Learn more about all these exciting developments in my session at Next’22 or try out one of our offerings today.

5053

Of your peers have already watched this video.

2:50 Minutes

The most insightful time you'll spend today!

Case Study

FedEx Ground Makes Talent Recruitment More Effective with AI

FedEx Ground is a package shipping company and is a subsidiary of FedEx. It wanted to make hiring easier, and more intuitive so that it could hire the best people.

“We need to have every advantage we can to recruit and retain talent. That’s what led us to the work with Google and its capabilities,” says Matt Tokorcheck, VP, Operations, Support and Engineering, FedEx Ground

The challenge was the narrow slotting of job roles. The openings were listed under specific headings which revolved around job types or departments–and if applicants didn’t fit or understand those categories, they didn’t apply.

Take, for example, applicants that came from the military. “Many of my fellow service members and veterans expressed difficulty in finding a job post the military because a lot of the skill sets that they’ve developed and honed over their military career aren’t as useful in the civilian world,” says David Henderson, Industrial Engineer, FedEx.

So FedEx Ground decided to work with Google Cloud’s AI-powered talent solution.

“As a job seeker when you come to our career site to search for jobs, that search is powered by Jibe and the Google Jobs API. And it really matches the keywords that a job seeker inputs with the jobs that are available at FedEx Ground, says Shailesh Bokil, MD, Talent Acquisition and Planning, Fedx Ground.

This makes job hunting a very intuitive experience for applicants.

“When I type into the search bar, I was immediately prompted to input my MOS, which is your military occupational specialty. And what it (the system) does is it takes the skills that are developed while serving in that MOS0 and matches them with skill sets that employers are looking. When I input 12A (an MOS), immediately I was getting results back for various engineer positions.

To find out more about how FedEx Ground employs AI-powered talent solution, watch the video.

4420

Of your peers have already watched this video.

28:30 Minutes

The most insightful time you'll spend today!

Explainer

Data Leaders in 2021 and Beyond: How to Prepare

As technologies and workplace environments are quickly evolving, how people and businesses leverage data in their day-to-day workflow is also changing.

Data leaders are an emerging force navigating and charting pathways forward towards new horizons for how people and organizations experience data.

In this presentation, Pedro Arellano, Product Marketing Director, Looker, will cover three areas:

  • Underline that the value of data within an organization is no longer up for debate. Everybody needs data, and everybody’s job has the potential to be improved with data.
  • Three trends that will help put into context how we’ve arrived at this particular moment of such great potential for data.
  • Why data leaders, almost regardless of their title, are now in a position to really influence the direction of organization like never before.

Learn what’s guiding the thinking of data leaders in 2021 and beyond.

Blog

Dataplex: Google Cloud’s Intelligent Data Fabric to Manage Data and Analytics at Scale

4861

Of your peers have already read this article.

2:30 Minutes

The most insightful time you'll spend today!

Google Cloud's latest offering, Dataplex is an intelligent data fabric that helps centrally manage, monitor, and govern data across data lakes, data warehouses and data marts. Read to understand how Dataplex could be relevant for your business.

Enterprises are struggling to make high quality data easily discoverable and accessible for analytics, across multiple silos, to a growing number of people and tools within their organization. They are often forced to make tradeoffs—to move and duplicate data across silos to enable diverse analytics use cases or leave their data distributed but limit the agility of decisions.  

Today we are excited to announce Dataplex, an intelligent data fabric that provides a way to centrally manage, monitor, and govern your data across data lakes, data warehouses and data marts, and make this data securely accessible to a variety of analytics and data science tools.  

Dataplex provides an integrated analytics experience, bringing together the best of Google Cloud and open source tools, so you can rapidly curate, secure, integrate, and analyze data at scale. With built-in data intelligence using Google Artificial Intelligence (AI) and machine learning (ML) capabilities and a flexible consumption model, you can now spend less time wrestling with infrastructure and more time focused on driving business outcomes.

integrated analytics experience.jpg

Dataplex enables you to:

  • Achieve freedom of choice to store data wherever you want for the right price/performance and choose the best analytics tools for the job, including Google Cloud and open source analytics technologies such as Apache Spark and Presto.
  • Enforce consistent controls across your data to ensure unified security and governance
  • Take advantage of the built-in data intelligence using Google’s best in class AI/ML capabilities to automate much of the manual toil around data management and get access to higher quality data.

Early customers like Equifax, Loblaw, and ANZ are excited about using Dataplex to address data management complexity.

“Dataplex will greatly simplify the existing analytics workflows within Equifax with its unified data fabric and single interface for policy management and governance across all our analytics data. Its built-in data discovery and data quality features will ensure that our data scientists and analysts always have access to high quality data that they can trust. Dataplex aligns well with our enterprise data strategy and we are excited to partner with Google Cloud on this.”
Kumar Menon, SVP, Data Fabric & Decision Science Technology, Equifax.

“Loblaw is Canada’s food and pharmacy leader, and we are excited to be an early adopter of Dataplex. We could significantly benefit from Dataplex as it provides a single pane of glass for end-to-end data management and governance. We are particularly interested in improving platform resilience and data quality by detecting anomalies as early as possible in the data pipeline with the help of Dataplex.”
Elton Martins, Senior Director of Data Insights & Analytics, Loblaw

“We are undergoing a major data transformation at ANZ, bringing together our various data assets and building a cohesive data ecosystem for customer benefit. Dataplex’s vision and capabilities align well with our current data strategy to build a unified data fabric for all our analytics and AI/ML use cases. We are excited to partner with GCP on Dataplex and test the product in private preview.”
Ashish Shekhar, Head of Technology – Enterprise Analytics & Applied AI, ANZ 

Dataplex is built for distributed data. We are starting with data stored in Google Cloud Storage and BigQuery, with support for other data sources coming soon. It provides a workflow-driven experience helping you build an open data platform and make data easily accessible to your end users while ensuring your policies and best practices are consistently enforced.

Organizing and curating your data

One of the core tenets of Dataplex is letting you organize and manage your data in a way that makes sense for your business, without data movement or duplication. For that, we are providing logical constructs like lakes, data zones and assets. Those constructs enable you to abstract away the underlying storage systems and become the foundation for setting policies around data access, security, lifecycle management, and so on. 

For example, you can create a lake per department within your organization (Retail, Sales, Finance, etc.) and create data zones that map to data readiness and usage (landing, raw, curated_data_analytics, curated_data_science, etc.). 

Once you have your lakes and zones setup, you can now attach data to these zones as assets. You can add data from different types of storage (e.g. GCS Bucket and BigQuery dataset) under the same zone. You can also attach data across multiple projects under the same zone.

Organizing and curating your data.jpg

You can ingest data into your lakes and zones using the tools of your choice including services such as DataflowData FusionDataprocPub/Sub or choose from one of our partner products. Dataplex provides built-in 1-click templates for common data management tasks. 

Securing your data 

Dataplex enables you to define and enforce consistent policies across your data, irrespective of where it physically resides. Data owners can easily set up policies for specific data domains based on business needs, without thinking about where the data is stored while data stewards get global visibility into governance policies and permissions across their data. 

You can apply security and governance policies for your entire lake, a specific zone, or an asset. Dataplex maps the policies to the underlying storage and pushes down permissions to the storage layer to provide end-to-end secure data access. Additionally, you can secure not just data but also related artifacts like notebooks, scripts, and models using the same set of access policies.

Making high quality data available for analytics and data science

One of the biggest differentiators for Dataplex is our data intelligence capabilities using Google’s best in class AI/ML technologies.  As you bring the data under management, Dataplex will automatically harvest the metadata for both structured and unstructured data, with built-in data quality checks. All of the metadata is automatically registered in a unified metastore, and made available for search and discovery. It is also published to BigqueryDataproc Metastore, and Data Catalog – ensuring that you have the same consistent data context and access across your tools.

For example, when you write parquet files to a Google Cloud Storage bucket – Dataplex will automatically extract metadata of these files, detect a tabular schema, including hive-style partitions, run data quality checks, and make this data queryable in BigQuery as an external table and from any open source or partner application – with the same consistent security and access policies you defined at the logical data layer. 

Your data scientists and analysts now have secured access to this data that meets your quality bar and governance rules via the tools of their choice, without needing any additional processing.

dataplex.jpg

One-click access to collaborative analytics

Dataplex provides fully managed, one-click analytics environments enabling you to use the power of Apache Spark and BigQuery with support for other engines coming in the future. 

As data administrators, you now have the flexibility to pre-configure these environments with the right cost and financial governance measures without taking on the overhead of managing and maintaining the infrastructure required to power these environments. You can easily configure different environments for different types of workloads and share it with multiple users using their IAM credentials. Dataplex manages the provisioning, monitoring, scaling, and shutdown of these environments. 

As data scientists, analysts, and engineers, you now have a turn-key experience to run your analysis using notebooks and a SQL workbench. You can search for notebooks and scripts alongside data, save and share your work with other users, and schedule your notebooks or scripts for recurring workloads – all using the same integrated experience within Dataplex.

access to collaborative analytics.gif

Building an open platform with industry leaders

We are partnering with industry leaders such as AccentureCollibraConfluentInformaticaHCLStarburstNVIDIATrifacta, and others to build an open platform to power analytics at scale. Our partners are excited about the capabilities that Dataplex will provide:

“Collibra is excited to partner with Dataplex to provide data governance and data quality for consistent controls across distributed data. Pairing Collibra’s multi-cloud and hybrid solution with Dataplex allows enterprises to securely open up access to more, higher quality data for users and analytics using a single unified view.”
—Jim Cushman, Chief Product Officer, Collibra

“Dataplex builds on Google Cloud’s commitment to open source by integrating with Apache Kafka®, a leading open source platform for event streaming. We at Confluent, the platform for data in motion that completes Apache Kafka® to be enterprise-ready, are excited to partner with Dataplex to enable customers to bring together distributed, real-time data and build a unified data fabric for end to end analytics.”
—Paul Mac Farland, VP, Head of Customer Solutions and Innovation, Confluent

“We are excited to partner with Google Cloud’s Dataplex team as we look to provide our joint customers an integrated and open data fabric for analytics at scale. Extending Dataplex’s data management and data quality capabilities with Starburst Enterprise will accelerate time to value for enterprises looking to connect distributed data without having to move data.”
—Justin Borgman, CEO and Co-founder, Starburst

Next Steps

Dataplex is now available in preview for a select number of customers. For more information, visit our website or watch the recording. If you would like to sign up, please click here.

3133

Of your peers have already watched this video.

17:00 Minutes

The most insightful time you'll spend today!

Blog

How FLYR and Google Cloud Help Airlines Forecast Demand and Set Prices

FLYR Labs is an international team of industry experts and specialists in revenue management that works to bring in intelligence to the airlines companies. FLYR uses machine learning and AI to help predict demand and optimize price so that every airline is operating its complete capacity. Watch the video from Architecting with Google Cloud to deep-dive into a use case with FLYR involving the use of historic data, competitors data and future information to build model for outputting demand, set prices and optimize revenue. You can can even have a quick view of the FLYR ML platform!

More Relevant Stories for Your Company

Case Study

City of San Jose Ensures Critical Services Reach Community Using AI Translation

San José is one of the most diverse U.S. cities, with residents speaking more than 100 languages. Several years ago, we set out to improve city community interactions through more equitable service management and delivery. This demanded a new approach to automating the intake of requests from a majority population

Blog

Global Communication Made Easy with Translation Hub

Hello! ¡Hola! 你好! नमस्ते! Bonjour! Being greeted in your own language can instantly put a smile on your face, and organizations around the world recognize that messages become more inclusive and powerful when translated into many languages. As the Head of Product Management for Translation AI at Google Cloud, I

Blog

An Expert’s Opinion on What Early-stage Startups Must Know

As lead for analytics and AI solutions at Google Cloud, my team works with startups building on Google Cloud. This puts us in the fortunate position to learn from founders and engineers about how early-stage startups’ investments can either constrain them or position them for success, even at the seed

Trend Analysis

2022 Healthcare Trends: Healthcare Data, M&As, Better Patient Care, AI in Drug Development & Strategic Partnerships

The COVID-19 pandemic continues to push the healthcare and life sciences industry in entirely new ways. In record time, we’ve witnessed public health officials, vaccine developers, equipment manufacturers, and essential workers take life saving actions—regularly putting their own lives at risk—to respond to the exceptional challenges of our time.  Yet

SHOW MORE STORIES