Google Cloud's Role in Minimizing Memory Errors Impact for SAP Customers - Build What's Next
Blog

Google Cloud’s Role in Minimizing Memory Errors Impact for SAP Customers

3739

Of your peers have already read this article.

2:00 Minutes

The most insightful time you'll spend today!

To minimize far-reaching effects of memory errors on customers' business, Google Cloud's Memory Poisoning Recovery (MPR) capabilities can protect SAP customers running HANA against unplanned and expensive downtime!

Every cloud system begins with high-quality hardware infrastructure. Sometimes, however, hardware breaks — and when it happens, our most important goal is to minimize the impact on our customers and their cloud workloads.

Memory errors are the most common type of hardware failure, and they’re also one of the most challenging in terms of their impact on production workloads and system reliability. That’s why we’re excited to share what Google Cloud has been doing to minimize the impact of memory errors. If your business runs SAP HANA in the cloud, this is an important innovation —  one that Google Cloud is proud to deliver to our customers.

Memory errors: A big problem with a long history

First things first: Memory errors are a high priority because they happen often. And when they happen, the disruption can have far-reaching effects on your customers and your business. 

In 2009, Google Cloud published the first major study on memory reliability. We found an average error rate of over 8% per year in DIMM modules installed in production systems. Given that each generation of DDR RAM packs more capacity into smaller packages, it’s safe to think that memory hardware has become less reliable since then.

Memory error impacts: They could be worse, but they’re far from good

What happens when a system detects a bad segment in a DIMM module? While data loss or corruption from memory errors is not common, some errors are correctable but some are not, potentially resulting in a critical system failure..  

Modern CPUs are equipped with error-correcting memory features and are very good at correcting simple errors with ECC (Error Correction Code). The challenge is that most of the software that runs on a host system — whether it’s a hypervisor, a virtual machine, an operating system, a database or an application — will crash instantly when it encounters an uncorrectable memory error. In a cloud environment, this kind of crash can take down cached data and even data saved to a local SSD. The crashed applications will recover, but the process means several minutes of downtime. The more data you have, the longer this process will take.

Sometimes, that’s merely an inconvenience. Other times, it’s a very big deal. A Google Cloud customer running business-critical SAP applications and an in-memory HANA database might measure downtime costs well over $10,000 per minute in lost revenue and other direct impacts. Many HANA databases load into terabytes of memory, and it can take an hour or longer to get everything restarted and back to normal after a crash. For SAP HANA, a fast recovery with up to 10 minutes of downtime requires a redundant replica provisioned all the time, doubling the cost.

And statistically speaking, when a HANA instance occupies almost all of the memory on a host system, it’s also the most likely application to stumble across a memory error. You can see why this would be a problem.

 The ‘victim neighbor’ VM challenge

There’s a final problem to consider when a memory error takes out production applications: what we call the “victim neighbor” issue.

In any cloud, a single physical host is a multi-tenant environment that might run dozens of VMs, potentially owned by dozens of different customers. A memory error won’t just crash the VM actually using the bad section, it will crash every VM running on the system. That’s a standard VM response to memory errors on a host system, and it will happen to any VM architecture available on the market today to avoid memory corruption. 

Overall, this “victim neighbor” effect accounts for more than 90% of the VMs that get knocked down by a memory error on a physical server. That’s a huge blast radius for such a common problem.

A practical solution to memory-error impacts

You can see why managing this problem is a big deal for Google Cloud. While we know that some failures are inevitable, we have developed another way to tackle the problem. Google Cloud already maintains some unique and valuable tools, such as Live Migration, that help our customers minimize unplanned downtime.When we integrate these tools with recent work that leverages error-handling capabilities built into CPUs (courtesy of Intel) and into certain applications (in particular, SAP HANA), we get a solution that dramatically reduces downtime and disruptions related to memory errors — in many cases, to the point where customers won’t even know there was a problem.

The Google Cloud solution: Memory poisoning recovery

At a big picture level, we refer to our solution as Memory Poisoning Recovery (MPR). It combines some existing Google Cloud capabilities, some new capabilities, and some important third-party capabilities at the CPU (Intel) and application (SAP HANA) levels. MPR can be broken down into two main processes:

Memory Error Isolation 

  • Step 1: We hardened our VM technology to be more robust against memory errors. We intercept and analyse the memory error coming from the system. Then we flag the signaled region of a memory DIMM with an uncorrectable error as “poisoned”. 
  • Step 2: Then we trigger processes to keep track of these “poisoned” regions and the VMs they affect so they can’t affect data integrity. 

Memory Error Recovery

  • Step 3: Then we notify the Guest OS & the MCE-aware applications that a memory error has been recorded, in a manner that allows the applications to execute application relevant memory error handling.
  • Step 4: At the same time we communicate with Google Cloud Live Migration to begin moving guest VMs off the affected host. This ensures customers are running on a healthy host which reduces the probability of more uncorrectable errors happening and avoids further downtime.

Below is a simple visual of how this all works:

memory poisoning recovery.jpg

How MPR makes life better for customers

Let’s look again at the different groups of Google Cloud customers involved in a memory error scenario and how we can help them achieve a happier ending after a crash — starting with the customer running the VM and application that actually triggered the memory error. 

Customer Group: MCE-Aware SAP HANA with Fast Restart enabled on a VM directly affected by a memory error.

Customer Group: Customers running other, non MCE Aware applications on a VM directly affected by a memory error

Next, our “victim neighbors” group probably won’t even know there was a problem with the host system. Google Cloud Live Migration will move them to a new host, instantly and automatically, and avoid the crash-and-restart scenario.

Customer Group: Customers running other, any application on a VM not directly affected by a memory error

Simple steps for taking advantage of MPR

Our MPR capabilities will be available on our Google Cloud memory-optimized Compute Engine second generation instances in Q4 of 2021. We’ll continue to roll out the capability during the months ahead to additional instances and look for new ways to work with applications that adopt a MCE Aware architecture.

Most customers in the “victim neighbor” category will not need to lift a finger to experience the benefits. By marrying our Live Migration feature to some awareness of those MCE signals, we ensure that it hears the alarm first and gets a critical head start on the migration process before issues begin with the guest VMs. Our customers land safely on a new host, and their applications keep running.

For our SAP customers running HANA, MPR is all about protecting against loss. Unplanned downtime for a HANA environment is incredibly expensive, the recovery process from a hard crash is extremely long, and the business disruptions can be truly damaging to the business. Thanks to MPR, all of that cost and worry can get compressed almost to nothing — with Fast Restart reducing what can be an hour or more of downtime to a matter of seconds.

But our SAP customers have to take a critical first step to claim these benefits. Fast Restart is a crucial piece of the MPR solution, and it is not enabled by default. Configuring your SAP HANA instance for Fast Restart involves changing a few configuration settings; the process is fast, easy, and doesn’t involve risk. 

Finally, if you’re not running your workloads — SAP or otherwise — on Google Cloud, consider the benefits of running on a cloud that mitigates a hardware reliability issue affecting businesses of every size and industry. And consider the value of tools like Live Migration that already help Google Cloud customers improve uptime and reduce risk.

Hardware failures happen, and they probably always will. But we’re proving how valuable it can be to avoid the bad things that usually happen when memory failures occur. Right now, only Google Cloud has a practical solution to this very difficult problem.    

Learn more about Fast Restart for SAP HANALive Migration and other key Google Cloud capabilities for your SAP environment.

Blog

Explore The New Era of Flexibility: Streamlined AWS-to-Google Cloud Migration

1345

Of your peers have already read this article.

3:30 Minutes

The most insightful time you'll spend today!

Explore how Google Cloud's Migrate to Virtual Machines enables smooth AWS-to-Google Compute Engine migration, with minimal changes, downtime, and risk, while maximizing scalability and flexibility. Read more!

As an IT leader, you’re asked to do it all: innovate and optimize your tech stack for business outcomes — all while being secure and compliant. It takes heroic efforts to achieve innovation and progress while also tightening budgets and teams. This is why many of you are considering migrating applications to Google Cloud, for benefits like scalability and flexibility, security and compliance, disaster recovery and business continuity, and cutting-edge technologies at lower costs.

To help you do this, we suggest using Google Cloud’s Migrate to Virtual Machines – part of Migration Center. This managed cloud service lets you lift and shift workloads at scale to Google Cloud Compute Engine with minimal changes and risk.

And recently, we rolled out our latest release which introduces support for workload migration from AWS to Google Compute Engine. With this addition, you can now migrate both your on-prem and AWS workloads at scale. This means centralized management for your end to end workload migration journey from both sources via Cloud Console or APIs. 

Simple and easy migrations from AWS and VMware sources to Google Cloud

Migration of AWS EC2 instances directly to Google Compute Engine using Migrate to Virtual Machines follows a well established and easy journey, which means a minimal learning curve for users who are already migrating workloads from VMware. Workload migration is agent-less, which means you do not need to access or alter workloads as a prerequisite for migration, allowing you to execute zero-touch migrations. Migrate instance data with no interruptions to the running workload at the source for a fast cutover to Google Cloud. In addition, our end-to-end cloud console interface surfaces your AWS EC2 inventory, migrations, and groups so you can execute migrations without ever leaving the cloud console interface. 

Large-scale migrations 

Completing a large-scale migration project in a timely manner calls for careful planning and streamlined migration sprints. The Migrate to Virtual Machines’ Groups construct enables you to group source VMs together in the planning phase. When it’s time to execute the planned migration, VM Groups let you execute migration operations on a group level, or on a subset of the group, streamlining the process at scale.

Minimal downtime and risk

Application uptime is key to keeping your business running. Every migration with this latest release of the service periodically replicates data from the source workload to the destination without manual steps or interruptions to the running workload, minimizing workload downtime and enabling fast cutover to Google Cloud. You can also launch non-disruptive migration tests — referred to as test-clones — to help you validate that these workloads will work properly in the cloud before cutting over. This helps avoid issues that might have otherwise been costly or disruptive to your business. 

How the service works

Migrations simply work, at scale, in a managed service fashion. With Migrate to Virtual Machines, there’s no requirement to provision or manage migration-specific resources in the cloud. The service uses replication-based migration technology to lift and shift workloads from source environments to Google Cloud. The Migrate Connection replicates source VM disk snapshots in the background with no interruption to the source workload. Replicated data is encrypted in transit and at rest, and when you instantiate a migrating VM using a test-clone or cut-over, the service seamlessly adapts your source VM operating system to boot and run natively in the cloud — including configuring network settings and deploying Google Cloud guest packages. 

The migration journey of an EC2 instance — or VMware VM — to Google Cloud is comprised of the following steps:

1. Onboarding a source VM for migration: Onboard one or more VMs for migration from the source environment fleet.

2. Configure landing zone target: You can migrate an instance to any Google Cloud project in your environment and update landing zone details at any time before executing a test-clone or cutover.

3. Initiate VM data replication of source workload: Migrate to Virtual Machines periodically replicates instance disks to the cloud with no interruption to the source instance. You can control replication frequently and pause or resume at any point in time. 

4. Test migrating instance: Test-clone creates a copy of your source instance in the defined landing zone to validate the migrating instance in the cloud before executing a cut-over. You can repeat the test-clone multiple times to multiple landing zones for thorough validation

5. Cutover migrating instance: Cutover operation shuts down your source instance and then performs the short final sync to Google Cloud. Migrated VM is instantiated in the target landing zone.

Getting started with Migrate to Virtual Machines 

It’s quick and easy to start migrating your AWS EC2 instances and on-premises VMs today:

  1. Enable the vmmigration API in a Google Cloud project 
  2. Create an AWS source in your environment
  3. Onboard and initiate replication of instance data from source 
  4. Set migrating instance target details. 
  5. Perform non-disruptive tests of your migrating instance using test-clone
  6. Cutover your instance to the cloud with minimal down time

You can also visit our website to learn more about Migrate to Virtual Machines. If you know you have to migrate in 2023 but aren’t sure how to get started, you can sign up for a free discovery and assessment of your current IT landscape so we can help craft the ideal migration plan for you and your business.

Blog

Know the Leaders of Google Cloud Public Sector Community

5017

Of your peers have already read this article.

1:30 Minutes

The most insightful time you'll spend today!

To mark its recent foray into the public sector under the visionary leadership of three extraordinary individuals, Google Cloud Public Sector celebrates their contribution and role in driving new product initiatives to the industry.

At Google Cloud, being a strategic partner is part of our DNA. Whether it’s listening closely to our customers, helping to build team skills for innovation or simply being there (since we know the cloud is 24/7), we get excited about working hands-on with customers to deliver new solutions. 

As we look to solve decades-old challenges with new technologies in workforce productivity, cybersecurity, and artificial intelligence/machine learning, we know that we are only as good as the people behind the technology. Today, we’re proud to spotlight a few of the inspiring folks behind Google Cloud Public Sector and celebrate their recognition in the industry. 

Melissa Adamson, Head of Government Channels at Google Cloud, has been named to the highly respected Women of the Channel list for 2021. This annual list recognizes the unique strengths, leadership and achievements of female leaders in the IT channel. The women honored this year pushed forward with comprehensive business plans, marketing initiatives and innovative ideas to support their partners and customers.

Melissa was brought on to build the Public Sector channel from scratch. The initial focus was building the channel for the US government team and has since expanded to include education, Canada and Latin America.

Having a career background at both Microsoft and Accenture, Melissa leveraged her extensive professional network to build the organic partnerships needed to accelerate the Public Sector partner ecosystem. This helped her drive two key wins (US Postal Service and PTO) and personally recruit top cloud partners in the industry. Melissa loves card games and is learning a new language.

Todd Schoeder, Director of Global Public Sector Digital Strategy, was recently featured in the “Top 20 Cloud Executives to Watch in 2021” by Wash Exec. Recognized for his work in helping customers navigate through the impact of COVID-19 and developing innovative solutions to meet mission challenges, he says: “New partnerships are required to solve for the problems of the future. Challenges that were previously thought of as insurmountable, too risky or expensive, are actually quite the opposite — as long as you have the right partner that is working in your best interest with you.”

Josh Marcuse, Head of Strategy & Innovation, received his second Wash100 Award for leading a digital transformation team that works to drive the development of public sector solutions, including cyber defense, smart cities, and public health.

Josh has launched services to support collaborative team operations including Workspace for Government and an artificial intelligence-based customer service platform to support remote work needs. His work also includes leading Google Cloud’s partnerships with organizations to improve data sharing in the public health community, contact tracing activities, and supporting research efforts across national laboratories. 

Like Melissa, Josh was brought on to build a new team dedicated to strategy and innovation. This team’s purpose is to bring an intense focus to public sector mission outcomes and the public servants who own them. Josh spent a decade pushing digital modernization and workforce transformation at the U.S. Department of Defense, and co-founded the Federal Innovation Council at the Partnership for Public Service, and now brings that domain expertise to supporting government workers who are driving digital transformation.

Join us in celebrating these folks for their leadership and contributions!

Blog

TELUS and Google Cloud Partner to Move Towards a More Sustainable Future

4481

Of your peers have already read this article.

2:30 Minutes

The most insightful time you'll spend today!

Google Cloud and Telus have come together to make the planet healthier by ensuring that their operations are as environmentally responsible as possible. Find out how you can leverage innovative technologies to empower sustainable business practices.

Environmental sustainability is a key priority for TELUS, a world-leading communications technology company. It continues to rank in the top 100 most sustainably managed companies in the world, and seeks to make a healthier planet for all by leveraging its global-leading technology, compassion to drive social change and reduce our collective carbon footprint through innovative technologies and sustainable business practices.

TELUS surpassed its sustainability objectives in 2019 and is now on a journey to procure all of its electricity from renewable or low-emitting sources by 2025. Next, it aims to achieve net carbon neutrality for its operations by 2030. TELUS has also been named to the Dow Jones Sustainability Index for 21 consecutive years, a feat unmatched by any other North American telecom or cable company. In 2021, it became the first company in Canada to release a Sustainability-linked bond (SLB) framework and complete an SLB offering, formally linking TELUS financing to its environmental performance.

“We’ve spent the last decade becoming a global leader in sustainability, helping make the planet healthier by ensuring that our operations are as environmentally responsible as possible,” said Geoff Pegg, Head of Sustainability and Environment at TELUS.

In part, TELUS’ strategy is focused on three key areas:

  1. Seek the best renewable energy options available
  2. Focus on migrating workloads to the cloud
  3. Embrace a multiplier effect through the use of sustainable partners

Renewable energy impact

Part of this environmental responsibility involves investing heavily in renewable energy sources through power purchase agreements (PPAs) that help renewable energy providers like wind farms and solar companies develop their infrastructure. TELUS executed PPAs with four Alberta-based solar and wind facilities to provide 100 per cent of its electricity load demand in a province where one-third of the grid is powered by coal.

As a technology company, electricity represents a large portion of TELUS’ energy needs: 80 percent of the operational carbon footprint comes from the power requirements for TELUS’ network and administrative buildings, Pegg explains. While TELUS is using renewable energy sources and low-emitting energy grids to power its buildings and network, there’s also the often-forgotten part of the carbon emissions equation: the energy it takes to power data centers. As the International Energy Agency recently reported, data centers represent 1 percent of the global electricity demand and that figure is expected to keep rising as the world increases usage of data-heavy technologies.

“It’s probably no surprise that everyone, whether you’re a business or a consumer, is concerned about reducing carbon emissions,” said Chris Talbott, the Google Cloud Sustainability Lead. “A lot of us think about the carbon emissions associated with our cars or with the electricity that powers our homes, but oftentimes we forget about the carbon emissions that come from the digital services that we use or the networks required to deliver that data.”

As a leader in sustainability, how can TELUS meet the energy demands of its customers while also protecting the environment? One way is through the company’s previously announced collaboration with Google Cloud. The two companies are working together to build a more sustainable world through technology and reduce TELUS’ carbon footprint, create value along the entire supply chain, and optimize industry solutions for social impact through data analytics and machine learning.

Taking a cloud first approach — reducing carbon emissions with green cloud computing

Google became carbon neutral in 2007 and has achieved 100 per cent renewable energy matching every year since 2017. Google has invested in renewable energy to match the electricity we use across our entire operations, including Google Cloud, meaning every workload that TELUS runs on Google Cloud has been matched with renewable energy purchases.

“The operational carbon footprint of running anything on Google Cloud is zero,” Talbott said. Also, by working with Google, TELUS gets the benefit of economies of scale using less electricity. Not only is TELUS leveraging Google data centers, it’s also relying on the digital collaboration made possible by Google Workspace to reduce the amount of travel required by employees attending meetings in different offices. Collaboration tools like Google Meet can reduce the carbon footprint of in-person conferences by 94 percent.

Google compensates for the environmental footprint of any electricity used in the data center and out to the edge network. “You can feel pretty good about using Google Meet because it’s carbon-neutral,” Talbott said.

Multiplier through sustainable partnerships — green cloud computing radiates out

By supporting TELUS in its environmental sustainability efforts, Google Cloud is also enabling TELUS to do the same for its various partnerships. For example, powered by Google Cloud’s infrastructure and data analytics capabilities, TELUS is partnering with Picacity (formerly NXN Digital) and Google Cloud to deliver an ecosystem of integrated smart technologies that enable cities to improve the lives of their residents.

From dynamic traffic signaling that reduces congestion and emissions, to data analytics that create smarter, more efficient city planning, the partnership is transforming the way municipalities operate in our increasingly digital world.The partnership is built on four foundational pillars of infrastructure and environmental sustainability, intelligent transportation, public safety and security, and health. In the case of intelligent transportation, this means sensors, cameras, and other devices are built into or near roads, sidewalks, and bike paths to provide data for innovative software to improve traffic flow in real time. The data can then foster informed decisions about infrastructure, city planning, fleet optimization, and public safety.

All of these environmental measures may seem small when compared with the enormity of the problem that is climate change, but as Talbott said, “Change begins with the small decisions we make every day such as paying attention to the practices of companies that we’ve come to rely on daily in the modern world. They may seem small and in the margins, but at scale, this is how we can make a real impact.”

3628

Of your peers have already watched this video.

1:30 Minutes

The most insightful time you'll spend today!

Case Study

How McKesson Gains Insights by Running SAP on Google Cloud

McKesson, a 185-yeal old, $200 billion, Fortune 6 pharmaceuticals and health information technology company, with over 80,000 employees migrated their SAP solution to Google Cloud for advanced healthcare analytics.

With changing consumer expectations, the company needed to change its architecture to be able to better serve its customers. And the old on-premise infrastructure was hindering the company from being able to meet those expectations.

Watch this video as Andrew Zitney, SVP and CTO at McKesson, explains how SAP on GCP helps the 180-year old company modernize and meet the changing expectations.

Whitepaper

Why Indian Enterprises Need to Embrace The Cloud-First Imperative to Accelerate Digital Transformation

DOWNLOAD WHITEPAPER

3790

Of your peers have already downloaded this article

3:30 Minutes

The most insightful time you'll spend today!

Digital transformation is rewriting the rules of business both in India and worldwide. Digital customer experiences deliver easy, effective, and emotional touchpoints that focus operations on what the customers value. Around half of Indian decision makers prioritize the improvement of CX and the simplification of operations, as top priorities in their business agenda, according to a Forrester Consulting study of 360 business and technology decision makers of Indian enterprises.

According to the study, forward-thinking enterprises are increasingly turning to cloud to support their business as they attempt to keep pace with evolving customer needs. As a result, cloud has become a strategic priority, and ensuring its support in the marketplace will only enable digital business and accelerate innovation.

The study reveals that:

  • Public cloud is a key enabler for the transformation of digital business.
  • Security, inconsistent monitoring tools, and legacy applications are top barriers to public cloud expansion.
  • Enterprises are expanding their adoption of the public cloud and want to gain a competitive edge.

Download this study to understand why more and more organizations are moving applications to the cloud in order to take advantage of scalability, lower capital costs, ease of operations, and the resilience offered by the public cloud.

More Relevant Stories for Your Company

Blog

Accelerating Success: Tips and Techniques for Optimizing and Scaling Your Startup

At Google Cloud, we want to provide you with the access to all the tools you need to grow your business. Through the Google Cloud Technical Guides for Startups, leverage industry leading solutions with how-to video guides and resources curated for startups. This multi-series contains 3 chapters: Start, Build and

Case Study

SevenRooms: Building a profitable, sustainable future in food and beverage with Google Cloud

Finding ways to increase customer loyalty and profitability in a post-COVID world is top of mind for hotels, bars, and restaurants. Unfortunately, many food and beverage services providers struggle to deliver the personalized experiences that keep guests coming back for more. The reality is that most traditional hospitality apps offer

Trend Analysis

APAC’s Retail Digital Pulse by Google Cloud and IDC Retail Index

In the post COVID-19 pandemic world, some retail businesses have aced digital transformation while others are still lagging behind. The Google Cloud commissioned IDC Retail Insights analyzed over 1, 108 retailers across seven nations in the Asia Pacific region and across eight different segments (online, drugstores, speciality shops, convenience stores,

Case Study

How Companies can Improve Scalability, Flexibility, and Reliability While Reducing Costs: Tips from Route4Me

Google Cloud Results Improves application performance by 8x to 12x; customers can create increasingly complex optimized driving routes in single-digit secondsImproves customer satisfaction via increased reliability and greater application performanceFocuses on adding value to customers by improving software and algorithms, not infrastructure managementSaves 5x in infrastructure costs In 2009, Dan

SHOW MORE STORIES