Tau VMs Joins Google Cloud to Offer Cost-effective Performance of Scale-out Workloads

5606
Of your peers have already read this article.
2:00 Minutes
The most insightful time you'll spend today!
Scale-out workloads demand the best combination of performance and price to bring down the cost of delivering applications, all while providing an excellent user experience. We are excited to announce a new virtual machine (VM) family, Tau VMs, coming to Google Cloud. Tau VMs extend Compute Engine’s VM offerings with a new option optimized for cost-effective performance of scale-out workloads.
T2D, the first instance type in the Tau VM family, is based on 3rd Gen AMD EPYCTM processors and leapfrogs the VMs for scale-out workloads of any leading public cloud provider available today, both in terms of performance and workload total cost of ownership (TCO). The x86 compatibility provided by these AMD EPYC processor-based VMs gives you market-leading performance improvements and cost savings, without having to port your applications to a new processor architecture.
As illustrated below, Tau VMs offer 56% higher absolute performance and 42% higher price-performance (est. SPECrate2017_int_base) compared to general-purpose VMs from any of the leading public cloud vendors.


SPECrate is a trademark of the Standard Performance Evaluation Corporation. More information available at www.spec.org

What our customers and partners are saying
Snap
“At Snap, it is critical for our business to continue improving our scale-out compute infrastructure for key Snapchat capabilities like AR, Lenses, Spotlight and Maps,” said Cody Powell, Senior Engineering Manager, Snap Inc. “We were impressed when we tested Google Cloud’s new Tau VMs with Google Kubernetes Engine. While it’s early days, we believe we can gain double digits in infrastructure performance improvements for key workloads—enabling us to do more with less and invest even more in new features for our amazing Snapchat community.”
Twitter
“High performance at the right price point is a critical consideration as we work to serve the global public conversation,” said Nick Tornow, Platform Lead, Twitter. “We are excited by initial tests that show potential for double digit performance improvement. We are collaborating with Google Cloud to more deeply evaluate benefits on price and performance for specific compute workloads that we can realize through use of the new Tau VM family.”
DoiT
“DoiT partners with leading cloud vendors who are focused on growth and cost optimization,” said Yoav Toussia-Cohen, CEO, DoiT International. “In our preliminary testing of Google’s new Tau VMs with the Coremark benchmark, we were thrilled to see the incredible performance at 50% better than a comparable ARM-based offering from another leading public cloud. With Tau VMs, Google Cloud has set a new bar for price-performance, making the cloud even more accessible to digital-native companies. We are excited to bring Google’s Tau VMs to our joint customers.”
Designed for demanding scale-out workloads
Tau VMs bring the benefit of Google’s long-standing experience engineering platforms for scale-out workloads to our customers. They come in multiple predefined VM shapes, with up to 60vCPUs per VM, and 4GB of memory per vCPU. They offer up to 32 Gbps networking bandwidth and a wide range of network attached storage options, making Tau VMs ideal for scale-out workloads including web servers, containerized microservices, data-logging processing, media transcoding, and large-scale Java applications.
Google Kubernetes Engine support
Google Kubernetes Engine (GKE) is the de facto standard for organizations looking for advanced container orchestration, delivering the highest levels of reliability, security, and scalability. GKE supports Tau VMs on day 1, helping you optimize price-performance for your containerized workloads. You can add Tau VMs to your GKE clusters by specifying the T2D machine type in your GKE node-pools.
Pricing
Tau VMs will be priced to support significant TCO and price-performance improvements for your cloud applications. A 32vCPU VM with 128GB RAM will be priced at $1.3520 per hour for on-demand usage in us-central1.
Coming soon to a Google Cloud region near you
If you are interested in trying out T2D VMs when they become available in Q3 2021 please sign-up here.
How Lowe’s SRE Team Decreases Mean-time-to-recovery (MTTR)

3363
Of your peers have already read this article.
1:30 Minutes
The most insightful time you'll spend today!
Editor’s Note: In a previous blog, we discussed how home improvement retailer Lowe’s was able to increase the number of releases it supports by adopting Google’s Site Reliability Engineering (SRE) framework on Google Cloud. Lowe’s went from one release every two weeks to 20+ releases daily, helping meet its customer needs faster and more effectively. Today, the Lowe’s SRE team shares how they used SRE principles to decrease their mean-time-to-recovery (MTTR) by over 80 percent.
The stakes of managing Lowes.com have never been higher, and that means spotting, troubleshooting and recovering from incidents as quickly as possible, so that customers can continue to do business on our site.
To do that, it’s crucial to have solid incident engineering practices in place. Resolving an incident means mitigating the impact and/or restoring the service to its previous condition. The average time it takes to do this is called mean time to recovery (MTTR). Tracking this metric helps us stay on top of the overall reliability of our systems at Lowe’s, while simultaneously improving the speed with which we recover. Our goal is to keep the MTTR metric as low as possible, so that failures don’t negatively impact our business. Here are the four areas we addressed to drive holistic improvement in our MTTR.
Lowe’s incident reporting process
To reduce MTTR, we created a seamless incident reporting process following SRE principles. Our incident reporting process is a workflow that starts at the time an incident occurs, and ends with an SRE captain who closes the action items after a postmortem report. With this approach, we are able to limit the number of critical incidents. The reporting process involves three core components: monitoring, alerting, and blameless postmortems.
Monitoring and alerting
Having proper monitoring and alerting in place is crucial when it comes to incident management. Monitoring and alerting tools let you detect issues as soon as they occur, and notify the right person in the shortest possible time to take action. From a measurement standpoint, we track this as our mean time to acknowledge (MTTA). This is the average time it takes from when an alert is triggered, to when work on the issue begins.
At the time of an incident, our monitoring and alerting tools notify the on-call SRE first responder via PagerDuty in the form of a phone call, text message and email. Our SRE software engineering team has done a lot of automation to enable various Service Level Indicator (SLI) alerts and Service Level Agreement (SLA) notifications. The on-call SRE then initiates a triage call with our service/domain stakeholders to resolve the incident. As a result, we reduced our MTTA from 30 minutes in 2019, to one minute – a 97 percent decrease.
Blameless postmortems: learning from incidents
A postmortem is a written record of an incident, its impact, the actions taken to resolve it, the root cause and the follow-up actions to prevent the incident from recurring (see example here). A blameless postmortem builds on that and is a core part of an SRE culture, and our culture at Lowe’s. We ensure that individuals are not singled out, and the outcome for all postmortems are directed toward learnings and process improvement.
For us, the postmortem process is the biggest part of our incident workflow. When an SRE creates a new postmortem report, the first step is to conduct a postmortem session with domain stakeholders to review the report. The postmortem then goes into the review stage and gets reviewed by more stakeholders in our weekly postmortem meeting. In the final stage of this process, the SRE captain will close the report once everyone in the weekly meeting agrees that the report is complete.
To conduct a successful postmortem, it is critical to keep the focus on identifying gaps and issues with the system and operations processes, rather than an individual, and generate concrete actions to address the problems we’ve identified. To ensure this, we follow a couple of best practices:
- We start by gathering the facts from the person who identified the problem, and each SLI owner has to identify a gap or the next SLI upstream owner who created the impact for them.
- Every SLI owner is provided full opportunity to present their case, and identifying the issue is done as a community exercise.
- Once action items and process changes are identified, an owner is nominated to complete the actions, or they will volunteer.
- For easy reference, we publish and store postmortems in our incident knowledge base. This process helps SREs continuously improve as future incidents arise.
Continuous Improvement
Encouraging a culture of honest, transparent and direct feedback that you need for blameless postmortems is often an iterative process that needs sponsorship from executives, empowering incident captains to lead the entirety of the discussion and outcomes. Running successful postmortems, and completing action items from them, needs to be recognized and accounted for in SRE performance objective assessment. As shared in Google’s SRE book, the best practice is to ensure that writing effective postmortems is a rewarded and celebrated practice, with leadership’s acknowledgement and participation. This is possibly the hardest part to accomplish in an effective postmortem during a cultural transformation unless you have full buy-in from leadership.
However, it’s all well worth it. This process is a key part of how we were able to improve our MTTR over time—from two hours in 2019 to just 17 minutes!
Our SRE incident reporting process has also transformed how our company solves issues. By streamlining this workflow from alerting, to solving an issue, to blameless postmortems, we have reduced our MTTR by 82 percent and our MTTA by 97 percent. Most importantly, our team is learning from every incident and becoming better engineers as a result. Visit the SRE Google Cloud website to learn more about implementing SRE best practices in the cloud.
Acknowledgement
Special thanks to Rahul Mohan Kola Kandy, Vivek Balivada, and the Digital SRE team at Lowe’s for contributing to this blog post.

5560
Of your peers have already downloaded this article
1:30 Minutes
The most insightful time you'll spend today!
Protecting a global network against persistent and constantly evolving cyber threats is one of the most important challenges faced by Google Cloud. So, how does Google’s global network protects seven different global businesses, each with over 1 billion customers, including popular Google services such as Google Search, YouTube, Maps, and Gmail?
The answer is a multi-step process, which is constantly changing and evolving to stay one step ahead of the malicious hackers and attackers. For instance, it’s network communications protocols—the rules that enable communications between systems—change multiple times per second to make malicious intrusions much harder.
Data in Google Cloud is encrypted both in transit and at rest. Google’s network capacity far exceeds any traffic load it hosts to thwart and DDoS attack. In addition, it has numerous other products, tools, and processes at work to provide defense in depth.
Download this e-book to get a detailed overview of Google Cloud’s approach to security and privacy.

5679
Of your peers have already downloaded this article
5:30 Minutes
The most insightful time you'll spend today!
Modernizing apps on the cloud isn’t an “all or nothing” decision. Businesses want the option to modernize on-premises or choose multi-cloud solutions that meet their needs. That’s why we created a new solution for running apps anywhere – simply, flexibly, and securely. Embracing open standards, Anthos lets you run your applications, unmodified, on existing on-prem hardware investments or in the public cloud. So that you write once and deploy anywhere.
Download this report and find out how to:
- Decouple infrastructure and applications with containers and Kubernetes
- Decouple cloud teams from one another so they can work independently
- Meet the challenges of microservice management using service mesh
- Implement a zero-trust security model to enforce more granular controls while maintaining a consistent user experience
Unlocking Economic Potential: Cloud FinOps

3251
Of your peers have already read this article.
4:00 Minutes
The most insightful time you'll spend today!
Built for a CapEx world, most organizations’ finance systems aren’t set up to take advantage of cloud’s dynamic, OpEx-driven consumption patterns.

GETTY
It may not be a household name yet, but chances are you’ve crossed paths with OpenX today. OpenX, a leader in programmatic advertising, operates one of the world’s largest ad exchanges, serving over 250 billion ad requests per day, connecting more than 30,000 brands and reaching nearly one billion consumers. To make it happen, in 2019, OpenX migrated entirely out of its data centers and became the first major ad-exchange platform to move completely to the cloud.
The OpenX CTO, Paul Ryan, knew that this cloud transformation initiative had the potential to increase costs faster than its revenues. To be successful, he needed his engineering, finance, and business teams to forge a new “cost-aware” culture, complete with effective cost visibility and controls. In other words, he needed Cloud FinOps — an operational framework and cultural shift that brings technology, finance, and business together to drive financial accountability and realize business benefits through cloud transformation.
Ryan laid out a cloud migration roadmap that included cost governance and controls around project ownership, established cost responsibility with engineering teams to accurately forecast cloud consumption, and challenged developers to lower per-unit costs — while at the same time improving performance, scalability, speed and global reach.
It worked! In just 9 months, OpenX reduced their per-unit cost by over 60%. The framework allowed OpenX to launch new regions in a matter of days, reduce their time to market for new features by over 50%, and complete their migration in record time — seven months! “We are now able to stop worrying about legacy infrastructure and focus more on our growth categories,” said Ryan. “Our tech stack is getting smarter and more sophisticated by the day, and we have the flexibility to scale our infrastructure in real-time as the business scales and evolves.”
Unblocking Cloud’s Potential
Cloud holds the key to a successful digital transformation. In fact, McKinsey forecasts that by 2030, the Fortune 500 alone may realize over $1 trillion of EBITDA value drivers associated with public cloud enablement. But unlike OpenX, many companies struggle to achieve near-term value objectives from their cloud investments. Surveys reflect that more than 30% of cloud spend in 2021 was wasted or inefficient, while upwards of 80% of CIOs have yet to achieve the business benefits of migrating to the cloud.
Traditional IT finance processes are ill-suited for cloud infrastructure: Traditional planning and budgeting processes are challenged to address dynamic consumption patterns and complex migrations. Centralized IT budgets using traditional allocations fail to provide the necessary visibility into sources of cost overruns. CapEx-focused cost controls have little ability to manage largely OpEx-driven spend. Trend-based forecasting often inaccurately predicts cloud costs. And developer teams lack access to cost-aware architecture patterns to deploy the applications more efficiently.
Enter Cloud FinOps
At Google we’ve worked with many companies, like OpenX, to help organizations realize the transformational benefits of the cloud by cultivating a culture of transparency and embedding agile processes to manage costs. We’ve distilled these learnings into a Cloud FinOps operational framework that gives organizations the financial governance and accountability they need to grow their business sustainably.

GOOGLE CLOUD
At a high level, a Cloud FinOps approach depends on five key areas:
- Accountability and Enablement
Accountability and enablement aim at instilling a cost-conscious culture across the organization. Oftentimes, this means standing up a cross-functional and dedicated team with members from technology, finance and engineering to establish cloud financial best practices and governance. In various organizations, we’ve seen this through an extension of a Cloud Center of Excellence, a Cloud Business Office or simply a Cloud FinOps team. Enablement focuses on empowering IT, finance and business leaders through training to help them better understand the economics of cloud services and the strategies to efficiently deploy and manage them. Cloud financial training guides teams on how to design cost-effective cloud environments, for example, embracing ”cloud-native” design principles such as auto-scaling/elasticity and Infrastructure as a Code.
- Measurement and Business Value Realization
Effective measurements not only create awareness and enable agile processes, but also support a culture that celebrates success and rewards teams for achieving business objectives. As such, measurement in the service of business value realization is about developing a comprehensive set of long-term benefits and cost KPIs to quantify the total net value of the return on digital transformation. Organizations often start with cost-related KPIs and eventually evolve those KPIs into business value metrics that are mapped to targeted business outcomes.
- Cloud Cost Optimization
Cloud cost optimization is an iterative and continuous process that provides a consistent methodology to manage cloud consumption cost-effectively. There are three key areas of optimization:
Resource optimization – Model cost-effective cloud usage based on utilization and consumption patterns.
Pricing optimization – Manage cloud spend through a continuous analysis of various pricing models. In a Google Cloud context, that might mean Committed Use Discounts, BigQuery flat rate reservations, etc.
Architecture optimization – Build applications with a cost-aware architecture by leveraging newer generation compute instances (like Tau VMs, which offer an industry-leading 42% better price-performance versus comparable offerings), or using managed services and serverless technology to offload operational overhead.
For example, video hosting, sharing and services platform provider Vimeo built transcoding pipelines by using Google Cloud Spot VMs to optimize their infrastructure spend. To do so, they created fault-tolerant workloads that could withstand preemptions, and in exchange, got up to a 91% discount compared to using regular on-demand instances.
- Planning and Forecasting
In the cloud, accurately forecasting your finances requires rethinking of traditional approaches to depreciation and trend-based forecasting of maintenance and licensing costs. One way to improve the accuracy of your dynamic cloud needs is to use workload-specific forecasting models that leverage a combination of trend-based models for steady-state workloads, driver-based models for scaling applications, as well as monthly variance analysis. In other words, you can define cloud budgets and forecasts by monitoring cloud consumption trends, allocating cloud cost pools with a proper tagging strategy that’s mapped to a chart of accounts in a general ledger, and conducting a cost-benefit analysis based on cloud infrastructure, implementation, and support costs.
- Tools and Accelerators
Without the proper tools and processes in place, understanding and managing cloud costs can be complex — and this especially true as organizations scale their business in the cloud. By deploying proper cloud cost management tools and accelerators such as Looker Cloud Cost Management and automation scripts to set guardrails and enforce cost control policies, organizations can effectively manage and track cloud spend with access to near-real-time billing and cost data to make better informed business decisions.
The key objectives of Google Cloud Cost Management tools are to make it as simple as possible for organizations to get visibility into their current and forecasted costs with built-in reporting and customizable dashboards; help drive greater accountability for cloud spending across the organization by providing flexible ways to organize cloud resources and allocate costs; provide strong financial governance controls to reduce the risk of overspending; and offer intelligent recommendations for optimizing cloud costs and usage.
Start Saving with Cloud
Businesses are continuously seeking to better operate and manage their cloud environments and the need is ever increasing to transparently manage cloud spend, optimize costs, and obtain their desired business agility. By enhancing your Cloud FinOps capabilities and adopting principles of continuous cost optimization, you too can accelerate the business value of cloud computing.
Special thanks to Bruce Warner, Daniel Petibone, Nihar Jhawar and FinOps Foundation community for their contributions and sharing their domain expertise to this important Cloud FinOps topic.
How Ather Energy is leveraging the Cloud to build and scale smart mobility solutions for India

8369
Of your peers have already read this article.
8:30 Minutes
The most insightful time you'll spend today!
In 2013, long before the world was discussing clean energy and sustainable practices, two IIT Madras graduates — Swapnil Jain and Tarun Mehta — had an idea to develop India’s first-ever electrical scooter.
This was at a time when auto manufacturers were still focusing on fossil-fuel-driven vehicles and ‘eco-friendly’ mobility solutions were more a trendy alternative catering to a niche market.
The duo founded Ather Energy in 2013 and launched their first fully-electric scooter, the Ather S340, in Bengaluru in 2016. Since then, the company has released several new models into the market and is planning to expand to eight more cities by the end of the year.
To support the smooth running of their vehicles, lower costs, improve time to market, and create great customer experience, Ather turned to Google Cloud.
Read the Full Story on YourStory
More Relevant Stories for Your Company

Google Cloud Next ’22 to Commence in October: Block Your Calendar!
We’re excited to announce that Google Cloud Next returns on October 11–13, 2022. Join us for keynotes from industry luminaries and engage live with Google developers. Explore dynamic content across various learning levels, and dive deep into technologies and solutions spanning the Google Cloud and Google Workspace portfolios. Participate in breakout sessions,

Case Study: How Texas’ Largest Grocery Chain Successfully Modernized its Legacy Mainframes
H-E-B, like many enterprises, is moving away from legacy mainframes in favor of microservices and public cloud infrastructure. With hundreds of applications powering their 100+ year-old grocery business (with more than 400 stores in Texas and Mexico), H-E-B needs to be confident that the platform they are building will provide

Google Cloud’s Autism Career Program to Nurture Neurodiverse Talent
My passion for neurodiversity began 10 years ago, when I became involved with Els for Autism, an organization that works with children and adults who have autism, as well as their families. At the time, I had a friend who was struggling to find resources for his son with autism. The

How the City of Memphis Uses Technology to Identify 75 Percent More Potholes
At 340 square miles, the City of Memphis is among the largest in the United States in terms of land area. Memphis has over 6,800 lane-miles of city streets, enough to drive back and forth to Los Angeles four times. Keeping these streets well maintained and safe for citizens and visitors is






