Airbus: Taking the flight to a brighter future with Google Cloud and Google Workplace

3561
Of your peers have already read this article.
1:30 Minutes
The most insightful time you'll spend today!
“Any device, anytime, anywhere.” A cohort of CIOs within Airbus believed that the cloud, combined with new ways of working, could provide the foundation for this vision. Google Workspace and Google Cloud have played a pivotal role in helping Airbus realize this new path, transforming security, data management, and collaboration along the way.

Adopting a secure-by-design approach
In adopting Google Workspace and Google Cloud Airbus needed to ensure a robust, zero-trust security model that works across the entire organization, even when employees are working outside the office. Google Workspace provides a single login that enables secure access to data, based on device and user information, as well as contextual inputs that inform the security risk of each login and user action. Airbus admins also use Google Workspace to define trust rules that govern what information and files can be shared within and outside the organization, making it easy for employees to comply with best practices from anywhere.
Encryption also plays a central role in keeping information secure and private. By default, Google Workspace uses the latest cryptographic standards to encrypt all data at rest and in transit. Google Workspace also offers client-side encryption, which Airbus uses for their most sensitive projects, giving them authoritative control over their data as the sole owner of their encryption keys.
And to ensure the organization is protected against hackers, Airbus has implemented sharding as a standard practice, thereby splitting data across multiple servers and data centers. Because the company works with incredibly sensitive information—including government and military information—the ability to locate data all within European data centers continues to be a necessity.
Powerful data management
Managing an enormous volume and variety of data, Airbus needs to ensure complete compliance with internal policies, as well as with external standards, like the General Data Protection Regulation (GDPR). Given this context, Airbus requires a solution that has strong built-in governance controls. Airbus also leverages the Drive labels feature, along with manual classification, to ensure that every file added to Google Drive is tagged and labeled correctly. In turn, these labels define the loss-prevention policies assigned to each file.
Staying connected during the pandemic
Before the pandemic, nearly every employee spent their workdays at an Airbus facility. When remote work became mandatory, the company made the pivot to Google Chat and Google Meet as an essential part of supporting real-time and asynchronous collaboration. Gmail also played a significant role in secure, anywhere-anytime communication, with its built-in anti-spam and anti-malware protections. Customizable filters let administrators protect against suspicious attachments, untrustworthy links, and countless other forms of malicious content. While Gmail blocks more than 99.9% of spam and phishing messages from ever reaching users’ inboxes, more advanced security measures like sandboxing can be put in place for specific use cases.
Protecting more than just data
Google Cloud’s sustainability efforts are as equally important to Airbus as data security. Google Cloud has been working to keep its climate footprint and those who use its services as low as possible, with all of Google currently carbon neutral and with a goal to run on carbon-free energy 24/7 at all of our data centers by 2030. And with our smarter, more efficient data centers, we’re already on that path with more than six times the computing power for the same amount of electrical power we used 5 years ago.
Building the future of work
By combining Google Workspace with Google Cloud, Airbus has been able to live up to its vision of “any device, anytime, anywhere.” The new model is a core foundation for its evolving future of work. Not only has Airbus adopted a zero-trust model across the organization, it’s also transformed how data is secured, managed, and accessed by employees working across a broad range of locations. The new flexible approach has also led to changes in how collaboration happens. Google is deeply gratified to have supported Airbus as they implement these changes and we continue to be proud that we’re running the cleanest cloud in the industry.
Want to learn more about how Google Workspace helps businesses like yours do more while keeping your data secure? Read this whitepaper to find out about our zero-trust model and other ways we protect organizations.
Google Announces New Cloud Region in Toronto

3403
Of your peers have already read this article.
1:30 Minutes
The most insightful time you'll spend today!
For over a decade, we’ve been investing in Canada to become a go-to cloud partner for organizations across the country. Whether they’re in financial services, media and entertainment, retail, telecommunications or the public sector, a rapidly growing number of organizations located or operating in Canada are choosing Google Cloud to help them build applications better and faster, store data, and deliver awesome experiences to their users, all on the cleanest cloud in the industry. To support this growing customer base, we’re excited to announce that the new Google Cloud region in Toronto is now open.
As you’d expect, we’re thrilled about this news, but we aren’t the only ones that have been looking forward to this launch. We asked some of our customers operating in Canada for their take on the upcoming cloud region. Here’s what they had to say:
“Our alliance with Google is truly distinctive in the Canadian market as we are working together to co-innovate and create new services for key industries, including communications technology, healthcare, agriculture, security, and the connected home. The new cloud region in Toronto marks another key milestone that will propel TELUS’ digital leadership by further leveraging the scalability, reliability and cost effectiveness of Google Cloud to support improved customer experience and build stronger, healthier and more sustainable communities.”—Hesham Fahmy, Chief Development Officer, TELUS
“We’re simplifying, modernizing and digitizing Scotiabank to enhance the customer experience for our 25 million customers across the globe. By leveraging powerful cloud-based services including Google Cloud, we’re able to put the most advanced software engineering, data analytics and machine learning tools in the hands of our talented employees. We welcome Google Cloud’s investment in Toronto and look forward to the opportunities the Toronto Cloud Region will present to our Technology team.”
—Michael Zerbs, Group Head Technology & Operations, Scotiabank
“Cloud technologies—and the access to scalable compute, rich geospatial datasets and smart analytics tools—will be critical contributors to support climate action and sustainable policy decisions. At Natural Resources Canada, scientists and researchers are applying innovative digital solutions to support Canada’s natural resource sector. The new Google Cloud region in Toronto will provide our scientists, technologists and researchers with the products and services necessary to turn Earth data into actionable insights.”
—Vik Pant, PhD, Chief Scientist and Chief Science Advisor, Natural Resources Canada
“At Accenture, we bring together technology and human ingenuity to create and respond to change. We’re thrilled to join forces with Google Cloud and their newest region in Toronto with an important mutual goal: to accelerate cloud innovation in Canada. Our clients already know us for our deep industry intelligence, cloud-first expertise and market-renowned delivery. We’re now combining that with Google’s human-centric design to bring even more opportunities to our clients across all industries.”
—Jeffrey Russell, President of Accenture in Canada.
“We are thrilled to see Google’s commitment to Canada. We look forward to helping our joint customers transform their operations, leveraging Google Cloud’s latest data center in Toronto. At Deloitte, we believe cloud is THE opportunity to reimagine everything.”
—Terry Stuart, Deloitte Chief Digital Officer, Canada.
“As Canadian organizations increasingly leverage cloud to transform their businesses, we are excited about the new opportunities that the Toronto Google Cloud region brings to the market. We look forward to continuing our strong partnership with Google Cloud to bring customized and innovative solutions that help Canadian companies fully realize the value of cloud technology, so that they can compete and win on the global stage.”
—Andrew Caprara, Chief Operating Officer, Softchoice
Toronto joins 27 existing Google Cloud regions connected via our high-performance network, helping customers better serve their users and customers throughout the globe. In combination with our Montreal region, customers now benefit from improved business continuity planning with distributed, secure infrastructure needed to meet IT and business requirements for disaster recovery, while maintaining data sovereignty.

The new region launches with three zones, allowing organizations of all sizes and industries to distribute apps and storage to protect against service disruptions, and with our core portfolio of Google Cloud Platform products, including Compute Engine, App Engine, Google Kubernetes Engine, Bigtable, Spanner, and BigQuery.
We’re working to bring you new cloud products and capabilities in Canada, and our goal is to allow you to access those services quickly and easily—wherever you might be in the country. The past year has proved how important easy access to digital infrastructure, technical education, training and support are to helping businesses respond to the pandemic. We’re particularly proud of the teams who faced the unique challenges of building a cloud region during this time to help our customers and community accelerate their digital transformation.
To support all of our users, customers and government organizations in Canada, we’ll continue to invest in new infrastructure, engineering support and solutions. We’re currently hosting our first ever Google Cloud Accelerator Canada to bring the best of Google’s programs, products, people and technology to startups doing interesting work in the cloud. We’ve recently received Protected B accreditation with Canadian Centre for Cyber Security, which is crucial for healthcare, education, and regulated industries adopting cloud services. We’re also pleased to announce the preview of Assured Workloads for Canada—a capability which allows you to secure and configure sensitive workloads in accordance with your specific regulatory or policy requirements.
For help migrating to Google Cloud, please contact our local partners. For additional details on Google Cloud regions, please visit our locations page, where you’ll find updates on the availability of additional services and regions. You can always contact us to help you get started or access our many educational resources. We’re excited to see what you build next with Google Cloud.
How a Mid-Sized Firm is Shrinking Inventory Carryovers 50% with AI

7431
Of your peers have already read this article.
5:30 Minutes
The most insightful time you'll spend today!
For more than 10 years, NMK Textile Mills has manufactured bed linens for major retailers in the United States and Canada. As e-commerce exploded, the company’s co-founder saw an opportunity to grow the business. So, in addition to manufacturing bed linens wholesale for retail customers, NMK Textile Mills reworked its complete supply chain to manufacture products for California Design Den, which became an e-commerce retailer selling its own fashion-forward products directly to consumers online.
With California Design Den’s push into e-commerce, it became apparent the SMB company (with about 250 global employees) faced the same tough supply chain questions as large retail customers, including maintaining enough inventory to meet customer demand in a timely, efficient way.
California Design Den depended upon a myriad of systems to track its complex forecasting and reordering processes. Team members typically planned inventory manually using desktop spreadsheet software, which could lead to excess inventory. Accurately forecasting demand and supply was essential to the company’s financial success—but it was also a challenge.
A couple years ago, California Design Den partnered with Pluto7, a technology solutions provider that offers a software as a service (SaaS) called Planning In A Box. Leveraging Google Cloud Platform machine learning and artificial intelligence, Planning In A Box intelligently helps predict demand and balances it with supply.
But that was just the beginning. “Along the way, we realized that to compete with larger retailers, make quicker decisions, and move faster, we needed to go further,” says Deepak Mehrotra, Co-founder and Chief Adventurer at California Design Den.
With guidance from Pluto7, California Design Den began migrating its database to Google Cloud Platform. Using Google BigQuery, Google Compute Engine, Google Cloud SQL, and Google Cloud Storage, and experimenting with Google Cloud Vision and Google Cloud AutoML, the company is reducing inventory carryovers by more than 50%, improving the accuracy of demand planning quarter over quarter, and gaining granular insights into how individual SKUs are performing.
“We would need an army of data scientists to make faster decisions on pricing and inventory levels. With Google Cloud Platform machine learning and artificial intelligence, we don’t need that. We can make much faster pricing decisions to optimize profitability and move inventory.”
—Deepak Mehrotra, Co-founder and Chief Adventurer, California Design Den
No need for army of data scientists
“Using Google Cloud Platform machine learning and AI was essential for California Design Den if it was to compete successfully with larger retailers,” Deepak says.
For example, tastes and fashions in bed linens can change quickly and consumer prices fluctuate. With more than 2,500 SKUs, it wasn’t possible for California Design Den’s team to continuously monitor product demand and experiment with competitive pricing in real time.
“We would need an army of data scientists to make faster decisions on pricing and inventory levels,” says Deepak. “With Google Cloud Platform machine learning and artificial intelligence, we don’t need that. We can make much faster pricing decisions to optimize profitability and move inventory.”
By integrating all its data onto Google Cloud Platform, California Design Den’s team gains deeper insights into product sales over time, which in turn helps the company improve demand planning by better determining which styles to manufacture and sell in the future.
“When experienced employees leave, their knowledge leaves with them. By having all data in one place, and with machine learning and AI, California Design Den can go back in its history, look at products made or sold years ago, and analyze product performance.”
—Deepak Mehrotra, Co-founder and Chief Adventurer, California Design Den
Merging visuals with data
Before Google Cloud Platform, team members had to dig through spreadsheets and run scenarios to get a sense of how particular products had sold. The next step was to perform keyword searches across the company’s photo library in the cloud to find each product’s image. From there, a team member would insert the product images into a presentation, along with relevant data points, to provide a report for stakeholders on how particular styles performed.
Today, California Design Den, with the help of Pluto7, is integrating its entire product image library with its database on Google Cloud Platform. Experimenting with Google Cloud Vision and Google Cloud AutoML, California Design Den is moving towards a day when team members can run sales scenarios and get deep background data on individual product performance while viewing images of the relevant products.
Merging product visuals with data will help designers and team members better understand sales patterns over time and in context. In the past, making correlations between things like which sheet colors sold well in California, compared to how the same sheet colors performed on the East Coast, was something that California Design Den employees primarily did in their heads.
“When experienced employees leave, their knowledge goes with them,” says Manjunath Devadas, Founder and CEO at Pluto7. “By having all data in one place, and with machine learning and AI, California Design Den can go back in its history, look at products made or sold years ago, and analyze product performance.”
“We are literally growing the complexity of our business on all levels, including designing, manufacturing, selling, reordering, inventory holding—everything.”
—Deepak Mehrotra, Co-founder and Chief Adventurer, California Design Den
Reimagining supply-demand balancing
Pluto7’s mission statement is to democratize supply demand balancing with machine learning and AI. California Design Den is a case in point, as the combination of Planning In A Box and Google Cloud Platform gives the company greater control over its destiny.
“Big retailers used to tell us what to manufacture and how much they would pay for it,” says Deepak. “That was our primary business, and if we didn’t accept the terms, a competitor would.” Today, in addition to continuing to make products for retailers, California Design Den can design, make, and sell a variety of designs for itself, including custom and limited-edition products, thanks to Google Cloud Platform and Pluto7 software offerings. The payoff is not only in having a more diversified business. California Design Den also receives more favorable profit margins by selling its own products.
“We are literally growing the complexity of our business on all levels, including designing, manufacturing, selling, reordering, inventory holding—everything,” Deepak says. “We can make smaller batches. We can connect directly with consumers. We can identify the missing pieces—what should we produce next, when, and how much? We otherwise couldn’t afford the level of talent it would take to do this.”
Cutting through the noise
Google and Pluto7 software helped California Design Den reduce inventory carryovers by more than 50%. Inventory tracking and distribution, along with insights and visibility into product sales, are faster, more efficient, and accurate. Google Cloud Platform flexible pricing, speed, reliability, security, and scalability enable California Design Den to stay relevant and be more competitive.
In addition to benefiting from Google Cloud Platform, California Design Den relies on G Suite—also part of Google Cloud—to enhance collaboration among its global teams. Previously, the company’s email server would sometimes crash, due to the heavy load of sharing product photos and other data. “Gmail and Google Drive handle the everyday demands on the business effortlessly and reliably,” Deepak says.
“Google machine learning and AI enable us to cut through all the noise from raw data, so we can see what’s important. We can focus on analytics to guide us to success today and in the future.”
—Deepak Mehrotra, Co-founder and Chief Adventurer, California Design Den
The company is exploring additional ways to leverage Google Cloud Platform in the near future. For example, one possibility is to import customer reviews from sites where products are sold into Google BigQuery, and to use that data to perform sentiment analysis via Google Cloud Natural Language. It could provide another valuable data source to help California Design Den’s team decide where to focus future designs.
“Google machine learning and AI enable us to cut through all the noise from raw data, so we can see what’s important,” Deepak says. “We can focus on analytics to guide us to success today and in the future.”
Unlocking Economic Potential: Cloud FinOps

3249
Of your peers have already read this article.
4:00 Minutes
The most insightful time you'll spend today!
Built for a CapEx world, most organizations’ finance systems aren’t set up to take advantage of cloud’s dynamic, OpEx-driven consumption patterns.

GETTY
It may not be a household name yet, but chances are you’ve crossed paths with OpenX today. OpenX, a leader in programmatic advertising, operates one of the world’s largest ad exchanges, serving over 250 billion ad requests per day, connecting more than 30,000 brands and reaching nearly one billion consumers. To make it happen, in 2019, OpenX migrated entirely out of its data centers and became the first major ad-exchange platform to move completely to the cloud.
The OpenX CTO, Paul Ryan, knew that this cloud transformation initiative had the potential to increase costs faster than its revenues. To be successful, he needed his engineering, finance, and business teams to forge a new “cost-aware” culture, complete with effective cost visibility and controls. In other words, he needed Cloud FinOps — an operational framework and cultural shift that brings technology, finance, and business together to drive financial accountability and realize business benefits through cloud transformation.
Ryan laid out a cloud migration roadmap that included cost governance and controls around project ownership, established cost responsibility with engineering teams to accurately forecast cloud consumption, and challenged developers to lower per-unit costs — while at the same time improving performance, scalability, speed and global reach.
It worked! In just 9 months, OpenX reduced their per-unit cost by over 60%. The framework allowed OpenX to launch new regions in a matter of days, reduce their time to market for new features by over 50%, and complete their migration in record time — seven months! “We are now able to stop worrying about legacy infrastructure and focus more on our growth categories,” said Ryan. “Our tech stack is getting smarter and more sophisticated by the day, and we have the flexibility to scale our infrastructure in real-time as the business scales and evolves.”
Unblocking Cloud’s Potential
Cloud holds the key to a successful digital transformation. In fact, McKinsey forecasts that by 2030, the Fortune 500 alone may realize over $1 trillion of EBITDA value drivers associated with public cloud enablement. But unlike OpenX, many companies struggle to achieve near-term value objectives from their cloud investments. Surveys reflect that more than 30% of cloud spend in 2021 was wasted or inefficient, while upwards of 80% of CIOs have yet to achieve the business benefits of migrating to the cloud.
Traditional IT finance processes are ill-suited for cloud infrastructure: Traditional planning and budgeting processes are challenged to address dynamic consumption patterns and complex migrations. Centralized IT budgets using traditional allocations fail to provide the necessary visibility into sources of cost overruns. CapEx-focused cost controls have little ability to manage largely OpEx-driven spend. Trend-based forecasting often inaccurately predicts cloud costs. And developer teams lack access to cost-aware architecture patterns to deploy the applications more efficiently.
Enter Cloud FinOps
At Google we’ve worked with many companies, like OpenX, to help organizations realize the transformational benefits of the cloud by cultivating a culture of transparency and embedding agile processes to manage costs. We’ve distilled these learnings into a Cloud FinOps operational framework that gives organizations the financial governance and accountability they need to grow their business sustainably.

GOOGLE CLOUD
At a high level, a Cloud FinOps approach depends on five key areas:
- Accountability and Enablement
Accountability and enablement aim at instilling a cost-conscious culture across the organization. Oftentimes, this means standing up a cross-functional and dedicated team with members from technology, finance and engineering to establish cloud financial best practices and governance. In various organizations, we’ve seen this through an extension of a Cloud Center of Excellence, a Cloud Business Office or simply a Cloud FinOps team. Enablement focuses on empowering IT, finance and business leaders through training to help them better understand the economics of cloud services and the strategies to efficiently deploy and manage them. Cloud financial training guides teams on how to design cost-effective cloud environments, for example, embracing ”cloud-native” design principles such as auto-scaling/elasticity and Infrastructure as a Code.
- Measurement and Business Value Realization
Effective measurements not only create awareness and enable agile processes, but also support a culture that celebrates success and rewards teams for achieving business objectives. As such, measurement in the service of business value realization is about developing a comprehensive set of long-term benefits and cost KPIs to quantify the total net value of the return on digital transformation. Organizations often start with cost-related KPIs and eventually evolve those KPIs into business value metrics that are mapped to targeted business outcomes.
- Cloud Cost Optimization
Cloud cost optimization is an iterative and continuous process that provides a consistent methodology to manage cloud consumption cost-effectively. There are three key areas of optimization:
Resource optimization – Model cost-effective cloud usage based on utilization and consumption patterns.
Pricing optimization – Manage cloud spend through a continuous analysis of various pricing models. In a Google Cloud context, that might mean Committed Use Discounts, BigQuery flat rate reservations, etc.
Architecture optimization – Build applications with a cost-aware architecture by leveraging newer generation compute instances (like Tau VMs, which offer an industry-leading 42% better price-performance versus comparable offerings), or using managed services and serverless technology to offload operational overhead.
For example, video hosting, sharing and services platform provider Vimeo built transcoding pipelines by using Google Cloud Spot VMs to optimize their infrastructure spend. To do so, they created fault-tolerant workloads that could withstand preemptions, and in exchange, got up to a 91% discount compared to using regular on-demand instances.
- Planning and Forecasting
In the cloud, accurately forecasting your finances requires rethinking of traditional approaches to depreciation and trend-based forecasting of maintenance and licensing costs. One way to improve the accuracy of your dynamic cloud needs is to use workload-specific forecasting models that leverage a combination of trend-based models for steady-state workloads, driver-based models for scaling applications, as well as monthly variance analysis. In other words, you can define cloud budgets and forecasts by monitoring cloud consumption trends, allocating cloud cost pools with a proper tagging strategy that’s mapped to a chart of accounts in a general ledger, and conducting a cost-benefit analysis based on cloud infrastructure, implementation, and support costs.
- Tools and Accelerators
Without the proper tools and processes in place, understanding and managing cloud costs can be complex — and this especially true as organizations scale their business in the cloud. By deploying proper cloud cost management tools and accelerators such as Looker Cloud Cost Management and automation scripts to set guardrails and enforce cost control policies, organizations can effectively manage and track cloud spend with access to near-real-time billing and cost data to make better informed business decisions.
The key objectives of Google Cloud Cost Management tools are to make it as simple as possible for organizations to get visibility into their current and forecasted costs with built-in reporting and customizable dashboards; help drive greater accountability for cloud spending across the organization by providing flexible ways to organize cloud resources and allocate costs; provide strong financial governance controls to reduce the risk of overspending; and offer intelligent recommendations for optimizing cloud costs and usage.
Start Saving with Cloud
Businesses are continuously seeking to better operate and manage their cloud environments and the need is ever increasing to transparently manage cloud spend, optimize costs, and obtain their desired business agility. By enhancing your Cloud FinOps capabilities and adopting principles of continuous cost optimization, you too can accelerate the business value of cloud computing.
Special thanks to Bruce Warner, Daniel Petibone, Nihar Jhawar and FinOps Foundation community for their contributions and sharing their domain expertise to this important Cloud FinOps topic.
How Lowe’s SRE Team Decreases Mean-time-to-recovery (MTTR)

3361
Of your peers have already read this article.
1:30 Minutes
The most insightful time you'll spend today!
Editor’s Note: In a previous blog, we discussed how home improvement retailer Lowe’s was able to increase the number of releases it supports by adopting Google’s Site Reliability Engineering (SRE) framework on Google Cloud. Lowe’s went from one release every two weeks to 20+ releases daily, helping meet its customer needs faster and more effectively. Today, the Lowe’s SRE team shares how they used SRE principles to decrease their mean-time-to-recovery (MTTR) by over 80 percent.
The stakes of managing Lowes.com have never been higher, and that means spotting, troubleshooting and recovering from incidents as quickly as possible, so that customers can continue to do business on our site.
To do that, it’s crucial to have solid incident engineering practices in place. Resolving an incident means mitigating the impact and/or restoring the service to its previous condition. The average time it takes to do this is called mean time to recovery (MTTR). Tracking this metric helps us stay on top of the overall reliability of our systems at Lowe’s, while simultaneously improving the speed with which we recover. Our goal is to keep the MTTR metric as low as possible, so that failures don’t negatively impact our business. Here are the four areas we addressed to drive holistic improvement in our MTTR.
Lowe’s incident reporting process
To reduce MTTR, we created a seamless incident reporting process following SRE principles. Our incident reporting process is a workflow that starts at the time an incident occurs, and ends with an SRE captain who closes the action items after a postmortem report. With this approach, we are able to limit the number of critical incidents. The reporting process involves three core components: monitoring, alerting, and blameless postmortems.
Monitoring and alerting
Having proper monitoring and alerting in place is crucial when it comes to incident management. Monitoring and alerting tools let you detect issues as soon as they occur, and notify the right person in the shortest possible time to take action. From a measurement standpoint, we track this as our mean time to acknowledge (MTTA). This is the average time it takes from when an alert is triggered, to when work on the issue begins.
At the time of an incident, our monitoring and alerting tools notify the on-call SRE first responder via PagerDuty in the form of a phone call, text message and email. Our SRE software engineering team has done a lot of automation to enable various Service Level Indicator (SLI) alerts and Service Level Agreement (SLA) notifications. The on-call SRE then initiates a triage call with our service/domain stakeholders to resolve the incident. As a result, we reduced our MTTA from 30 minutes in 2019, to one minute – a 97 percent decrease.
Blameless postmortems: learning from incidents
A postmortem is a written record of an incident, its impact, the actions taken to resolve it, the root cause and the follow-up actions to prevent the incident from recurring (see example here). A blameless postmortem builds on that and is a core part of an SRE culture, and our culture at Lowe’s. We ensure that individuals are not singled out, and the outcome for all postmortems are directed toward learnings and process improvement.
For us, the postmortem process is the biggest part of our incident workflow. When an SRE creates a new postmortem report, the first step is to conduct a postmortem session with domain stakeholders to review the report. The postmortem then goes into the review stage and gets reviewed by more stakeholders in our weekly postmortem meeting. In the final stage of this process, the SRE captain will close the report once everyone in the weekly meeting agrees that the report is complete.
To conduct a successful postmortem, it is critical to keep the focus on identifying gaps and issues with the system and operations processes, rather than an individual, and generate concrete actions to address the problems we’ve identified. To ensure this, we follow a couple of best practices:
- We start by gathering the facts from the person who identified the problem, and each SLI owner has to identify a gap or the next SLI upstream owner who created the impact for them.
- Every SLI owner is provided full opportunity to present their case, and identifying the issue is done as a community exercise.
- Once action items and process changes are identified, an owner is nominated to complete the actions, or they will volunteer.
- For easy reference, we publish and store postmortems in our incident knowledge base. This process helps SREs continuously improve as future incidents arise.
Continuous Improvement
Encouraging a culture of honest, transparent and direct feedback that you need for blameless postmortems is often an iterative process that needs sponsorship from executives, empowering incident captains to lead the entirety of the discussion and outcomes. Running successful postmortems, and completing action items from them, needs to be recognized and accounted for in SRE performance objective assessment. As shared in Google’s SRE book, the best practice is to ensure that writing effective postmortems is a rewarded and celebrated practice, with leadership’s acknowledgement and participation. This is possibly the hardest part to accomplish in an effective postmortem during a cultural transformation unless you have full buy-in from leadership.
However, it’s all well worth it. This process is a key part of how we were able to improve our MTTR over time—from two hours in 2019 to just 17 minutes!
Our SRE incident reporting process has also transformed how our company solves issues. By streamlining this workflow from alerting, to solving an issue, to blameless postmortems, we have reduced our MTTR by 82 percent and our MTTA by 97 percent. Most importantly, our team is learning from every incident and becoming better engineers as a result. Visit the SRE Google Cloud website to learn more about implementing SRE best practices in the cloud.
Acknowledgement
Special thanks to Rahul Mohan Kola Kandy, Vivek Balivada, and the Digital SRE team at Lowe’s for contributing to this blog post.
An Expert’s Opinion on What Early-stage Startups Must Know

6443
Of your peers have already read this article.
3:00 Minutes
The most insightful time you'll spend today!
As lead for analytics and AI solutions at Google Cloud, my team works with startups building on Google Cloud. This puts us in the fortunate position to learn from founders and engineers about how early-stage startups’ investments can either constrain them or position them for success, even at the seed level. In this post, I want to share a few of the best practices to keep in mind as you’re building.
Understand your value proposition before diving into a technology stack
If you’re launching a startup in the cloud, you’re no doubt thinking about a technology stack, but it’s important to step back a bit and think carefully about the major value proposition that your startup offers to your customers. That value proposition is going to fundamentally drive the kind of technology that you should pick.
For example, does your system need processing in real time, or can it be done in a batch mode? Can you rely on once-a-day insights or do the insights have to come in as events happen?
Additionally, what kind of latency will your customers face? That latency makes your value proposition either usable or unusable. Early on in Google’s development, leaders realized that no one was going to wait more than a few hundred milliseconds for a web page to show them their results, and that realization drove the technology decisions that have allowed Google to scale from being a startup in a garage to being a trillion dollar company. Your startup needs to define its value to customers with this level of specificity before it can build a technology stack suited to its needs.
Focus on customer interactions
A few companies have gracefully pulled off big IT pivots that reshaped their value proposition. Netflix, for example, moved from mostly sending DVDs through the mail to becoming a streaming service and major content producer. That’s a huge shift in the user experience and the technology stack necessary to support it, even if the underlying value proposition (i.e., get content to customers) was broadly the same. But it’s also an outlier. If you’re planning for potential changes of this magnitude, rather than focused on getting your value proposition to users, you probably need to sharpen what that value proposition is.
Specifically, you need a clear vision of how customers will access and interact with your business. Typically, they’ll do so over a website or a mobile app, but there are still so many variables.
Are customers going to transmit documents? If so, in what format? Is handwriting supported or is input limited to typing? Can they use images for optical character recognition? Will it mostly be forms? Will the data be structured or unstructured? If all that sounds a little overwhelming, don’t worry, it’ll seem simpler by the end of this article—but also be aware: we’re just getting warmed up.
Imagine that most of your customers will access your business via voice, so you know you’ll want to prioritize conversational workflows. That’s a start—but dig deeper. Even if we suppose you’re usingDialogflow, a Google Cloud conversational AI platform that lets you build and deploy virtual agents, we’re still not really seeing the value proposition. How will all this work, from the beginning of a typical full customer interaction to the resolution? How many interactions will have to be facilitated over low-bandwidth connections, for example? When it comes to user interactions, make sure you can see an end-to-end use case.
Another example: you’re building a retail website, and one of your end-to-end use cases involves the customer asking if a certain amount of a given product is in stock, whether it’s one unit of the product, ten or hundreds. If the product is not sufficiently stocked, you want your app to offer similar items that are. Will your technology stack support this end-to-end use case?
These considerations are not an argument for premature optimization. There’s value in moving fast, getting minimum viable products to users, and then iterating. But in the early stages, you only get one chance to start on the right foot—and how you navigate that chance will influence a lot of dollars and effort down the road. You need to make sure you have business use cases, not just an idea, before you can start designing a technology stack.
Here’s how to get in the right frame of mind. Pick three use cases: two that are “bread and butter” and one that is technologically complex. Make sure your proposed technology stack can support all three, end to end.
Default toward higher levels of abstraction
Now that we’re in the right frame of mind, we’re ready to think about the technology stack more directly.
As a startup, you’ll need to conserve resources, and to do that, you’ll want to build at the highest level of abstraction possible for your value proposition. For example, you probably don’t want your people setting up clusters. You don’t want them configuring things if they can use a fully managed service. You want them focused on building your prototype, not managing infrastructure.

This focus has definitely informed how we create products at Google Cloud, as our canonical data stack—Pub/Sub, Dataflow, BigQuery, and Vertex AI—consists of auto-scaling and serverless products.
But management of infrastructure is not the only place where you should err toward a less-is-more philosophy.
When it comes to architecture, choose no-code over low-code and low-code over writing custom code. For example, rather than writing ETL pipelines to transform the data you need before you land it into BigQuery, you could use pre-built connectors to directly land the raw data into BigQuery. That’s no code right there. Then, transform the data into the form you need using SQL views directly in the data warehouse. This is called ELT, and it is low code. You will be a lot more agile if you choose an ELT approach over an ETL approach.
Another place is when you choose your ML modeling framework. Don’t start with custom TensorFlow models. Start with AutoML. That’s no-code. You can invoke AutoML directly from BigQuery, avoiding the need to build complex data and ML pipelines. If necessary, move on to pre-built models from TensorFlow Hub, HuggingFace, etc. That’s low-code. Build your own custom ML models only as a last resort.

Focus on getting your vision to market, not chasing technology hype
The goal is to pick the right technology stack for bringing your vision to market, generating value for customers, conserving resources, and maintaining flexibility for growth. Early IT investments should usually gravitate toward things that preserve flexibility, such as managed services built on standard protocols or open APIs, but they needn’t always rush to the flashiest technologies. The answer isn’t always ML, for example. The answer might be heuristics to start, with a path to ML once you have collected enough data. You want to make sure that your intelligence layer has enough abstraction so you can mark it up with simple rules at first, but then replace it with a more robust system as you go along.
Launch and iterate fast with these principles
The preceding discussion is a reminder that your most expensive resource is your people—and that you really want them to be focused on building your prototype, minimum viable product or production app You want to launch fast and iterate fast, and the only way you can do that is by focusing on the things that differentiate you.
But regardless of the technologies you use, the bottom line is the same: follow these four principles.
- Figure out your major value proposition and design your tech stack around it.
- Be very careful about user interactions. User experience is super important; you need to make sure you deliver the kind of experience that your customers have grown to expect.
- When you’re building, pick the highest possible level of abstraction possible—the most fully managed tools and no-code/low-code frameworks that give you the functionality that you need.
- Instead of choosing new or flashy technologies, consider if you can build a “good enough” minimum viable product quickly and come back to a better implementation later.
To learn more about why startups are choosing Google Cloud, click here.
More Relevant Stories for Your Company

Taking Partnership forward: Google Cloud VMware Engine Now in VMware Cloud Universal
As the pace of digital transformation accelerates, our partnership with VMware continues to focus on helping customers successfully navigate their cloud journey and achieve their business objectives through seamless and rapid migration of business critical VMware workloads. We announced the general availability of Google Cloud VMware Engine in May 2020.

The Transformative Journeys of Financial Firms on Google Cloud: Watch Video
The reliance on cloud-based architectures, high performance computing, big data and more are accelerating in the banking, capital, insurance and financial services industries. Google Cloud had a strong role in transforming many businesses especially in the pandemic to smoothly transition into the digital space and understand their customers. Two years

Lucent Bio: Boosting collaboration and sustainability with Google Workspace
From electrifying transportation to shifting the grid to renewable energy, environmental sustainability is one of the greatest challenges of our generation. A critical but often forgotten goal is the development of sustainable agricultural practices, especially given increasing water shortages and soil degradation around the world. Lucent Bio was born to

Google Cloud’s Role in Minimizing Memory Errors Impact for SAP Customers
Every cloud system begins with high-quality hardware infrastructure. Sometimes, however, hardware breaks — and when it happens, our most important goal is to minimize the impact on our customers and their cloud workloads. Memory errors are the most common type of hardware failure, and they're also one of the most






