Rubin Observatory Leverages Google Cloud to Power Astronomical Research - Build What's Next
Case Study

Rubin Observatory Leverages Google Cloud to Power Astronomical Research

6060

Of your peers have already read this article.

2:00 Minutes

The most insightful time you'll spend today!

The Vera C. Rubin Observatory launched its the first preview of new Rubin Science Platform (RSP) for an initial cohort of astronomers, powered by Cloud Storage, Google Kubernetes Engine (GKE), and Compute Engine to drive astronomical researches.

This week, the Vera C. Rubin Observatory is launching the first preview of its new Rubin Science Platform (RSP) for an initial cohort of astronomers. The observatory, which is located in Chile but managed by the U.S. National Science Foundation’s NOIRLab in Tucson, AZ and SLAC in California, is jointly funded by the NSF and the U.S. Department of Energy. The platform provides an easy-to-use interface to store and analyze the massive datasets of the Legacy Survey of Space and Time (LSST), which will survey a third of the sky each night for ten years, detecting billions of stars and galaxies, and millions of supernovae, variable stars, and small bodies in our Solar System.

The LSST datasets are unprecedented in size and complexity, and will be far too large for scientists to download to their personal computers for analysis. Instead, scientists will use the RSP to process, query, visualize, and analyze the LSST data archives through a mixture of web portal, notebook, and other virtual data analysis services. An initial launch with simulated data, called Data Preview 0, builds on the Rubin Observatory’s three-year partnership with Google to develop an Interim Data Facility (IDF) on Google Cloud to prototype hosting of the massive LSST dataset. This agreement marks the first time a cloud-based data facility has been used for an astronomy application of this magnitude.

Bringing the stars to the cloud

For Data Preview 0, the IDF leverages Cloud StorageGoogle Kubernetes Engine (GKE), and Compute Engine to provide the Rubin Observatory user community access to simulated LSST data in an early version of the RSP. The simulated data were developed over several years by the LSST Dark Energy Science Collaboration to imitate five years of an LSST-like survey over 300 square degrees of the sky (about 1,500 times the area of the moon). The resulting images are very realistic: they have the same instrumental characteristics, such as pixel size and sensitivity to photons, that are expected from the Rubin Observatory’s LSST Camera, and they were processed with an early version of the LSST Science Pipelines that will eventually be used to process LSST data. “This will be the first time that these workloads have ever been hosted in a cloud environment. Researchers will have an opportunity to explore an early version of this platform,” says Ranpal Gill, senior manager and head of communications at the Rubin Observatory.

Broadening access for more researchers

Over 200 scientists and students with Rubin Observatory data rights were selected to participate in Data Preview 0 from a pool of applicants that represents a wide range of demographic criteria, regions, and experience level. Participants will be supported with resources such as tutorials, seminars, communication channels, and networking opportunities—and they will be free to pursue their own science at their own pace using the data in the RSP. 

“The revolutionary nature of the future LSST dataset requires a commensurately innovative system for data access and analysis paired with robust support for scientists,” says Melissa Graham, lead community scientist for the Rubin Observatory and research scientist in the astronomy department at the University of Washington. “I’m personally excited to enhance my own skills by using the RSP’s tools for big data analysis, while also helping others to learn and to pursue their LSST-related science goals during Data Preview 0.” 

At the same time, the fact that the RSP is hosted in the cloud provides researchers at smaller institutions access to state-of-the-art astronomy infrastructure that is comparable to that of the largest national research centers.

The launch benefits the observatory too: the development team can learn what researchers are interested in while also testing and debugging the platform. Graham says that “the platform is still in active development so researchers using it will be able to follow along in the progress, and provide feedback on ways that we can optimize the development of the tools.”

Next steps

The LSST aims to begin the ten-year survey in 2023-24 and expects it to include 500 petabytes of data. Through the cloud, Google aims to help make this extraordinary project scalable and accessible to researchers everywhere. To learn more about Data Preview 0, watch this video.


Want to ramp up your own research in the cloud? We offer research credits to academics using Google Cloud for qualifying projects in eligible countries. You can find our application form on Google Cloud’s website or contact our sales team.

Whitepaper

Cloud as an Innovation Platform in Capital Markets

DOWNLOAD WHITEPAPER

5638

Of your peers have already downloaded this article

5:30 Minutes

The most insightful time you'll spend today!

Public cloud, big data, and AI technologies offer competitive advantages and cost savings for capital markets firms ready to make the transition. This paper discusses the three phases capital markets firms go through in transitioning to public cloud, and the workloads, benefits, and cultural changes that characterize the three phases:

Infrastructure Optimizers: The first step on the public cloud journey, where firms focus on migrating specific workloads to save costs.

Cautious Strategists: Firms build on the success of their first public cloud migrations, and begin to change the way they develop technology to increase cost savings and start taking advantage of capabilities only available on public cloud.

Transformative Innovators: Firms shift to a fully public cloud-enabled mentality, and fully leverage the flexibility and agility of the public cloud to build industry-changing solutions and attract top IT talent.

Additionally, we reveal the five things that capital markets innovators who have advanced to the transformation phase do well in their adoption of cloud, big data, and AI technologies across the front, middle, and back office functions.

Blog

Google Cloud’s Data Analytics May Recap

6899

Of your peers have already read this article.

5:00 Minutes

The most insightful time you'll spend today!

Apart from the inaugural Data Cloud Summit, Google Cloud's Data Analytics and Management solutions have made waves with recognition as a leader in Cloud Data Warehouse and Streaming Analytics domain, new innovations and releases. Read what's next!

May was a very busy month for data analytics product innovation. If you didn’t have the chance to attend our inaugural Data Cloud Summit, video replays of all our sessions are now available so feel free to watch them at your own pace. 

In this blog, I’d like to share some background behind the innovations we released in May, why we built them the way we did, and the type of value they can bring your company and your team.

But first, a huge thank you!

This week, we had the honor to announce that Google has been named a Leader in The Forrester Wave™: Streaming Analytics, Q2 2021 report. Forrester gave Dataflow a score of 5 out of 5 across 12 different criteria, stating: “Google Cloud Dataflow has strengths in data sequencing, advanced analytics, performance, and high-availability”. 

Google has more than a decade of experience in building real-time and internet-scale systems for its own needs, and we are excited to see that our ability to provide customers with a reliable, scalable, and performant platform is bearing fruit. 

This announcement comes on the back of the release of The Forrester Wave™: Cloud Data Warehouse, Q1 2021 report, which also named Google Cloud as a Leader.

We couldn’t be more excited about the recognition and appreciate all your feedback and trust in the work that we do to support your goal in accelerating data-powered innovation.

Innovation galore

Your feedback and your passion is the fuel that drives our ambition to deliver more and better services to you. That’s why, this year, we didn’t want to wait until Google Cloud Next to share some great products we have been working on. On May 26, our team announced a slew of new products, services and programs. Watch a quick summary below:

https://youtube.com/watch?v=DG1mOPMXJvw%3Fenablejsapi%3D1%26

Meeting you where you are

An important design principle behind all of our services is “meeting you where you are”. This means we aim to provide you with the tools and software you need to innovate on your own terms. Here are three new services that will help you do just that:

Datastream

Datastream, our new serverless change data capture (CDC) and replication service, allows your company to synchronize data across heterogeneous databases, storage systems, and applications reliably and with minimal latency to support real-time analytics, database replication, and event-driven architectures. Datastream delivers change streams from Oracle and MySQL databases into Google Cloud services such as BigQueryCloud SQLCloud Storage, and Cloud Spanner, saving time and resources while ensuring your data is accurate and up-to-date. 

  • Under the hood, Datastream reads CDC events (inserts, updates, and deletes) from source databases, and writes those events with minimal latency to a data destination. It leverages the fact that each database source has its own CDC log—binlog for MySQL and LogMiner for Oracle—which it uses for its own internal replication and consistency purposes. 
  • Datastream integrates with purpose-built and extensible Dataflow templates to pull the change streams written to Cloud Storage, and create up-to-date replicated tables in BigQuery for analytics. It also leverages Dataflow templates to replicate and synchronize databases into Cloud SQL or Cloud Spanner for database migrations and hybrid cloud configurations. 
  • Datastream also powers a Google-native Oracle connector in Cloud Data Fusion’s new replication feature for easy ETL/ELT pipelining. By delivering change streams directly into Cloud Storage, customers can leverage Datastream to implement modern, event-driven architectures.

Looker and BigQuery Omni on Microsoft Azure

Research on multi cloud adoption is unequivocal — 92% of businesses in 2021 report having a multi cloud strategy. We want to continue supporting your choice by providing the flexibility you need to see your strategy through. 

  • This past month, we introduced Looker, hosted on Microsoft Azure. For the first time, you can now choose Azure, Google Cloud, or AWS for your Looker instance. You can also self-host your Looker instance on-premises.
  • We also introduced BigQuery Omni for Azure, which along with last year’s introduction of BigQuery Omni for AWS, will help you access and securely analyze data across Google Cloud, AWS, and Azure. 

The cost of moving data between cloud providers isn’t sustainable for many, and it’s still difficult to seamlessly work across clouds. BigQuery Omni represents a new way of analyzing data stored in multiple public clouds, which is made possible by BigQuery’s separation of compute and storage. By decoupling these two, BigQuery provides scalable storage that can reside in Google Cloud or other public clouds, and stateless resilient compute that executes standard SQL queries. 

  • Unlike competitors, BigQuery Omni doesn’t require you to move or copy your data from one public cloud to another, where you might incur egress costs. You also benefit from the same BigQuery interface on Google Cloud, enabling you to query data stored in Google Cloud, AWS, and Azure without any cross-cloud movement or copies of data. 
  • BigQuery Omni’s query engine runs the necessary compute on clusters in the same region where your data resides. For example, you can query Google Analytics 360 Ads data stored in Google Cloud and query logs data from your ecommerce platform and applications that are stored in AWS S3 and/or Microsoft Azure. 

Then, using Looker, you can build a dashboard that allows you to visualize your audience behavior and purchases alongside your advertising spend. 

Dataplex

We understand that most organizations still struggle to make high-quality data easily discoverable and accessible for analytics, across multiple silos, to a growing number of people and tools within their organization. 

They are often forced to make tradeoffs. For instance, moving and duplicating data across silos to enable diverse analytics use cases or leaving their data distributed but limiting the agility of decisions. 

  • Dataplex provides an intelligent data fabric that enables you to centrally manage, monitor, and govern your data across data lakes, data warehouses, and data marts, while also ensuring data is securely accessible to a variety of analytics and data science tools. 
  • One of the core tenets of Dataplex is letting you organize and manage your data in a way that makes sense for your business, without data movement or duplication. For that, we provide logical constructs like lakes, data zones, and assets. These constructs enable you to abstract away the underlying storage systems and become the foundation for setting policies around data access, security, lifecycle management, and so on. 
  • For example, you can create a lake per department within your organization (e.g. Retail, Sales, Finance, etc.) and create data zones that map to data readiness and usage (e.g. landing, raw, curated_data_analytics, curated_data_science, etc.). 

Once you have your lakes and zones setup, you can attach data to these zones as assets. You can add data from different types of storage (e.g. GCS Bucket and BigQuery dataset) under the same zone. You can also attach data across multiple projects under the same zone. You can ingest data into your lakes and zones using the tools of your choice, including services such as Dataflow, Data Fusion, Dataproc, Pub/Sub, or choose from one of our partner products. Dataplex comes with built-in 1-click templates for common data management tasks. 
To find out more about Dataplex, head to cloud.google.com/dataplex or watch the video below:

https://youtube.com/watch?v=bbFeAt7cw1g%3Fenablejsapi%3D1%26

Helping you innovate everyday

Sharing data is hard. Traditional data sharing techniques use batch data pipelines that are expensive to run, create late arriving data, and can break with any changes to the source data. These techniques also create multiple copies of data, which brings unnecessary costs and can bypass data governance processes. They also fail to offer features for data monetization, such as managing subscriptions and entitlements. Altogether, these challenges mean that organizations are unable to realize the full potential of transforming their business with shared data.

Analytics Hub

To address these limitations, we are introducing Analytics Hub, a new fully managed service that helps organizations unlock the value of data sharing, leading to new insights and increased business value. 

This new service is built on the tremendous experience and feedback we have received over the years. For example, BigQuery has had cross-organizational, in-place data sharing capabilities since its inception in 2010—and the functionality is very popular. Over a 7-day period in April, we had over 3,000 different organizations sharing over 200 petabytes of data. These numbers don’t include data sharing between departments within the same organization.

One week in the life of data sharing in BigQuery

Analytics Hub takes sharing to the next level, making it easy for you to publish, discover, and subscribe to valuable datasets that you can combine with your own data to derive unique insights. 

This includes: 

  • Shared datasets: As a data publisher, you create shared datasets that contain the views of data that you want to deliver to your subscribers. Data subscribers can search through the datasets that are available across all exchanges for which they have access and subscribe to relevant datasets. In addition, the publisher can track subscribers, disable subscriptions, and see aggregated usage information for the shared data.
  • Curated, self-service data exchanges: Exchanges are collections used to organize and secure shared datasets. By default, exchanges are completely private, but granular roles and permissions make it easy to deliver data to the right audience—whether internal or public. 

This is just the beginning for Analytics Hub. Please sign up for the preview, which is scheduled to be available in the third quarter of 2021.

Dataflow Prime

At Google Cloud, we have the great privilege of working with some of the most innovative organizations in the world. And this work provides us with a unique perspective into the future of big data processing. Dataflow Prime is a new platform based on a serverless, no-ops, and auto-tuning architecture that brings unparalleled resource utilization and radical operational simplicity to big data processing. This new service introduces a large number of exciting capabilities but I’d like to highlight three key aspects of the product:

  • Vertical Autoscaling: Dataflow Prime dynamically adjusts the compute capacity allocated to each worker based on utilization, detecting when jobs are limited by worker resources and automatically adding more resources. Vertical Autoscaling works hand in hand with Horizontal Autoscaling to seamlessly scale workers to best fit the needs of the pipeline. As a result, it no longer takes hours or days to determine the perfect worker configuration to maximize utilization. 
  • Right Fitting: Each stage of a pipeline typically has a different resource requirement than the others. Until now, either all workers in the pipeline would have had the higher memory and GPU, or none of them would. Pipelines either had to waste resources or suffer slower workloads. Right Fitting solves this problem by creating stage-specific pools of resources, optimized for each stage. 
  • Smart Recommendations: Smart Recommendations automatically detects problems in your pipeline and shows potential fixes. For example, if your pipeline is running into permissions issues, a Smart Recommendation will detect which IAM permissions you need to enable to unblock your job. If you are using an inefficient coder in your job, Smart Recommendations will surface more performant coder implementations that can help you save on costs.

What’s next

We’re excited to hear your thoughts and feedback about all these exciting new services. I would also highly recommend that you connect with members of the community to learn more about their story and journey. A good example to start with is the Data To Value customer panel we produced at our inaugural Data Cloud Summit with the Chief Data Officers of Keybank and Rackspace. You can watch it for free below:

https://youtube.com/watch?v=ITI2Q3MkxuA%3Fenablejsapi%3D1%26

4686

Of your peers have already watched this video.

10:30 Minutes

The most insightful time you'll spend today!

Explainer

How Google’s Customer Data Platform Helps Retail Brands Offer Data-driven , Personalized CX

Retail companies need customer insights to deliver personalized experiences that impact revenue generation and cost savings. Watch how Google Cloud’s customer data platform helps brands integrate and build holistic view of data in silos to drive marketing and customer service success.

Blog

A Year of Going Carbon-free! Google’s Road to Sustainability Looks Promising

3728

Of your peers have already read this article.

2:00 Minutes

The most insightful time you'll spend today!

Google Cloud announced its sustainability goal to turn fully carbon-free by 2030. To mark the progress of the data centres on the road to sustainability, Google Cloud releases 2020 carbon-free energy percentages (CFE%). Read to know more.

Last year, we announced our most ambitious sustainability goal yet: to operate everywhere on 24/7 carbon-free energy by 2030. We’ve set this goal to ensure that Google Cloud continues to be the cleanest cloud in the industry, and to show that full-scale decarbonization of electricity use is possible. 

Since setting our target, we’ve made tremendous progress in how we trackbuy, and use electricity; advocate for clean energy policies; and support the development of new technologies to help us reach this goal. And we’ve done it all while maintaining a commitment to transparency that we hope will make it easier for other organizations wishing to fully decarbonize their operations as well.

In the spirit of transparency, today we’re releasing the 2020 carbon-free energy percentages (CFE%) for all Google data centers, as well as overall progress on the road to our 2030 goal: In 2020, Google achieved 67% round-the-clock carbon free energy across all its data centers, up from 61% in 2019. In other words, of all the electricity consumed by Google data centers in 2020, two-thirds of it was matched with local, carbon-free sources on an hourly basis.

Though we saw a significant jump in global CFE% in 2020, we expect the numbers to vary from year to year. Ultimately, CFE% is dependent on the amount of new clean energy that comes online in a given year; we may even occasionally see short-term drops in the numbers. What’s most important is that we continue to maintain a long-term trajectory toward our 2030 goal. With meaningful progress in clean energy policy, technologies, and transactional models, we believe 24/7 carbon-free energy is achievable.

Tracking these numbers also allows us to give Google Cloud customers greater control in their own sustainability efforts. Earlier this year we announced Google Cloud Region Picker, a system that helps our customers assess factors like cost, speed, and CFE% as they choose where to run their applications. 

To outline some of the events that have helped us get to 67% CFE%, we’ve developed an animation that shows every hour of electricity use in 2020, at all our Google data centers around the world. 

Visualizing clean energy: every hour, every day, everywhere

Imagining what every hour in a year looks like is hard enough (there are 8,760 of them, in case you’re wondering). With 23 data centers and 25 cloud regions around the world, we’re aiming to source clean energy for over 200,000 operational hours each year.

https://youtube.com/watch?v=f9ecEokcFlk%3Fenablejsapi%3D1%26

The “A year in carbon-free energy” animation points out significant projects that came online in 2020 to bring our data centers closer to operating entirely on round-the-clock carbon-free energy. It also reflects an unparalleled level of transparency about our carbon-free energy data, showing hour-by-hour where we need to develop new clean energy projects, advocate for policy changes, and in some cases, look to new technologies that can help fill in the gaps left by variable renewable resources. 

In the animation, you’ll notice sites with a lot of green at midday (e.g. in Chile or the U.S. Southeast) – a sign that solar is making a big contribution. Other data centers, such as our facilities in the U.S. Midwest, rely more heavily on wind power and are subject to seasonal fluctuations in wind speed. 

Preventing the worst impacts of climate change will require decarbonizing the world’s electric grids, as fast as possible. Google is committed to doing as much as possible to clear a path for others and drive collective action to achieve this goal. We’re thrilled to be in good company as we move, together, toward a carbon-free future.

Blog

Google Cloud Connector for SAP LaMa Helps Make Most of the Multi & Hybrid-cloud Strategies

3489

Of your peers have already read this article.

1:30 Minutes

The most insightful time you'll spend today!

Google Cloud Connector for SAP LaMa is an important contribution for organizations dealing with complexities associated with hybrid-cloud and multi-cloud application strategies. Learn how connector extends LaMa functionality to SAP systems on cloud.

As an SAP user, you may already be familiar with SAP Landscape Management (SAP LaMa) — a specialized product that supports centralized management of your SAP landscape. SAP LaMa is useful for a wide range of SAP customers and use cases, but it’s especially important to large enterprises that manage very large, diverse, and often geographically dispersed SAP landscapes.

Google Cloud recently released a free connector that our SAP customers can deploy on an existing or new SAP LaMa instance. It’s a potentially important contribution for enterprises that need to address the growing complexity associated with multi-cloud and hybrid-cloud application strategies.

Let’s take a look at how SAP LaMa works, how it fits into a customer’s SAP landscape, and what kinds of benefits they can expect. We’ll also run through some adjustments that SAP admins will need to make in order to launch the connector.

SAP LaMa: Bringing simplicity to SAP landscapes

The Google Cloud Connector for SAP LaMa is distributed as a Java archive — provided free of charge — that customers can deploy on a SAP LaMa system that already exists on premises, or on another cloud, or as part of a new SAP LaMa deployment on Google Cloud. SAP LaMa itself is a centralized SAP management tool designed to simplify, automate, and orchestrate a variety of management and administration tasks across your entire SAP landscape. Some examples of use cases for LaMa include:

  • Getting a single, comprehensive overview of your organization’s full SAP landscape
  • Automating day-to-day system administration tasks, like system refreshes
  • Executing mass stop/start operations across your landscape
  • Deploying custom operations and workflows within the SAP application environment

SAP admins can reclaim a significant amount of time by creating SAP context sensitive automation routines that can execute without human intervention, and by completing system maintenance tasks more quickly and consistently. SAP LaMa also elevates service quality by using automation to remove manual intervention (and thus human error) from the system admin process.

The Google Cloud Connector builds bridges for SAP customers

The Google Cloud connector functions as a trigger that extends LaMa functionality to SAP systems deployed on Google Cloud. The connector enables standard SAP LaMa execution scenarios, such as:

  • Listing Google Cloud projects, zones, and VM instances within SAP LaMa’s landscape overview
  • Mass stop/start of Google Compute Engine (GCE) instances, using either the SAP LaMa user interface or the scheduler
  • System clone, copy, and DB refresh with Post Copy Automation
  • Resizing of machine types
  • Relocating SAP application instances to another VM instance
  • SAP HANA failover to a replicated SAP HANA HA system  via the SAP Host Agent

By supporting native SAP automation and orchestration, the Google Cloud connector contributes to the overall value customers get from using SAP LaMa. What may be just as important, however, is how the connector makes life easier for customers running Google Cloud as part of a hybrid cloud or multi-cloud SAP landscape. That’s an increasingly important advantage, given recent research revealing that 82% of enterprises are currently deploying hybrid clouds.1

Before getting started with the Google Cloud connector, it’s important to be aware that you’ll have to run it on SAP LaMa 3.0 (the current version), and that you will need to have an SAP Landscape Management – Enterprise Edition licence provided by SAP to support all functionality.

In addition, getting access to the full SAP LaMa feature set for your Google Cloud SAP environment will require you to deploy SAP systems using SAP’s adaptive design principles. These include:

  • Using SID (SAP system ID) and instance-specific disk mount points
  • Using alias IP for virtual host mapping and portability
  • Configuring DNS to manage virtual hostname (aka. FQDN) for all VMs
  • Using a local NFS host to manage shared files, as well as copying, migrating, and managing solutions like Filestore and NetApp separately

Because meeting these requirements may require altering your production systems, you’ll have to assess the pros and cons of doing so to take full advantage of the Google Cloud Connector.

Google Cloud is constantly looking for ways to support our SAP customers in running more efficiently and in making the most of their hybrid cloud and multi-cloud strategies. Releasing the Google Connector for SAP LaMa is one more reflection of our commitment to creating value for customers any way we can. Learn more about Google Cloud offerings for SAP customers.

More Relevant Stories for Your Company

Case Study

Medical Data Breakthrough: Google Aids PicnicHealth’s Growth

In the fragmented world of U.S. healthcare, patients often have to wait in line or on hold, navigate multiple patient portals, and fill out numerous request forms—all in pursuit of their own medical history. Healthcare technology startup PicnicHealth is on a mission to put control back with the patient, where

Case Study

How Sri Lanka’s Largest Ride-hailing Company Fixed its App and Improved Business

PickMe is Sri Lanka's largest ride-hailing company. “(Almost) every Sri Lankan is our customer. We have passengers who use us on a daily basis. We have drivers who use the platform to make a living. So obviously, the ecosystem is pretty big,” says Jiffry Zulfe, Founder & CEO, PickMe. Before

Blog

Qlik and Google Cloud Combo: Extending Integration of SAP Data on BigQuery

If your organization is one of the 52% of SAP customers whose top analytics pain point is data integration1, Google Cloud has got you covered. By working with partners like Qlik, we are expanding our integration options and bringing real-time replication capability for SAP to BigQuery. Integrated data for accelerated insights

Blog

Rethinking retail with Google Cloud Retail Search

Cloud Retail Search, part of Discovery Solutions For Retail portfolio, helps retailers significantly improve the shopping experience on their digital platform with ‘Google-quality’ search. Cloud Retail Search offers advanced search capabilities such as better understanding user intent and self-learning ranking models that help retailers unlock the full potential of their

SHOW MORE STORIES