menu
arrow_back

Exploring the Lineage of Data with Cloud Data Fusion

—/100

Checkpoints

arrow_forward

Create a Cloud Data Fusion instance

Add Cloud Data Fusion API Service Agent role to service account

Import, Deploy and Run Shipment Data Cleansing pipeline

Import, Deploy, and Run the Delayed Shipments data pipeline

Exploring the Lineage of Data with Cloud Data Fusion

1 个小时 30 分钟 7 个积分

GSP812

Google Cloud Self-Paced Labs

Overview

This lab shows how to use Cloud Data Fusion to explore data lineage: the data's origins and its movement over time.

Cloud Data Fusion data lineage helps you:

  • Detect the root cause of bad data events
  • Perform an impact analysis prior to making data changes

Cloud Data Fusion provides lineage at the dataset level and field level, and is time-bound to show lineage over time.

  • Dataset level lineage shows the relationship between datasets and pipelines in a selected time interval.

  • Field level lineage shows the operations that were performed on a set of fields in the source dataset to produce a different set of fields in the target dataset.

For the purpose of this lab, you will use two pipelines that demonstrate a typical scenario in which raw data is cleaned then sent for downstream processing. This data trail from raw data to the cleaned shipment data to analytic output can be explored using the Cloud Data Fusion lineage feature.

Note: Currently, the Cloud Data Fusion Lineage feature is only available with the Cloud Data Fusion Enterprise Edition.

加入 Qwiklabs 即可阅读本实验的剩余内容…以及更多精彩内容!

  • 获取对“Google Cloud Console”的临时访问权限。
  • 200 多项实验,从入门级实验到高级实验,应有尽有。
  • 内容短小精悍,便于您按照自己的节奏进行学习。
加入以开始此实验