Data Quality (DQ) Processor using Snowflake

The DQ Processor stage focuses on applying validation rules, identifying data quality issues, and managing issue resolver workflows based on the profiling results and defined constraints.

This topic describes how to create a data quality processor pipeline that reads data from Snowflake, applies data quality rules to validate and resolve identified issues, and writes the processed results back to Snowflake for downstream use.

The data quality processor pipeline has the following nodes:

Snowflake (data lake) > Snowflake (data quality node) - DQ Processor

Prerequisites

To create or run a data quality processor job using Snowflake, you must complete the following prerequisites:

  • Get access to a Snowflake data lake configuration listed under Configuration > Cloud Platform Tools & Technologies > Databases and Data Warehouses.

  • Ensure that the Statistics Common Table is configured in the Snowflake instance to store statistical data and metadata information.

To create a data quality processor job

  1. Sign in to the Calibo Accelerate platform and navigate to Products.

  2. Select a product and feature. Click the Develop stage of the feature and navigate to Data Pipeline Studio.

  3. Add the Data Lake stage > Snowflake node, then configure the data lake node.

  4. Add the Data Quality stage > Snowflake > DQ Processor node.

  5. Connect the data lake and data quality nodes to each other.

    Snowflake Data Quality pipeline

  6. Click the data quality processor node and complete the following steps to create a data quality processor job:

Related Topics Link IconRecommended Topics What's next?Databricks Templatized Data Integration Jobs