Data Quality
Data Quality has gained paramount importance because businesses today use data for decision making. Bad quality data can lead to imperfect reports, which in turn lead to misguided conclusions, increased operational costs, and problems for downstream users of data. With data coming from disparate sources and with no control over the quality of the source data, it becomes necessary to transform and cleanse the data to make it suitable for data analysis. It is crucial that data that is used for analytical purposes provides accurate inputs, else inaccurate data can lead to incorrect output or results. Flawed data leads to faulty insights for business critical decisions that can cost companies time, money, and resources.
Data quality provides an aggregate score of the overall quality of the data and provides a percentage rating that shows how accurate the data is.
What is Data Quality?
Some important elements that define the quality of your dataset include:
-
Accuracy - Checks the level to which data conforms to a defined standard. For example, if the date is required to be mentioned in the mm:dd:yy format and in instead mentioned in the dd:mm:yy format, then this data is inaccurate.
-
Completeness - Checks if the data has all the required values with no missing information. For example, a complete record of address of an employee guarantees that the employee is reachable.
-
Consistency - Checks if the information that is stored and used at multiple instances matches. For example, if some phone records are stored with international code separately and some are prefixed with the international code, then there is need for consistency.
-
Uniqueness - Checks if a record has a single instance in a data set.
-
Validity - Checks the degree to which data matches with business rules or definitions accurately. It also checks if data conforms to the correct, accepted formats and values of the dataset fall within the proper range.
Data Quality in Data Pipeline Studio
Data Quality in Data Pipeline Studio helps you assess, validate, and monitor the quality of your datasets. You can profile datasets, define and execute data quality rules, identify data quality issues, and analyze data quality results.
Data Pipeline Studio supports Data Quality for datasets stored in S3, Unity Catalog data lake, and Snowflake. The Data Quality workflow varies depending on the data lake and technology you use.
Data Quality Workflows
Depending on your specific use case and the technology preference of your organization, you can use one of the following workflows for data quality:
The four-stage approach is available when Databricks is used as the Data Quality technology and S3 is used as the data lake.
The workflow consists of the following stages:
Data Profiler → Data Analyzer → Data Validator → Data Issue Resolver
Each stage performs a specific Data Quality function:
-
Data Profiler - this is the stage which runs an analysis on a sample piece of data against selected parameters like completeness, validity, character count and so on which gives you statistics of the data on those parameters. If you have used validity as a constraint in the data profiler job, then after output of the profiler job is available you can create a validator job. The validator results provide more information about the data pattern in the selected columns.
-
Data Analyzer - in this stage an analysis is performed on the complete dataset, based on the selected constraints. This job runs in two parts.
-
First you create and run a data analyzer job by adding the required constraints.
-
Once the job run is complete you create a data validator job. In this job you can add the constraints of the data analyzer job as well as new constraints as per your requirement.
-
-
Issue Resolver - in this stage you further enhance the quality of data by resolving some of the data-related issues that are found. You can achieve this by running the data through various constraints to improve the quality of the data, in the following ways:
-
Handling duplicate data
-
Handling missing data
-
Handling outliers
-
Specifying the partitioning order
-
Handling string operations
-
Handling case sensitivity
-
Replace selective data
-
Handle data against master table
-
The integrated approach is available for Unity Catalog data lake and Snowflake.
This approach combines multiple Data Quality capabilities into fewer stages:
Data Integration → DQ Processor
The Data Integration stage combines Data Ingestion and Data Profiling, enabling you to ingest and profile datasets as part of a single workflow.
The DQ Processor combines the Data Analyzer, Data Validator, and Data Issue Resolver capabilities into a single processing stage. It enables you to create, execute, and evaluate data quality rules.
The DQ Processor also supports automated data quality rule generation, which can help streamline the process of creating rules for a dataset.
See Data Quality (DQ) Processor using Unity Catalog.
See Data Quality (DQ) Processor using Snowflake
After the DQ Processor executes the rules, you can use the DQ Results Dashboard to analyze the data quality results and identify data quality issues.
Are any data quality stages mandatory?
The simple answer to this question is no. The data quality stages mentioned above are optional and none of the stages is a prerequisite for the next one. You can use each stage independently as long as you have a valid usecase and the required data for that particular stage. For example, you can do profiling of data for a better understanding of your dataset. Alternately, you can skip this stage and directly use data analyzer to perform a complete analysis of your dataset. If you have data which is already analyzed, you can directly use the issue resolver stage to resolve some of the issues with that data.
| What's next? Databricks Data Profiler |