Data Quality

Data Quality has gained paramount importance because businesses today use data for decision making. Bad quality data can lead to imperfect reports, which in turn lead to misguided conclusions, increased operational costs, and problems for downstream users of data. With data coming from disparate sources and with no control over the quality of the source data, it becomes necessary to transform and cleanse the data to make it suitable for data analysis. It is crucial that data that is used for analytical purposes provides accurate inputs, else inaccurate data can lead to incorrect output or results. Flawed data leads to faulty insights for business critical decisions that can cost companies time, money, and resources.

Data quality provides an aggregate score of the overall quality of the data and provides a percentage rating that shows how accurate the data is.

What is Data Quality?

Some important elements that define the quality of your dataset include:

  • Accuracy - Checks the level to which data conforms to a defined standard. For example, if the date is required to be mentioned in the mm:dd:yy format and in instead mentioned in the dd:mm:yy format, then this data is inaccurate.

  • Completeness - Checks if the data has all the required values with no missing information. For example, a complete record of address of an employee guarantees that the employee is reachable.

  • Consistency - Checks if the information that is stored and used at multiple instances matches. For example, if some phone records are stored with international code separately and some are prefixed with the international code, then there is need for consistency.

  • Uniqueness - Checks if a record has a single instance in a data set.

  • Validity - Checks the degree to which data matches with business rules or definitions accurately. It also checks if data conforms to the correct, accepted formats and values of the dataset fall within the proper range.

Data Quality in Data Pipeline Studio

Data Quality in Data Pipeline Studio helps you assess, validate, and monitor the quality of your datasets. You can profile datasets, define and execute data quality rules, identify data quality issues, and analyze data quality results.

Data Pipeline Studio supports Data Quality for datasets stored in S3, Unity Catalog data lake, and Snowflake. The Data Quality workflow varies depending on the data lake and technology you use.

Data Quality Workflows

Depending on your specific use case and the technology preference of your organization, you can use one of the following workflows for data quality:

Are any data quality stages mandatory?

The simple answer to this question is no. The data quality stages mentioned above are optional and none of the stages is a prerequisite for the next one. You can use each stage independently as long as you have a valid usecase and the required data for that particular stage. For example, you can do profiling of data for a better understanding of your dataset. Alternately, you can skip this stage and directly use data analyzer to perform a complete analysis of your dataset. If you have data which is already analyzed, you can directly use the issue resolver stage to resolve some of the issues with that data.

Related Topics Link IconRecommended Topics What's next? Databricks Data Profiler