Find data quality issues and clean your data in a single line of code with a Scikit-Learn compatible Transformer.
-
Updated
Dec 13, 2023 - Python
Find data quality issues and clean your data in a single line of code with a Scikit-Learn compatible Transformer.
Run greatexpectations.io on ANY SQL Engine using REST API. Supported by FastAPI, Pydantic and SQLAlchemy as best data quality tool
🦆 Blazing Fast and highly customizable Github Action to setup a DuckDb runtime
SQL based data profiling & data quality checks, which will help you to perform data profiling & data quality checks on SQL database at table & database level.
A lightweight simple data quality testing tool.
DataBridge Quality Control
FinAUDIT is an AI-powered financial data health and compliance system that automatically audits datasets against global regulatory standards (GDPR, Visa CEDP, AML, PCI DSS, and Basel). It combines a deterministic 30-rule engine for rigorous data quality scoring with a Generative AI Analyst (Gemini) to provide natural-language answer.
dbt Datasphere Plugin is for integrating multiple open-source data quality frameworks into your dbt projects. It unifies Soda SQL, Great Expectations, Datafold, providing a single interface to configure and run data quality checks.
This project involves a comprehensive analysis to determine the top YouTubers in the UK for 2024, Using Excel, SQL and Power BI.
This repository contains a complete data lakehouse implementation using Docker. It showcases an end-to-end data pipeline with Apache Spark for ETL, MinIO and Delta Lake for storage, Airflow for orchestration, DQOps for data quality, and Superset for BI.
End-to-end data analytics project analyzing Amazon sales data using SQL and Power BI to uncover revenue drivers, seasonal trends, regional demand patterns, and discount-driven sales behavior through exploratory data analysis and interactive dashboards.
一站式解决测试数据痛点:从高质量模拟数据生成到严格的数据质量校验。
This project aims to import data from different sources in python and extracting insights about data quality using PANDAS library. Three different data sources are used in this dataset.
A fully‑modelled, Snowflake‑backed analytics pipeline for operational appointment intelligence across CareGrid clinics. This project builds a trustworthy, chain‑resolved no‑show metric, a canonical status model, and a data‑quality governance layer that surfaces scheduling inconsistencies instead of hiding them.
Data Governance Cockpit based on Great Expectations Engine
This project extracts and cleans raw YouTube data from excel-csv (Kaggle) through SQL and identifies the top-performing UK-based Influencers. Data Stack: Excel | Microsoft SQL Server | Power BI
An Apache Airflow data pipeline is designed to perform ELT operations, utilizing Amazon S3 and Amazon Redshift Serverless.
A personal project using an NLP Model to create graphical plots based on user inputted file and instructions (consisting of choice of columns to be used for the graphical plots) . Additionally, this tool also does data quality check of uploaded data file along with sentiment analysis of user-inputted text.
KPMG Data Analytics Consulting Virtual Internship
Ramblings of a curious mind
To associate your repository with the dataqualitycheck topic, visit your repo's landing page and select "manage topics."