MIT licensed ยท SQLite, SQL, Streamlit

Messy event data,
turned into numbers you can trust.

A complete analytics engineering pipeline: raw delivery events go through SQL-based ELT into canonical fact/dimension tables, get validated by 7 automated data quality checks, and land in a self-serve dashboard โ€” the full path from raw data to something a business team can act on.

19,960clean rows from 20,100 raw
7/7data quality checks passing
4canonical fact/dim tables
pipeline run

Raw events โ†’ canonical tables

SQL window functions and joins deduplicate and clean the data before anything downstream ever touches it.

raw_orders.csv  (raw, messy: duplicates, nulls, bad values)
       โ”‚
       โ–ผ
elt_pipeline.py โ€” SQL transforms in SQLite
       โ”‚
       โ–ผ
dim_restaurant ยท dim_date ยท fact_orders ยท mart_daily_region_metrics
       โ”‚
       โ”œโ”€โ”€โ–บ data_quality_checks.py โ€” 7 automated checks
       โ”‚
       โ””โ”€โ”€โ–บ dashboard.py โ€” Streamlit, self-serve

7 checks, run before anyone trusts the data

Declarative expectations in the same spirit as Great Expectations or Pydeequ โ€” CI-ready, exits non-zero on failure.

01
Uniqueness
order_id has no duplicates in fact_orders
02
Referential integrity
every fact row joins to a valid dimension key
03
Null-rate thresholds
key columns stay under an acceptable null %
04
Plausible ranges
subtotals and delivery times fall in sane bounds
05
No negative values
subtotals and quantities can't be negative
06
Row-count sanity
cleaned volume falls within an expected range of raw
07
Grain check
mart table has exactly one row per date + region

Results

What the pipeline actually produces, end to end.

20,100raw rows generated
19,960clean rows after dedup + filtering
7/7quality checks passing
4metrics on the dashboard

Run it

No warehouse account needed โ€” SQLite stands in for Snowflake/BigQuery locally.

bash
pip install pandas numpy streamlit
python3 generate_data.py        # ~20k raw rows, intentionally messy
python3 elt_pipeline.py         # builds analytics.db, canonical tables
python3 data_quality_checks.py  # 7 checks, exits 0 if all pass
streamlit run dashboard.py      # launches the self-serve dashboard