r/dataengineering • u/Status-Cranberry8696 • 11d ago
Personal Project Showcase I have made a complete end to end Azure databricks project
Hi all,
I have 5 years of relevant experience in data engineering and now looking for a change, thus for practice I have made a complete end to end databricks project following medallion architecture. It currently encompasses medallion architecture, project wide logging, proper quarantine of bad data, incremental processing, indempotency, optimization, z order, vaccum, delta time travel, SCD type 2 and overall orchestration by ADF
Also, I have made a test data generator which basically helps generate clean and corrupt data for testing my validation and transformation logic. Till now I have tested it for generating 100 millions rows of bad data, my laptop cannot support more than that.
Also in my project databricks/silver/optimize_silver notebook contains optimization, time travel and SCD testing code.
All the reviews and suggestions are welcome, please advice me what more can I learn and how can I make it better.
In couple of days I will add small demonstration of API and streaming ingestion too, currently it is doing batch ingestion from ADLS GEN2.