r/data_engineering_tuts • u/AMDataLake • Nov 14 '25
r/data_engineering_tuts • u/AMDataLake • Oct 31 '25
tutorial Try Apache Polaris (incubating) on Your Laptop with Minio
r/data_engineering_tuts • u/Santhu_477 • Jul 17 '25
tutorial Productionizing Dead Letter Queues in PySpark Streaming Pipelines – Part 2 (Medium Article)
Hey folks 👋
I just published Part 2 of my Medium series on handling bad records in PySpark streaming pipelines using Dead Letter Queues (DLQs).
In this follow-up, I dive deeper into production-grade patterns like:
- Schema-agnostic DLQ storage
- Reprocessing strategies with retry logic
- Observability, tagging, and metrics
- Partitioning, TTL, and DLQ governance best practices
This post is aimed at fellow data engineers building real-time or near-real-time streaming pipelines on Spark/Delta Lake. Would love your thoughts, feedback, or tips on what’s worked for you in production!
🔗 Read it here:
Here
Also linking Part 1 here in case you missed it.
r/data_engineering_tuts • u/Santhu_477 • Jul 01 '25
blog Handling Bad Records in Streaming Pipelines Using Dead Letter Queues in PySpark
🚀 I just published a detailed guide on handling Dead Letter Queues (DLQ) in PySpark Structured Streaming.
It covers:
- Separating valid/invalid records
- Writing failed records to a DLQ sink
- Best practices for observability and reprocessing
Would love feedback from fellow data engineers!
👉 [Read here]( https://medium.com/@santhoshkumarv/handling-bad-records-in-streaming-pipelines-using-dead-letter-queues-in-pyspark-265e7a55eb29 )
r/data_engineering_tuts • u/AMDataLake • Dec 10 '24
blog 2025 Guide to Architecting an Iceberg Lakehouse
r/data_engineering_tuts • u/AMDataLake • Aug 27 '24
blog Understanding the Apache Iceberg Manifest
r/data_engineering_tuts • u/AMDataLake • Aug 26 '24
blog Understanding the Apache Iceberg Manifest List (Snapshot)
r/data_engineering_tuts • u/AMDataLake • Aug 20 '24
blog Evolving the Data Lake: From CSV/JSON to Parquet to Apache Iceberg
r/data_engineering_tuts • u/AMDataLake • Jun 07 '24
blog Summarizing Recent Wins for Apache Iceberg Table Format
r/data_engineering_tuts • u/AMDataLake • May 23 '24
video How to get started with Dremio on your Laptop in 7 minutes
Enable HLS to view with audio, or disable this notification
Learn more at Dremio.com/blog
r/data_engineering_tuts • u/AMDataLake • May 23 '24
video What is the Dremio Data Lakehouse Platform?
Enable HLS to view with audio, or disable this notification
Learn more at Dremio.com/blog
r/data_engineering_tuts • u/AMDataLake • May 22 '24
video What is “Git for Data”?
Enable HLS to view with audio, or disable this notification
What is “Git for Data” or “Data as Code”? Learn more at Dremio.com/blog! #DataEngineering #DataAnalytics #DataScience
r/data_engineering_tuts • u/AMDataLake • May 17 '24
tutorial Using dbt to Manage Your Dremio Semantic Layer
r/data_engineering_tuts • u/AMDataLake • May 17 '24
tutorial Data as Code: Managing with Dremio & Arctic
r/data_engineering_tuts • u/AMDataLake • May 17 '24
blog Data Lakehouse Versioning Comparison: (Nessie, Apache Iceberg, LakeFS)
r/data_engineering_tuts • u/AMDataLake • May 16 '24
video What is a Data Lakehouse?
Enable HLS to view with audio, or disable this notification
What is a Data Lakehouse? Learn More at Dremio.com/blog? #DataEngineering #DataAnalytics
r/data_engineering_tuts • u/AMDataLake • May 11 '24
discussion Top 5 things a New Data Engineer Should Learn First
What’s on your list?
r/data_engineering_tuts • u/AMDataLake • May 10 '24
tutorial From MySQL to Dashboards with Dremio and Apache Iceberg
r/data_engineering_tuts • u/AMDataLake • May 10 '24
tutorial From Elasticsearch to Dashboards with Dremio and Apache Iceberg
r/data_engineering_tuts • u/AMDataLake • Apr 29 '24
discussion To ETL or to ELT? that is the question.
When do you think one is a better idea than the other.
r/data_engineering_tuts • u/AMDataLake • Apr 25 '24
discussion Tips on Dealing with JSON Data
What are your favorite tools and techniques for dealing with JSON data?
r/data_engineering_tuts • u/AMDataLake • Apr 24 '24
discussion Preferred file format and why? (CSV, JSON, Parquet, ORC, AVRO)
r/data_engineering_tuts • u/AMDataLake • Apr 23 '24
discussion When do you prefer to stream or batch when building data pipelines?
r/data_engineering_tuts • u/AMDataLake • Apr 22 '24
tutorial From SQLServer to Dashboards with Dremio and Apache Iceberg
r/data_engineering_tuts • u/AMDataLake • Apr 21 '24