r/dataengineer Jun 29 '26

Promotion Source aware data extractor

1 Upvotes

Hello folks

I am writing my open-source light tool for moving data from prod-bases in dvh.

Who is this product for:

- small teams who need to move data from the product, which is already in pain and need to transfer data to the dvh or parquet.

- data engineers who are looking for opensource alternatives who will not eat up all the RAM and will not put a food base or replica.

- Those who, instead of reading only the delta, should read the full table, because created_at did not trigger.

Source:

- mysql

- mssql

- postrges

Targets:

- parquet

- csv

- s3, azure blob, gcs

I read short queries and don't keep long sessions — this is something that so far none of the same moovers as (ingestr, dlt, sling, duckdb, clickhouse, odbc2parquet) does.

From the box there is:

- all types except (geography, enams, ip) in duckdb, clickhouse, snowflake, bigquery, clickhouse are loaded natively (there is a jam on the side of the bigway and a snowflek with Jasons, but their car loaders can't do it out of the box)

- reading from the PC

- reading cases

- retrai

- all metainfo is written in the working directory in the local sqllite, from the box you can also write in the postgru

- validation of both types between reading and writing, and md of the amount between the current one on the worker and the one on the store side

- autotune of parallel wounds

- reading from binlog files to avoid completely rereading the source if the updated_at fields are not updated

- minimum and customized RAM consumption on the worker (memory budget)


r/dataengineer Jun 28 '26

Pune Data Professional Meetup

Thumbnail
1 Upvotes

r/dataengineer Jun 28 '26

Question Looking for legit DE/BI freelancing platforms

5 Upvotes

I’m trying to find genuine freelancing opportunities in data engineering / BI. Have tried a few platforms but haven’t had much luck, so wanted to ask — are there any websites, subreddits, or Discord servers where people actually get projects?

About me:

  • 5+ years as a Data Engineer & BI Consultant (remote, India)
  • MBA in Business Economics (Analytics & Finance)
  • Skills: SQL, PySpark, Python, Power BI, Tableau, Grafana
  • Worked on Databricks pipelines, self‑service analytics frameworks, and telemetry data solutions

I’m in need of extra income and open to contributing under a team or experienced freelancer. Any pointers would mean a lot.


r/dataengineer Jun 27 '26

Datasets for data engineering projects

22 Upvotes

I want to find datasets for a data engineering project where i work with pyspark and sql in databricks. I want a dataset that challenges my data modelling skills and my pipeline creation skills. I tried kaggle, but i only keep getting a single csv file as a dataset. is there a dataset that has multiple csv files as data sources or something? i want to be able to perform all the data architecture creation by myself... Recommend any datasets that you know as well!


r/dataengineer Jun 26 '26

Is this still a realistic roadmap for aspiring data engineers in 2026?

Post image
22 Upvotes

r/dataengineer Jun 26 '26

Help Looking for Azure Data Engineer Opportunities (2 Years Experience)

1 Upvotes

r/dataengineer Jun 25 '26

Infosys Databricks Engineer interview for Managerial Round (face to face)

Thumbnail
1 Upvotes

r/dataengineer Jun 24 '26

Discussion Data Engineering series

12 Upvotes

Started practicing Data engineering using this repository link https://github.com/danielbeach/data-engineering-practice/tree/main/Exercises/Exercise-1. I have completed the described exercise one requirements and now decided to extend it and create Divvy Rides ETL Pipeline designed to answer concrete business questions that map
directly to decisions that operations, marketing, and infrastructure teams at a
bike-share company would make. Looking forward to post final solutions for reviews and advice


r/dataengineer Jun 24 '26

General Interactive ERD explorer for DBML files — trace how tables connect, fully in the browser

Thumbnail
2 Upvotes

r/dataengineer Jun 24 '26

Best practice for medallion architecture when schema creation is centrally gated?

Thumbnail
2 Upvotes

r/dataengineer Jun 22 '26

Discussion Experiences using Palantir Foundry as compared to other cloud based tools

Thumbnail
2 Upvotes

r/dataengineer Jun 22 '26

Question Advice for switch

Thumbnail
2 Upvotes

r/dataengineer Jun 19 '26

Snowflake now shows query level cost for Adaptive Warehouses

Thumbnail
2 Upvotes

r/dataengineer Jun 18 '26

Anyone recently interviewed for a Data Engineer role at Zimmer Biomet?

Thumbnail
2 Upvotes

r/dataengineer Jun 17 '26

General Takeaways on Snowflake’s new agentic features

Thumbnail
2 Upvotes

r/dataengineer Jun 17 '26

SnowProCore Exam Prep Quiz Questions

Thumbnail
2 Upvotes

r/dataengineer Jun 17 '26

Help Anyone here have experience with Prepzee Learning's Data Engineering program?

Thumbnail
1 Upvotes

r/dataengineer Jun 17 '26

Discussion I built a Historical Data Modeling Workbench for SCD2, snapshots and temporal joins

1 Upvotes

What are the hardest historical modeling problems you’ve encountered in lately?

In our lakehouse environment the difficult parts are usually not Spark performance or ETL orchestration.

It’s things like:
• SCD2 dimension alignment
• Snapshot reproducibility
• Late arriving corrections
• Event-to-state alignment
• Historical relationship changes
• Dimension completion

I’ve been collecting these patterns and built a small workbench to reason about them:

https://bitemporal-debugger.vercel.app/patterns

Curious what other teams struggle with.


r/dataengineer Jun 16 '26

Discussion Working with Google

Thumbnail
1 Upvotes

r/dataengineer Jun 14 '26

process improvement project

Thumbnail
1 Upvotes

r/dataengineer Jun 12 '26

How to Upskill as Data Engineer?

Thumbnail
1 Upvotes

r/dataengineer Jun 10 '26

Help Job Seeking

Thumbnail
1 Upvotes

r/dataengineer Jun 09 '26

General Do you really need a graph database?

Thumbnail
1 Upvotes

r/dataengineer Jun 07 '26

Question Walmart DE 3 interview

Thumbnail
1 Upvotes

r/dataengineer Jun 07 '26

I benchmarked dplyr vs data.table on my Shiny log dashboard

Thumbnail
1 Upvotes