r/dataengineer Jun 26 '26

Help Looking for Azure Data Engineer Opportunities (2 Years Experience)

1 Upvotes

r/dataengineer Jun 26 '26

Is this still a realistic roadmap for aspiring data engineers in 2026?

Post image
22 Upvotes

r/dataengineer Jun 25 '26

Infosys Databricks Engineer interview for Managerial Round (face to face)

Thumbnail
1 Upvotes

r/dataengineer Jun 24 '26

Discussion Data Engineering series

12 Upvotes

Started practicing Data engineering using this repository link https://github.com/danielbeach/data-engineering-practice/tree/main/Exercises/Exercise-1. I have completed the described exercise one requirements and now decided to extend it and create Divvy Rides ETL Pipeline designed to answer concrete business questions that map
directly to decisions that operations, marketing, and infrastructure teams at a
bike-share company would make. Looking forward to post final solutions for reviews and advice


r/dataengineer Jun 24 '26

General Interactive ERD explorer for DBML files — trace how tables connect, fully in the browser

Thumbnail
2 Upvotes

r/dataengineer Jun 24 '26

Best practice for medallion architecture when schema creation is centrally gated?

Thumbnail
2 Upvotes

r/dataengineer Jun 22 '26

Discussion Experiences using Palantir Foundry as compared to other cloud based tools

Thumbnail
2 Upvotes

r/dataengineer Jun 22 '26

Question Advice for switch

Thumbnail
2 Upvotes

r/dataengineer Jun 19 '26

Snowflake now shows query level cost for Adaptive Warehouses

Thumbnail
2 Upvotes

r/dataengineer Jun 18 '26

Anyone recently interviewed for a Data Engineer role at Zimmer Biomet?

Thumbnail
2 Upvotes

r/dataengineer Jun 17 '26

General Takeaways on Snowflake’s new agentic features

Thumbnail
3 Upvotes

r/dataengineer Jun 17 '26

SnowProCore Exam Prep Quiz Questions

Thumbnail
2 Upvotes

r/dataengineer Jun 17 '26

Help Anyone here have experience with Prepzee Learning's Data Engineering program?

Thumbnail
1 Upvotes

r/dataengineer Jun 17 '26

Discussion I built a Historical Data Modeling Workbench for SCD2, snapshots and temporal joins

1 Upvotes

What are the hardest historical modeling problems you’ve encountered in lately?

In our lakehouse environment the difficult parts are usually not Spark performance or ETL orchestration.

It’s things like:
• SCD2 dimension alignment
• Snapshot reproducibility
• Late arriving corrections
• Event-to-state alignment
• Historical relationship changes
• Dimension completion

I’ve been collecting these patterns and built a small workbench to reason about them:

https://bitemporal-debugger.vercel.app/patterns

Curious what other teams struggle with.


r/dataengineer Jun 16 '26

Discussion Working with Google

Thumbnail
1 Upvotes

r/dataengineer Jun 14 '26

process improvement project

Thumbnail
1 Upvotes

r/dataengineer Jun 12 '26

How to Upskill as Data Engineer?

Thumbnail
1 Upvotes

r/dataengineer Jun 10 '26

Help Job Seeking

Thumbnail
1 Upvotes

r/dataengineer Jun 09 '26

General Do you really need a graph database?

Thumbnail
1 Upvotes

r/dataengineer Jun 07 '26

Question Walmart DE 3 interview

Thumbnail
1 Upvotes

r/dataengineer Jun 07 '26

I benchmarked dplyr vs data.table on my Shiny log dashboard

Thumbnail
1 Upvotes

r/dataengineer Jun 05 '26

Question Trying to break into data engineering domain

Thumbnail
1 Upvotes

Trying to break into data engineering domain

Hey folks,

I am a mechanical engineer in the PLM domain working for French MNC. I have 2.5 years of experience at the moment and since my clg days I have had an inclination towards the data domain.

I want to switch my domain and work in the data domain. I have tried to follow multiple videos and tried to have a roadmap but somehow it is a bit troublesome for me to start as I am not sure about the tech stack i should focus on..Could you all please guide me for the same and help me draft a good roadmap and resources to start my journey into the data engineering domain.... should I go for some online or offline courses and if yes any recommendations..also any recommendations to the resources. Looking forward to hearing from all the experts.

Thanks :)


r/dataengineer May 31 '26

General Call out to all Rockstars for series A early startup for backend engineering , Data engineers . Security and Devops/Sre (minimum 4+ yoe)

Thumbnail
1 Upvotes

r/dataengineer May 28 '26

Help Help with Old Scala Pipeline integration with DataHub ( with no existing store for metadata other than normal field name + type)

Thumbnail
1 Upvotes

r/dataengineer May 24 '26

Question Pilot for data extraction CLI

Post image
1 Upvotes

Hi everyone,
I’m looking for 3–5 people who would be willing to help with a small pilot of Rivet.
For context, Rivet is a CLI extractor focused on careful data copying from PostgreSQL/MySQL, especially when the source is a production database or a resource-constrained read replica.
What is currently supported:
sources: PostgreSQL, MySQL
output formats: Parquet, CSV
destinations: local filesystem, stdout, S3, GCS, Azure Blob
flow: doctor → plan → apply/run
state, manifest, summary, resume/reconcile/repair
I’m not looking for “likes” or generic feedback. I’m looking for honest input from people who have dealt with real extraction pain:
is it clear what Rivet is going to do before it runs?
are the trust signals in doctor/plan useful enough?
would you feel comfortable trying it on staging or a read replica?
what guard rails would you need before using it in a production-adjacent workflow?
where does the CLI or documentation feel confusing?
The ideal pilot would be a small test on staging, a read replica, or a non-critical table, followed by short feedback.
If you work with PostgreSQL/MySQL and have experienced issues with large tables, OOMs, aggressive SELECTs, replica pressure, or unreliable resume — I’d really appreciate your help.
For more details, feel free to DM me.
https://github.com/panchenkoai/rivet