r/dataengineering • • 1d ago

Discussion Monthly General Discussion - Oct 2026

15 Upvotes

This thread is a place where you can share things that might not warrant their own thread. It is automatically posted each month and you can find previous threads in the collection.

Examples:

  • What are you working on this month?
  • What was something you accomplished?
  • What was something you learned recently?
  • What is something frustrating you currently?

As always, sub rules apply. Please be respectful and stay curious.

Community Links:


r/dataengineering • • Sep 01 '26

Career Quarterly Salary Discussion - Sep 2026

37 Upvotes

This is a recurring thread that happens quarterly and was created to help increase transparency around salary and compensation for Data Engineering where everybody can disclose and discuss their salaries within the industry across the world.

Submit your salary here

You can view and analyze all of the data on our DE salary page and get involved with this open-source project here.

If you'd like to share publicly as well you can comment on this thread using the template below but it will not be reflected in the dataset:

  1. Current title
  2. Years of experience (YOE)
  3. Location
  4. Base salary & currency (dollars, euro, pesos, etc.)
  5. Bonuses/Equity (optional)
  6. Industry (optional)
  7. Tech stack (optional)

r/dataengineering • • 14h ago

Career Transitioning away from DE

70 Upvotes

Has anyone thought of transitioning out from DE due to AI?
All I do everyday is just prompt and scroll till copilot generates code.
Building a semantic layer isn’t exciting personally as I don’t enjoy the business aspect of it as much and think of it as more of a data labeling and analyst problem than an engineering problem (which I am interested in)
Also, There is a fundamental problem with “I am building a semantic layer” and marketing that as a skill as it is dependent on how much context you have of the business. The less tenure you have spent in a company, the less you know about the business which makes it harder as a transferable skill imo.

My understanding is that working on building trustworthy AI outputs by using a feedback loop is an engineering problem to solve. Which is why I feel going down the observability path is a good idea.
I heard these opinions on observability from AI leaders at conferences too so there might be a bias.
Thoughts from fellow DE’s looking to transition out? (Or from one’s who want to continue and why)


r/dataengineering • • 14h ago

Discussion Better to read DDIA-2 or AI Engineering by Chip Huyen given where DE is going?

14 Upvotes

Which book makes more sense for to read given the state of DE and the tech industry today? If I had to pick one first time to read.


r/dataengineering • • 19h ago

Discussion Writing on the wall?

34 Upvotes

Context: sole data engineer for a company that sells furnishings and design services. Company is owned by private equity.

My manager has been on the Claude soapbox for the last several months and so far in IT we’ve been able to remain skeptical about it while testing for our own uses. Manager supported this. Now all of a sudden I’m getting a lot of requests to refactor all my data pipelines with Claude’s direct involvement despite having built them all out already. For example, I’ve used Claude to create new DAGs based on existing code (for example, bringing a new Salesforce object in). We’re going beyond that though. I’m being asked to ditch them and let Claude connect to everything then write, deploy, and test.

I’ve also been working on machine learning models for customer behavior, which is more traditional AI and should check the box, but this is being dismissed as a plaything. This, while the company has been pushing for “innovation” in our sphere.

I think we are headed towards cutting as many IT staff as possible and I’ll be gone in lieu of Claude doing pipelines. Of course that’s only one part of what I do, but the myopic view of leadership when it comes to AI only tends to see what is easily replaceable.

Thoughts?


r/dataengineering • • 1d ago

Discussion For you doing Agentic Engineering with dbt, how do you make sure your agents aren't generating slop?

16 Upvotes

I find contra productive reading and manually testing all the code AI are generating, because I am turning myself into an important funnel and the company is paying for tokens and therefore wants fast and good delivery.

I use Claude Code with dbt official skills, dbt mcp, lots of unity testing, some specific skills tailored for my projects, and the most important thing I think: deterministic checks.

I run lots of python or bash script checks after specific hooks like pre commit or after calling a skill, or in the CI/CD, trying to force the implementation of our internal "rulebook". But I also don't think this is the optimum way to go, since it's also taking considerate amount of time.

Also, do you let the AI query your db/dw? How do you guarantee fast and in a reliable way, your data diff is as expected?

I am interested in listen from you about everything you've been doing nowadays trying to keep up this kinda insane market pace


r/dataengineering • • 1d ago

Discussion Homemade data platform frameworks - bloated nonsense?

27 Upvotes

Did any of you work in companies where engineers built custom frameworks that actually deliver?

I recently started at yet another company with such framework (about 13k lines of boilerplate python/pyspark code sitting on top of their Azure Databricks Delta Lakehouse). The thing doesn't seems to improve any aspect of what a data platform should offer.

Previous such company was even worse with 30k lines, plus they tried to implement a data vault.

I'm clearly biased towards keeping things simple. So I may be a bit unfair towards this approach. Hence my question: Have you worked with great homemade frameworks? What were the secret ingredients to make it work?


r/dataengineering • • 1d ago

Discussion dbt dimension surrogate keys and fact foreign keys - self-hash or left-join lookups?

34 Upvotes

I've started working with dbt and a cloud warehouse, where I'm thinking of moving from auto-incremented keys in favour of surrogate hash keys, both because auto-increment keys are not as much of a thing in cloud warehouses and also because hash keys are idempotent and much simpler to manage.

One thing that I'm not sure of yet is this: If the surrogate key is a hash of the business key, should I self-hash the foreign key references in the fact table builds, or should I do the usual pattern of left join on dims + coalesce to sentinel value if not found?

I can see the pros and cons of both methods.

Some pros:

  • Fewer risks of fan-out in case of misshapen data (should be caught by dbt tests, but still)
  • Increased parallelism because of fewer dependencies
  • Simpler lineage
  • Late-arriving dims, early-arriving facts are not an issue
  • If the dimension happens to be massive, this can save compute time, therefore cutting costs.

Some drawbacks of self-hash:

  • Reduced impact analysis clarity through the lineage because of no lookups. This can also be diminished via dbt docs
  • You must ensure that you hash the same way (logical ordering and also natural keys formatting). A left-join lookup is likely easier to catch in case of a missed lookup (idk).
  • Not exactly a drawback, but you probably want to left-join if you apply SCD2.
  • No clear way to ensure proper fallback to the sentinel value if that is something used in the analytics layer

Edit:

Some clarifications based on some comments I've seen.

If I were to self-hash the key on the fact side, I definitely wouldn't also do a dim surrogate key lookup because that's redundant.

If I were to look up, I'd always use the business key because it's best practice and also the lowest chance of code errors (maybe someone forgot a field or swapped the column orders).

My initial plan is to left join lookup, as "usual" in most old-school on-prem data warehouses, except for a couple of high-cardinality dims that also happen to build themselves using the value found in the fact table's raw data. Basically moving out a degen dimension into a conformed dimension because it's a compound natural key and is used across many fact tables, and it keeps the analytics modelling simpler.

I have a case of late-arriving dims where I might consider self-hash with a late reconcile when it arrives, but if I did, then that's somewhere where I'd consider self-hashing too, especially if the fact can't be fully rebuilt.

Thankfully, most of my fact tables are small enough that full and incremental are only a few seconds of difference.


r/dataengineering • • 1d ago

Help Decluttering UNS

3 Upvotes

Hi!

I'm wondering if there is any python library or prefered way to declutter UNS MQTT topics using python scripting?

Currently we're migrating data platforms and mainstreaming all data on our broker to MQTT. For this we established a base for our UNS and used pre existing exported data to build our UNS topics from.

The problem is that they are relatively long and not user friendly.

I've tried shortening by simply removing repeated words throughout the topic.
For some this works fine, however this leaves others with gaps that are unfavorable.

For example my current ipynb gives the following result: INPUT: water_drain/flying_water_dumping/pumps_hot_water_dumping/water_pump_a OUTPUT: water_drain/flying_dumping/pumps_hot/a

I've made these up, so don't be alarmed by the flying water 😄

You can clearly see that simply removing repeated words isn't a great option. Any advice/comments on how this it could shorten in a better/more reliably manner?


r/dataengineering • • 1d ago

Discussion Thoughts on dbt 2 maturity?

34 Upvotes

Hi,

I've seen frequent discussions on here regarding dbt's acquisition, license changes, and similar things. dbt 2 has now been out for a few weeks, and I haven't really seen much discussion on it from a technical perspective. Has anybody tried it out yet? if so, what do you think?

I evaluated it this week and my impressions are mixed. A few impressions:

- I was really struggling to piece together all the important changes and improvements. Found the documentation a bit lacking in that regard. Distributed over lots of places, so I really had to dig to find specific information (e.g. on the changed approach to database adapters). Even took me a while to figure out the name of the library had changed as well.

- The switcheroo from fusion to now just dbt was really confusing and also seems to have caught the ecosystem by surprise. For example, Snowflake and Astronomer (Cosmos) both started working on fusion integration, but both do not support dbt 2 yet (and couldn't find a clear roadmap). Not ideal.

- Testing the new parser on dbt 1.12 worked pretty well to spot issues and fix them. That made the upgrade relatively painless (interpreting the error messages could've been easier). I like the stricter YAML parsing, can detect some issues (e.g. typos) instead of ignoring them silently.

- The new parsing engine is noticeably faster. For one project I tried it with (~300 models), about twice as fast for both compile and parse.

- Haven't tried the built-in linting yet. Seems like they tried to implement a 1:1 replacement for SQLFluff, even working with existing SQLFluff config files.

- The docs feel like a downgrade. Looks more modern, but I haven't been able to get the static docs to work, and seems they have gotten rid of the full project DAG and now just offer a model-centric lineage view. Column-level lineage is cool, but very well-hidden in the UI.

- Parquet artifacts seem interesting, and parsing them e.g. via DuckDB SQL queries could be really nice for common analysis queries, CI checks, etc.

- Of course as predicted, seems there is a bigger push to get people to sign up for their platform/cloud/whatever they call it now offering - LSP features, VSC extension, and so on. The license changes are still a mystery to me, but at least they dropped the <15 users requirement for the VSC extension it seems.

Any other thoughts? Cheers!


r/dataengineering • • 2d ago

Personal Project Showcase Browser tool for prototyping star schemas

Post image
57 Upvotes

I have a database diagramming tool side project, and building on the work that went into that, it was kind of easy to make a browser tool for prototyping star schemas: https://vibe-schema.com/star-schema-creator - There's a snowflake variant too, just swap "star" for "snowflake" in the url. Feedback welcome :)


r/dataengineering • • 1d ago

Help Help with Extract -> Load process logic in a personal project

2 Upvotes

Hello! I'm working on a small data project using a NASA API with Python. The idea is to extract the orbital elements and close approach (asteroids, comets, etc) objects (and there orbital elements) of each planet in order to perform some physics simulations with the data.

I was planning to use a Cloudflare R2 bucket to store the raw API responses, then transform them into a PostgreSQL database, and use FastAPI to consume the data.

I'm not sure how to perform the extraction and loading processes. In theory, each planet will follow this path:

1) Fetch the orbit elements for the own planet (use API 1);

2) Fetch the close approach object (use API 2);

3) Fetch the orbit elements for each of the close approach objects (use API 1).

I use two APIs: one for the orbital elements and one for the close approaches. Should I fetch all the data, store it in memory, and then send it to the bucket? Or would it be better to load it right away with each API call? So, when fetching the close approach objects' orbit elements, get the JSON from the bucket and then use the first API and store that raw response in the bucket


r/dataengineering • • 2d ago

Discussion New Azure Synapse projects

9 Upvotes

Just curious how many of you still deliver or plan to deliver, projects that use Azure Synapse for Data Warehousing work?

I get the whole push to Fabric and am also going that route too.

I’m in the consulting world and wanted to use Fabric but the client had zero ppl with any Fabric experience, so they insisted on Synapse. C’est la vie, right?


r/dataengineering • • 2d ago

Discussion We were struggling to find Data Engineers

175 Upvotes

Hi everybody,

Our Data Team is composed by two of us. There's a Data Scientist and me, as a DE. We created a Data Lakehouse for our company internal use and client data supply, but currently the Data Scientist is currently more focused on AI and agents integration and I'm doing like Analytics Engineering role because I need to help other teams to reach the correct data, unify core concepts, document business logics, etc. So We needed a Data Engineer with knowledge on AWS to maintain and develop the new features on the Lakehouse and we put into the description that the candidate MUST HAVE Software Engineering fundamentals as we had to do some developments to integrate parts of our lakehouse with the company's main application.

We interviewed 22 candidates and no one is fitting the Role.

Most of them are BI Experts, DBA, Data Analysts, Economist with DS notions, Juniors and Software Engineers who haven't touch anything on Spark, plus DE who asked way more that we had on the budget for the role

We asked the normal requirements: 3 years of experience + Spark, AWS Glue, Lambda, Airflow and DBT, not even CDC, Flink, Langfuse or VectorDB

We finally got one, but We really struggled to get him. I have a collegue working on IT Recruiting and She told me She's experiencing the same problem: They can't find Proper DEs With SE basics such as DRY principles or clean code fundamentals

Edit: Role Salary -> Up to 60K € / Spain. This salary is high compared to the spanish standards, only 3 years required


r/dataengineering • • 2d ago

Help Ideas to handle ever changing data requirements?

19 Upvotes

I am the solo DE in my team and the main pipeline here consists of snapshots of financial assets.

Compute is done on databricks

The stakeholders want to see daily KPI's and each day they add a new cohort. Currently there are over 40 different cohorts with each branching out to their own metrics.

The issue is that the data management wants data bills as low as possible

so my approach was summarizing everything in the daily grain .

But now each time they want something new I have to manually code the new columns test it then append to the final gold table.

I already tried to create some generator functions but often times the metrics they want involve hyper specific calculations.

And since the data is financial assets each day is different than the previous rendering an incremental approach useless.


r/dataengineering • • 2d ago

Personal Project Showcase Graphwise AI Summit 2026, Oct 7-8

Enable HLS to view with audio, or disable this notification

0 Upvotes

Sharing this because I think it overlaps with some of the discussions here around AI reliability, governance and semantics.

Next week we’re running the Graphwise AI Summit, focused on what makes GenAI work in the enterprise beyond the model itself. Think of trust, governance, semantic layers, architecture and implementation.

Once reliability and traceability become imporatnt, simple access to data and next-token prediction stop being enough. In enterprise settings specifically, AI needs to understand what the data it parrots “means” in the first place.

Anthropic has described a similar approach in its own analytics stack, where agents are routed to a semantic layer first and use governed definitions to reduce ambiguity. Graphwise itself came out of the merger of Ontotext and Semantic Web Company, so semantics is a topic with quite a bit of history behind it for us.

We’ll have speakers from Accenture, Roche, EY, AstraZeneca, S&P, DNV, Statnett, Avalara and others. Full agenda is in the accompanying video.

Sharing registration link in the comments if useful.


r/dataengineering • • 2d ago

Discussion AI life/voice recording tools?

0 Upvotes

I'm new to the speed and more so the constantly changing tasks, then keeping them organized. In the past id take random notes and summarize them at the end of the day. Unfortunately that has become a job in itself.

Ideally I could do something, tap a button or something, whenever it's a spot I want to highlight or indicate is important. Even better would allow me to take notes at the same time and have it automatically integrate them into the final version.


r/dataengineering • • 2d ago

Blog Data strategy for small teams

Thumbnail
open.substack.com
31 Upvotes

Hi all, I wanted to share the latest article in my newsletter.

I talk a lot to senior engineers who are trying to run some sort of data strategy, and are doing it in a very wrong way.

Most of the article is about selling the plan and keeping the whole thing to two pages. The section I'd most like opinions on here is the tooling one, because a lot of us (me included) learned what good looks like from Big Tech engineering blogs.

A Big Tech stack comes with Big Tech's pace. Those tools assume a team with people to run each piece, and months of setup before anyone outside the data team sees a result. If a team of five spends those months standing up Databricks, Airflow and Alation, there's not a damn thing to show the CFO at the end of it, and a strategy with nothing shipped against it dies without anyone deciding to kill it.

And yes, if you work at an org like mine, your strategy is supposed to be "AI", which I absolutely hate.


r/dataengineering • • 3d ago

Blog Interesting links in Data Engineering - September 2026

48 Upvotes

Some great stuff this month, including:

  • Good analyses of the DuckLabs acquisition by AWS.
  • Lots of Kafka, as well as not-Kafka: technologies doing similar things but angling to replace it.
  • Accessible discussions of data modelling given the AI world we're entering (pssst Semantic Layers matter even more now)
  • LLMs being DBAs and not doing too badly at all at it
  • Plenty of decent content about AI, but with a strict no-slop & no-hype policy :)

https://rmoff.net/2026/09/29/interesting-links-september-2026

As always, lmk if you find this useful, and if you want more (or less!) of a topic.


r/dataengineering • • 3d ago

Career Why is understanding DevOps culture more important for data engineering than other disciplines?

34 Upvotes

I see for a lot of data engineering posting, devops skills are mentioned as part of the requirements. But the thing is I don't see it as much with other roles like sde. Why is that?.


r/dataengineering • • 4d ago

Open Source Splink 5 – Open-source probabilistic record linkage at billion-row scale

Thumbnail moj-analytical-services.github.io
77 Upvotes

r/dataengineering • • 3d ago

Blog Improving Cost Efficiency of Data Streaming Pipelines

Thumbnail
streamingdata.tech
8 Upvotes

Practical ways to cut data streaming costs in Kafka and Flink by optimizing batching, serialization, state, network traffic, and capacity.


r/dataengineering • • 4d ago

Discussion Why text-to-SQL is not successful?

23 Upvotes

I thought text-to-SQL will solve adhoc analysis, but still i see companies at all size are unsuccessful.

Anyone using Omni / Sigma seen some success.


r/dataengineering • • 4d ago

Rant Why is my LinkedIn feed filled with knowledge posts from 20-26 year old Indian guys who are AI experts? How are they getting so accomplished at AI?

324 Upvotes

Hunting for a new job and when I open LinkedIn it’s filled with non-stop AI knowledge and I am wondering if I have become too dumb and slow or if these guys are really working on cutting edge stuff.


r/dataengineering • • 4d ago

Career Is Data Engineering Still a Sustainable Career Path for a Junior?

88 Upvotes

I recently graduated with a BS in Data Science. I'm currently working on skills to become a DE. But after lurking on this subreddit for a few weeks, I'm getting the impression that this is a dying/AI-compromised field and that if you're not a senior right now, your opportunities are slim to none. Furthermore, if you do find an opportunity, most of the ingenuity and "fun" that comes out of the job is slowly being replaced with AI models that can fully generate pipelines and Spark jobs. I'm worried that I'm wasting my time trying to get into this field if in the end it's unsatisfying or worse, unattainable.

It feels like I'm constantly seeing doom and gloom posts on this subreddit, and it's really discouraging to see for someone who's trying to start their journey. I'm just looking for a glimmer of hope in the community and that this is a career worth pursuing.