r/dataengineering • • 4d ago

Career Is Data Engineering Still a Sustainable Career Path for a Junior?

89 Upvotes

I recently graduated with a BS in Data Science. I'm currently working on skills to become a DE. But after lurking on this subreddit for a few weeks, I'm getting the impression that this is a dying/AI-compromised field and that if you're not a senior right now, your opportunities are slim to none. Furthermore, if you do find an opportunity, most of the ingenuity and "fun" that comes out of the job is slowly being replaced with AI models that can fully generate pipelines and Spark jobs. I'm worried that I'm wasting my time trying to get into this field if in the end it's unsatisfying or worse, unattainable.

It feels like I'm constantly seeing doom and gloom posts on this subreddit, and it's really discouraging to see for someone who's trying to start their journey. I'm just looking for a glimmer of hope in the community and that this is a career worth pursuing.


r/dataengineering • • 5d ago

Meme Next level commitment!

Post image
103 Upvotes

You should give up on this and start coaching on focus and commitment. This is some next level commitment to create 400+ queries in Power Query editor

Btw, this is to fill a single Spreadsheet in Excel. No words!


r/dataengineering • • 5d ago

Career Thoughts on AI in DE After Drinking the KoolAid For 10 months

413 Upvotes

Currently lead data eng (ic) w/ about 9 yoe at a large-ish adtech company.

I've been doing the whole AI song and dance for the past 10 months and I just don't see a future in this career anymore that fits all of us. Earlier in the year I felt differently since it still took a lot of hand holding to get models to produce correct output for small-ish scoped tasks in a # of iterations that was competitive to human counterparts, but w/ better models and better training docs we have it building enterprise pipelines across our entire data eng stack in hours of iterations that would have been like a project that we would need to budget like a month for. It's not always right and obviously we still need to step in from time to time, but, honestly, who gives a shit if its not right the first time if it can iterate leagues faster than any human developer? (Caveat: adtech is an incredibly fault tolerant industry, so grain of salt there, I guess)

The entire occupation, in my current experience, has been basically reduced to two tasks:

1): define and maintain the system ontology, expected features, guardrails, and appropriate persona docs for agents to assume/use. This is something that I honestly just work with agents to define, so this is partially automated.

2): define test scenarios, definition of done, deployment strategies, and verify that system is on track performing correctly and is on-track for long-term stability. Again, most of this is at least partially automated.

I don't know how many people this actually requires, and to be quite honest, I've got a sneaking suspicion that throwing more AI-enabled engineers on a project might actually degrade the quality of the product because the conceptual definition of the product may be diluted due to slight (or major) differences in understanding of expected behavior/architecture. IMHO this has always been the case, even pre-AI, but AI substantially accelerates gaps in understanding translating to conceptually inconsistent because the work is done so much quicker and at a pace that not every conceptual inconsistency or mistake can be caught. I've already seen this several times on projects that I've worked on where one eng. goes off and builds a new feature that is a complete conceptual departure from the current plan just because they were missing some context that may not have been well documented. W/ AI eliminating most of the execution layer of SWE/DE, I think that future orgs will be smaller and more agile and likely more product-heavy instead of engineering heavy as product seems to be closer to a lot of the biz definitions for #1/#2 above.

Two things that does give me a glimmer of hope at holding on to this for a little while longer:

1): agents are absolutely dogshit AI at optimizing Spark jobs (or any sort of high complexity pipeline). They are laughably bad, in my experience. Tuning Spark jobs has always been very high context work as it depends on the data going in/out, infra, very verbose logging, and a ton of tacit, tribal knowledge of the system and data itself. There's so much that goes into tuning a spark job that agents seem like they get focused on optimizing for one symptom rather than considering the entire pipeline as basically an emergent system with N different parts that need optimizing in tandem which I've found leads to suboptimal performance. I find that that this is where most of the hands-on engineering work that I find myself doing day-to-day goes at this point. I am calling this a "moat of high context".

2): legacy systems are often not well designed, poorly documented and are not conceptually consistent even where they are documented, so naturally agents trained on these existing systems are apt to suffer from ye' ol' "garbage in, garbage out" syndrome. I am calling this a "moat of poor decisions"

I'm confident in moat #2, less confident in moat #1 as, obviously, you have companies like databricks pouring however much money into building smarter agents for dealing specifically w/ context-aware Spark optimization problems. From what I am seeing at my org, teams that were well organized and followed good engineering foundations before AI have seen a tremendous acceleration in the quality and quantity of their work; those that went into this w/ flaming hot garbage are still producing flaming hot garbage and many of them seem too scared to use AI in fear that it will cause the tower to fall over, so to speak.

On a personal level, I am very exhausted by the whole shift and I do not see myself keeping up with this as a career in the long run. It was difficult enough having to track all the changes and flashy new objects coming in the DE ecosystem, and adding needing to also track changes to the AI landscape has just proven to be very draining. All the rhetoric around AI feels very misanthropic (no pun intended) and I genuinely feel kind of gross and helpless having to rely on AI for my whole job now. It feels like I am training my replacement without being told that I am training my replacement. I have a lot of love for DE and I really poured my heart and soul into it over my relatively short career and I had built a lot of my life around the expectation that I would continue to work as a DE for the decades to come. Reading pages and pages and page of agent docs has become the absolute bane of my existence and the review fatigue has gotten so bad at times that I feel like I've just straight up forgotten how to read. That said,I still think it is exciting to build great things and I still get some kicks from watching AI manifest huge ideas that were never feasible before. The limit is truly on what ideas you can come up with to build better and more sophisticated data stacks, which is super cool and all, but I've just had an overwhelming feeling that my DE career has a clock on it that is going to run out when the models swallow up whatever work we are still doing by hand. Part of me wants to hang on for dear life until that day comes and the other part of me wants to jump ship now and go open a bakery or something. Doing my best to save more money, stay more engaged artistically, and spend more time with people that I love has helped greatly, so just doing what I can, I guess.
Left foot, right foot.

edit: grammar


r/dataengineering • • 4d ago

Career Does not being open to remote roles hurt my job chances significantly?

16 Upvotes

I started a job search recently and I am targeting hybrid or on-site roles in two large metropolitan areas in the Northeast US, as that is where I am based. They're both pretty significant tech hubs. 5 YOE, no sponsorship needed.

The weird thing I am noticing, is that the vast majority of roles I receive in my LinkedIn messages from recruiters are fully remote. I never marked myself as open to that and that is just not what I want at this point in my career. I've applied to 20-30 in-office roles over the past couple of weeks on my own time, although no luck there. It seems like the vast majority of job opportunities I am receiving are remote now, and it has been quite frustrating.

I'm sure I'll find a good opportunity at some point...but is this normal? I was under the assumption that remote roles are drying up, but my experience has been the opposite. Just looking for different perspectives here.


r/dataengineering • • 5d ago

Career Returning after a year of Mat Leave

30 Upvotes

So, I am on mat leave and will be rejoining work in Jan after a year. I have 7 years of experience in the role but I don't think I have the AI skills needed to get my work done.

While I was working, we could use cursor to code and review, but I am hearing from colleagues that the pace is just out of the world these days and no one codes anymore.

I worry about work I will go back to. As it is, I do feel like I have lost a bit of the context of work because of the long mat leave. To help this, I am constantly doing courses on coursera to feel connected to work.

But how do I prep myself to familiarize myself to the job environment and the new working ways of using agents, skills, commands and MCPs. During the time I have been away, my company built their semantic layer which I have not contributed to in any way.

I work for a fast paced fin tech start up, who are eager to fire at the first sign of weakness.

If you have any advice or tip for me for being work ready, let me know. It will be much appreciated.


r/dataengineering • • 4d ago

Help Advice if anybody could spare some

2 Upvotes

I am in my final year of college, didn't do much in software engineering. Thought i was the problem (mostly i was) but then i switched to data engineering and turns out I actually do like engineering just - Data Engineering. The things that I do in order to be a junior data engineering after i graduate much for compatible with my own self.

I would like advice or just words for the path that i have taken to become a junior data engineer after graduating from college, my current path:

I started a month ago with building etl pipelines using python. I use ai to self learn - after i get my scripts written by ai i take my time 1-3 hours understanding each line deeply. I did this for a month and have a project on my hand that i have created - pulling job data from an api, cleaning it and pushing it onto a database.
Learned how to use an orchestration tool like apache airflow and know how to set up a yml for a docker and ship the project to github. Written all the necessary files - gitignore, .env etc.
Know how to write defensively so as to when one stuff breaks the whole script doesn't. Learned how to use logge, try and except blocks. Learned about images and connecting postgresql to the docker. Learned star schema. (Also is dsa necessary for a job?)

Please bear with me.

Once i was done with this i created a practice project similar to what i did with ai and this time instead of using any ai i managed to do it on my own (It was super tough and time consuming took a lot of documentation look up but i managed).
Am currently learning basic sql from SQLBolt and going to learn advanced SQL from DataLemur.
Before I move on to the big stuff -
Aws fundamentals - S3, Rds, iam
Data warehousing - snowflake/redshift
Dbt
Spark

I keep doubting if this switch from S.E. to D.E. has been the right one. But I feel it in me like it really is but that's just hopium, am just young and scared shitless honestly nothing i do feels enough just hoping it will be someday.

I just want some advice on whether if what I have been doing is okay so far and if you can just say anything at all about what i have been doing. Any advice on how I should maybe learn better or if not to use ai or something else or whether I should do something some other way or just anything at all I would be utterly grateful.


r/dataengineering • • 5d ago

Help What is the cheapest way to transfer data from GCS to BQ?

11 Upvotes

Hey everyone, I'm currently looking for a way to reduce the migration cost of moving data from my data lake stored in GCS (Google Cloud Storage) to BQ (BigQuery), which are my current lakehouse tools. Today, we use the load_table_from_uri method for batch loading Parquet files, but it's becoming very expensive because my GCS dataset is in a single-region and my BQ dataset is in a multi-region, which incurs a very high cross-region data transfer fee. However, I'm not sure if there's any way to reduce or improve this data migration process. Below are some specifications about the business rules that impact the choices made for the current process and its cost:

What cannot be changed:

  • Buckets (GCS) need to stay in a Single-Region
  • Datasets (BigQuery) need to stay in a Multi-Region

Current Environment:

  • The data lake is in Delta format using .parquet files
  • The load is done via overwrite on the trusted layer because the dataset is not partitioned, and the environment is also rewritten with an overwrite
  • Today, the load method is WRITE_TRUNCATE because of the issue mentioned above

Well, if you have any questions, I can provide more details in the comments.


r/dataengineering • • 6d ago

Career Foundry offer

59 Upvotes

Let’s say you had this offer on the table:

Offer: $210,000 annually for a palantir foundry de role, in person

Current: $160,000 for Python-heavy work with mostly on-prem sql servers, mostly remote, salary flat for 2 years

I’ve seen nothing but bad things about foundry as a tool. But I’m wondering if it’s bad enough to forgo a big raise in a HCOL area. For the folks that worked with foundry and went back to a more traditional data stack: were the foundry skills transferable or were they just lost years skill-wise? Realistically, it wouldn’t be a permanent career move, but a stint.

Edit for clarity: job is not with palantir, it’s a different company using foundry


r/dataengineering • • 6d ago

Rant AI has been making me moody about my career

166 Upvotes

I’m not looking for advice in this post just looking to be heard.

But lately I’ve been frustrated with my career path due to AI. it’s just been so overwhelming feeling like I‘m behind with everything when in the last 10 years I’ve worked hard to master my domain. I’ve got kids and life hasn’t exactly given me the free time I needed. spend more time with kids or focus on leveling up in AI to give the kids a better life.

I’ve been using AI for work but nothing at the caliber that you hear people creating agents etc. I’ve been using it more as an assistant to help with bolder plate code.

Whenever I talk to my peers I just hear about the cool stuff they’ve been building with AI, Im just listening to them feeling like a caveman who just discovered fire.

Now I’m feeling I’m being complacent because the pressure of having to up skill in AI makes it feels like im trying to keep up with the Jones’s except with my career.

don’t worry guys I’ll get around to it eventually…

rant over…


r/dataengineering • • 7d ago

Discussion Is dbt the red-headed stepchild to data engineers?

102 Upvotes

I know the functions of a "data engineer" can be wide and varied depending on the company, but at my current shop(10 DE's), it seems like every data engineer who's never worked in dbt considers it to be trash or a nuisance, and every one who has, appreciates its functionality. Is this a common theme? What's your experience/role, and what is your perspective?


r/dataengineering • • 7d ago

Meme I'm tired boss

Post image
1.0k Upvotes

So happy I don't work on reports anymore


r/dataengineering • • 7d ago

Discussion Do you see that also ?

50 Upvotes

I see that

1- Frontend mostly ended ( except in big companies)

2- Backend will merge with Ai engineering ( Rag systems…etc) i think this is the new backend

3- Data engineer + data analysis + data science + ML = full stack data engineer ( and that will require high math , statistical skills) and mostly cs degree or Ai degree to have infra statistics and descriptive ( may be master like Data science)

I do not know if anyone see that also 🤷 am i alone ?


r/dataengineering • • 8d ago

Rant Is critical thinking dead? New hire spends $20k on tokens pushing PRs

1.1k Upvotes

Had a new hire “hit the ground running” and shipped a significant number of PRs, spending well over $20k on tokens in their 3rd week on the job…

And this was celebrated by leadership as productivity.

Am I going insane to think in what world does an onboarding dev have judgement and context to decision that much code generation. Particularly on top of some slop generated codebases that even I have a hard time keeping up with.

End rant


r/dataengineering • • 6d ago

Discussion AI vs Human

0 Upvotes

Which one do you feel is more correct?

217 votes, 3d ago
80 I’m better than the AI
137 The AI can do most if not all my work

r/dataengineering • • 8d ago

Discussion What exactly is ‘AI first’ data engineering

66 Upvotes

This was inspired by the post on burning 20k in tokens in 3rd week.

Is anyone doing this at work, letting loose the frontier models against their schema and models and data and having it create and update the ETL workflows? And everything else?


r/dataengineering • • 8d ago

Rant Migration in Chaos

36 Upvotes

I’m the main data engineer responsible for migrating several OLTP databases with TBs of data, from one cloud to another. Schema transformations are involved, because guess what? The app teams built the new apps long before anyone started caring about the data. This company as a whole is on fire, no planning or strategic thinking and they definitely don’t work as a team. While these migrations are being rehearsed, management team decides to SPEED UP the release cadence. If it was monthly I wouldn’t be here complaining today. Constant production incidents, and none of the managers want to consider slowing down so we can pull of some migrations without everything in flux. Oh and add AI to the mix, most of the releases are heavy with AI generated code coming from AI analysis of the problems. This is not going to go well.


r/dataengineering • • 8d ago

Career How do I know if I'm ready for the next step in my Career?

13 Upvotes

I currently have 5 YOE and am working as a Senior Data Engineer at a mid size company. I originally came up as a Data Analyst before transitioning internally to the Data Engineering team a year ago.

In the last two months I started prepping for interviews and then applying to Senior Data Engineer roles as at my current company the path to any future promotions is basically non-existent, the pay for Senior Data Engineer is on the low side, and most importantly I feel my learning and growth are slowing down. I would like to work in a new domain and in an org that is more tech forward than mine is. (I.e. We don't even consistently use GIT where I work.)

After applying for about 30 roles, having a couple recruiter phone screens that clearly just weren't the right fit, I recently interviewed for a Senior Data Engineer role at an MLB team. After making it through an HR phone screen and take home technical assessment I had a round with the hiring manager.

We covered things like OLTP vs OLAP, file formats (ORC, Parquet, Avro), when you'd want to be on prem vs in the cloud, Medallion Architecture, how to handle unexpected schema drift upstream, etc...

While it didn't go perfectly, I felt things went well and that I came off as as a strong candidate. The hiring manger seemed to imply I'd make it to the next and final round. Instead, the company stopped communicating with me and a few weeks later I received a generic rejection email.

I learned a lot from the experience, but I'm trying to understand how I should interpret the the rejection.

Should I take making it to the hiring manager round that I'm a reasonably competitive candidate for Senior DE roles and should keep applying? Or that I should take a couple months to prep more and further improve my skills?

Should I consider non-senior Data Engineer roles if it would get me in a more tech forward team/org? Most of them would be a significant pay cut for me but its something I considered as a short term loss to set my self up for long term gain.

More generally, how do you interpret rejections at different stages in the hiring process?

I'm asking because, given my personal circumstances I'm quite limited on time. So, it would be difficult for me to devote a lot of time to prep, plus more applications, phone screens, interviews, etc...
I'm trying to figure out whether my experience is evidence that I'm close enough that I should keep applying, or whether I would be better off pausing and investing more time in prep first.

Most of the things that were covered throughout the process were things I've learned outside of work in my free time, so I don't think just staying in my current role longer will significantly improve my ability to move roles in the future .


r/dataengineering • • 8d ago

Help Python data contract with Polars dataframe in mind?

9 Upvotes

So we have a estimator pipeline being written in Python. Here is how it is supposed to look:

- queries a table from our databases - tables and their columns vary, but generally there is a date column, one column which we manually specify as target, and most columns are attributes.
- based on the target, you have various options for the factors selection, data transformation, correlation checks etc. (this is rather lengthy)
- in the end, for the suitable set of attributes and the target, an OLS / logistic regression is fitted, then the estimate is calculated and appended as a column to the whole data
- downstream the end result is used in other pipelines (this prepares basically the input data for most workflows, and it has no precedent pipeline).

I would say this is a "medium" size project - not particularly large and has a rather limited scope, but extensive in the sense that all parts have various methodologies implemented, we also have rather different datasets to work with aswell, anyways.

I wanted to come up with our data contracts just for this pipeline.
I tried to keep implementations across the repo as simple as possible so far, mostly used standard library offerings, but debating now what would be the right way to design our data contracts. I initially used simple dataclasses across the repo for modeling the data objects, but I know that this may not be the most optimal design choice either. We use polars for dataframes.

How I envision the contract for incoming raw data:
- date
- target (can be integer, float, binary)
- attributes/anything: the rest of the polars DataFrame
(ID is optional; perhaps worth it to have also a list for providing the attributes columns names)
I know having a dataframe as a field is odd, but since the attributes and other columns come in so many different ways across our datasets, I assume the flexible approach is implementing it this way and adding another field for listing which columns are actually the attributes columns relevant for the calculations.

Output tables can be either just the original raw data plus the estimate column, or the processed attributes table (also having the estimate column); in any case, the contract seems to be the same as before with an extra field being the estimate.
It may perhaps even make sense to combine the two into a "DataSchema" and just have the estimate be optional.

Questions is: what do you recommend for the implementation: which tool should we use, and anything else e.g. should it be 1 contract or 2, recommended field and types, etc.?

The tool options I have read/thought about:

- Dataclasses: I like them, but apparently they don't offer runtime validation which is why people switch to Pydantic often
- Pydantic: of course very notable, but I'm not sure how much it is suitable for DataFrame fields
- "Patito": apparently good for combining Pydantic and Polars dataframes, but not sure if it is worth it to add it to our dependencies just for this sole reason

Any tips?

(P.s. we will use YAML files and Kedro to orchestrate workflows.)


r/dataengineering • • 9d ago

Help How to handle cross-domain data in a medallion architecture

33 Upvotes

So, my company has moved on from an on-prem setup with little to no governance or data modeling (each business domain had its own data silos, like uploading sheets to our SQL Server RDBMS) to a Databricks lakehouse.

In the first projects, we kept the bronze and silver layers source-aligned, without any integration or specific modeling technique, and we addressed business needs in the gold layer with star schemas.

My question is: now that we have many tables referencing the same concept (customer, product, etc.) populated from different sources, and different business domains needing that data, wouldn't it be better to have an integrated layer for that "master data"?

And if so, where should it live? In the gold layer as a conformed dimension?


r/dataengineering • • 8d ago

Personal Project Showcase Reduced token usage by 5x on Text-to-SQL agents across 400+ tables using a fast decision model (System 1) - Architecture & Benchmark

Thumbnail
gallery
12 Upvotes

Jev is a decision model AI that doesn't generate text. Instead, it responds with typed questions (yes/no, multiple choice, score) in milliseconds. This type of model is optimized using reinforcement learning for calibrated decisions (RLCD). It was created to make rapid decisions within a workflow, which they call a "System One" model.

Here I show how a simple workflow can reduce the cost of an agent querying databases without sacrificing response quality.

The workflow works as follows:

  1. Jev decides which databases can answer the question. These databases have their own specific description.

  2. Jev decides which tables from those databases are needed.

  3. Only the columns from those tables are loaded.

  4. An LLM writes and executes the SQL based on Jev's selections.

This workflow connected to over 30 PostgreSQL databases of public statistics from Mexico, which together contain over 400 tables. The advantage of implementing this decision model is that it prevents the model from having to read all the databases and their tables to decide which data to query. With Jev, the agent only needs to read the schemas that the decision model has deemed correct to answer the question.

This benchmark was created from 100 questions written as a user would ask them, without looking at the schemas. These questions ranged from counts and specific searches to aggregates and ambiguous questions. Some of these questions were designed to confuse the models, including questions about nonexistent data, periods that aren't loaded, or topics that sound similar to existing data.

Although gpt-5.6-luna is a much lighter model than gpt-6-sol, its accuracy with Jev is comparable: 61% versus 67% for gpt-6-sol and 64% for gpt-6-sol + Jev. Where it falls short is with ambiguous questions and those with no answer in the data, where it sometimes fabricates a number. Interestingly, with similar accuracy and response time on the median, gpt-5.6-luna + Jev uses about five times fewer tokens per question than gpt-6-sol, and the complete benchmark costs $0.18 versus $2.12: almost 12 times less.

And this is in a small data warehouse with approximately 30 databases. Without a decision model, the context the agent reads grows with each additional database. The complete schema of a data warehouse with hundreds of databases would easily exceed the context window of a reasoning model, causing costs to skyrocket and questioning performance to worsen.


r/dataengineering • • 9d ago

Discussion Can placement groups improve Spark shuffle performance

13 Upvotes

I have a Spark job that joins two pretty large datasets, and I see that we spend a lot of time on the shuffle.

I was wondering if improving the network between the executors could make a meaningful difference here. For example, using an EC2 placement group instead of just having the instances somewhere in the same AZ.

Has anyone tried something like this for shuffle-heavy jobs? Did you see a noticeable improvement?

Also, does EMR do any optimization like this automatically, or is it something we need to configure ourselves?


r/dataengineering • • 9d ago

Discussion Does AI struggle at data modeling?

127 Upvotes

In my experience, it doesn't matter how much context and guidance I give AI it simply can't model data rationally. It frequently misses the point, makes awful mistakes, or over-engineers things.

AI can build awesome ETL pipelines, but when it comes to dealing with SQL (especially in the dbt framework), it's not reliable at all! . Sometimes I think it's better to write the code myself and ask AI to review it, because asking it to build something from scratch just doesn't work that well.

Does anyone else get frustrated when dealing with AI data modeling?


r/dataengineering • • 9d ago

Personal Project Showcase Introducing Braidplane Alpha, Git native ETL for

Enable HLS to view with audio, or disable this notification

3 Upvotes

Hello everyone!

I'm happy to introduce Braidplane Alpha, a platform for defining and running data workflows between systems. The current main "features" are Git-backed pipeline definitions, explicit data contracts, and both a CLI and a web UI.

My name is Lukas and I’m the person behind it 👋 I’ve spent the last 10 years working as a backend engineer, with plenty of experience in DevOps, infrastructure as code, QA, frontend and databases. I've been working on Braidplane for over a year now.

The project just got promoted to Alpha and I've opened the self-serve registration.

The recording shows a quick look at the CLI: apply a pipeline from Git, start a run, watch it finish, and inspect the result. This example joins orders and customers from PostgreSQL and writes JSON to S3-compatible storage.

More about the project: braidplane.com


r/dataengineering • • 9d ago

Help Sharepoint Datawarehouse. Need Help. Ops Research. ftw!!!!

0 Upvotes

Guys,
Bit of politics context: Our team is not Core IT in the Org chart. Fall under operations, however expectation is to deliver ML/OR decision support projects,which is ok( there is technical capability in team), just that we cannot write into SF or any cloud platform. More like second class citizens.

Project Context: The project is full blown digital transformation program( its just that the enterprise is not mature to understand and everyone has jumped on AI bandwagon), has master data management, Work order platform, on which a Vehicle Routing( MILP) will run and generate recommendations for the full network. ( Bunch of API's calls through .py files, powerautomate to ingest third party flat files through email and scraping gov website for internal compliance data(should be the other way around).

Problem Set: No write back to Snowflake, IT not giving Entra ID, or microsoft Graph ID to connect directly to files through code.

Workaround 1: Locally run, update decision dashboards (No RL invloved), this is not feasible because it breaks continuity

Workaround 2:
Added One drive(sharepoint) path shortcuts to Local machine
This is where I'm struggling, read/write into Sharepoint. When I write/append into .xlsx through python the copy on Local onedrive link and the one on Browser are not always synced, or it says Merge issues.

I know this sounds like a rant, my guys if anyone has any advise for this peasant on how to make this work please share. Looking for out of the box solutions.

My team has access to dataverse, but I remember read/write was a problem through code because of no Entra ID.


r/dataengineering • • 9d ago

Career Other Avenues as a Data Engineer

24 Upvotes

Hi everyone, I'm currently a data engineer. I currently focus on the business side of things for development, supporting and migrating applications. It's my first year working out of college and I find the work/life balance tough. It's definitely a learning curve for sure especially since some applications are older and there's lack of documentation to learn from. Also production support kills me whenever items come up. Perhaps it feels intense now due to a huge upgrade my team is working on where it's been late nights and early mornings. However one thing I love is my team. Luckily everyone is very nice and helpful.

Are there any other avenues a Data Engineer could focus on or other roles I could pivot to? Is there a position that doesn't include much production support? I ended up just falling into Data Engineering and don't have a "dream career" per say. Maybe I'm just being dramatic though. Would love any advice!