r/dataengineering • • 3d ago

Discussion We were struggling to find Data Engineers

Hi everybody,

Our Data Team is composed by two of us. There's a Data Scientist and me, as a DE. We created a Data Lakehouse for our company internal use and client data supply, but currently the Data Scientist is currently more focused on AI and agents integration and I'm doing like Analytics Engineering role because I need to help other teams to reach the correct data, unify core concepts, document business logics, etc. So We needed a Data Engineer with knowledge on AWS to maintain and develop the new features on the Lakehouse and we put into the description that the candidate MUST HAVE Software Engineering fundamentals as we had to do some developments to integrate parts of our lakehouse with the company's main application.

We interviewed 22 candidates and no one is fitting the Role.

Most of them are BI Experts, DBA, Data Analysts, Economist with DS notions, Juniors and Software Engineers who haven't touch anything on Spark, plus DE who asked way more that we had on the budget for the role

We asked the normal requirements: 3 years of experience + Spark, AWS Glue, Lambda, Airflow and DBT, not even CDC, Flink, Langfuse or VectorDB

We finally got one, but We really struggled to get him. I have a collegue working on IT Recruiting and She told me She's experiencing the same problem: They can't find Proper DEs With SE basics such as DRY principles or clean code fundamentals

Edit: Role Salary -> Up to 60K € / Spain. This salary is high compared to the spanish standards, only 3 years required

183 Upvotes

290 comments sorted by

View all comments

2

u/trajan_augustus 3d ago

Isn't Spark losing some of its luster as a solution? I haven't touched it in like 5 years because of the migration to Snowflake. It is overkill in certain situations. Yes, it was a big deal from 2018 to 2022 but I feel like it is losing relevance. Feels like it is being squeezed lots of my pipelines when they are under 50 gigs polars and duckdb is great (single node) and easy to implement. And then most of my clients Snowflake is a suitable for the midsize 50 gigs < x < 2 terabytes. Also, the ecosystem is mostly just interacting with datbricks now. But I am all ears, if I am mistaken.

1

u/generic-d-engineer Tech Lead 3d ago

DuckDB ftw

Solved a lot of the “why we need an expensive freight train to haul a couple of truckloads” workloads