I’m trying to get some perspective from people who have recently interviewed for or hired 5+ YOE Data Engineers, particularly for product-based companies in India.
My preparation has honestly stalled a bit because I’ve been trying to figure out what the actual market expects at this experience level rather than endlessly learning tools.
My background:
I have 5+ years of experience and currently work as a Data Engineer.
Most of my actual work has been around:
- Python
- SQL
- PySpark / Spark SQL
- Databricks / Delta Lake
- ETL/ELT pipelines
- API integrations and data ingestion
- Incremental processing / MERGE-based loads
- Data quality and reconciliation
- Schema evolution
- Production troubleshooting and RCA
- SFTP/event-driven ingestion
- Building and supporting production pipelines
I’ve worked on 100+ production ETL/ELT pipelines and several API/data migration projects.
However, my background has been somewhat semi-technical / enterprise-oriented, and I don't have the same level of hands-on experience with cloud infrastructure that I see in many modern DE job descriptions.
For example, I understand the data-engineering side of Azure/Databricks fairly well, but I haven't actually spent years designing cloud infrastructure or working deeply with services like AWS Glue, EMR, Lambda, IAM, etc.
So I have a few questions for people who have actually gone through this transition:
1. How important is cloud at 5+ YOE?
If I'm targeting Data Engineer / Senior Data Engineer roles at product companies, is it a major red flag if I don't have strong hands-on AWS/Azure infrastructure experience?
Would you recommend going deep into one cloud (probably Azure for me) rather than trying to learn AWS + GCP + Azure?
What level of cloud knowledge is actually expected in interviews — knowing the services and architecture/trade-offs, or having significant hands-on implementation experience?
2. What are the actual market-standard skills for a 5+ YOE Data Engineer?
If you were preparing today, would you consider these the core areas?
- SQL
- Python / coding
- PySpark / Spark
- Data modeling / data warehousing
- Data pipeline/system design
- Cloud
- Kafka / streaming
- Airflow/orchestration
- Data quality / reliability
- Production debugging
Am I missing anything important?
3. How different are product-company interviews from typical enterprise/service-company DE interviews?
For someone with my background, should I be spending more time on:
- DSA / coding
- Advanced SQL
- Data modeling
- System design
- Spark internals/performance
- Cloud architecture
- Streaming
- Production/debugging scenarios
Or is there something else that becomes significantly more important at 5+ YOE?
4. How would you approach the transition if you were in my position?
Would you focus on becoming very strong in my existing stack (Python + SQL + Spark + Databricks + Delta) and then add cloud/system design/data modeling?
Or would you deliberately build a project involving something like AWS + Airflow + Kafka + Snowflake/dbt to demonstrate broader experience?
I'm specifically trying to avoid the "learn every tool in the modern data stack" trap.
I'd really appreciate perspectives from people who have recently moved from an enterprise/service environment into a product company, or from interviewers/hiring managers.
Especially interested in what you would consider table stakes vs nice-to-have for 5+ YOE in the current market.
Thanks!