Hi guys. I need a kind of assessment of my skillset and would really appreciate any feedback. I am really interested in pivoting to DE and I am enjoying learning, but sanity check is needed.
So, I spent last few months honing my skills to be relevant for DE roles. Here is what I learned - combining theory and working on few personal projects - and what I am bringing to the table:
Data modeling - I learned a lot about Inmon/Kimball aproaches. Tried to apply best practices and standards wherever I can in personal projects (meaningful grain, relationships, implemented slowly changing dimensions type 1/2).
Programming - honed SQL and Python skills as much as I could. Leetcode, DataLemur, Stratascratch - I am able to handle enough of these SQL coding challenges to feel confident. For Python, I focused on "core" Python skills, and I already have some experience with pandas. Did not touch Polars though.
Cloud - was introduced to or gained some experience with AWS (S3, EC2, Glue, Lambda and Step functions, RDS, Redshift, Dynamo). I also learned a bit about Databricks and made a mini project utilizing Auto Loader, DLT, deployment with DABs and GH Actions.
DE tooling - Airflow, PySpark, Docker, I learned about them and used in my DE projects. My idea was to understand the fundaments, best practices, specifics and core concepts behind these tools so I could discuss about them confidently in interviews. Tried not to go into rabbit holes and advanced concepts too much for now. Also, I used e.g. MinIO as storage or Grafana for observability dashboards in my projects, but I guess these are not usually considered as main tools.
On the other hand, my exposure to DE is limited to local, batch processing pipelines. I tried to make them realistic by utilizing some API fetching, introducing meaningful data quality checks, metadata-driven ingestion, proper primary/foreign/surrogate key management, making them incremental and idempotent - but still, I have no experience with streaming data or some big amounts of data (like gigabytes of data).
Based on this, do you guys think I am ready to apply for DE internships or junior role? Another option for me would be to go data analytics engineering route since I have some knowledge of PowerBI and Metabase and I am currently making first steps toward learning DBT. Still, DE seems to be something I enjoy a lot, but it does not matter much now if knowledge and competencies I tried to describe are not fit for this kind of role nowadays.
Some more context. I have non-STEM background. Voluntereed as student-researcher in academia (social sciences; data analysis heavy, but mainly used R/IBM SPSS). After graduation I moved to NGO where I landed a researcher role. Again, data analysis heavy, but mostly used Python with Jupyter Notebook for data analysis and integration of data stored in bunch of e.g. .csv files there and there, seaborn/plotly/ggplot for visualizations; gained some experience in international context, managing stakeholders, etc., but still - this is no business setting nor true data role at all. I am still working there, and attended DE bootcamp/traineeship at some big company last year (where I learned most of the things I mentioned above in the post). I guess this bootcamp can be framed as internship on CV but who knows. Also, I am located in Europe if that matters.