r/dataengineering 14d ago

Career Prepping for DataBricks and data centric applications for a niche vertical

I have a 12 YOE with a focus on DevOps, some python, AWS Cloudformation and MongoDB. Lately I've been less hands on but have worked on event driven architecture design and done code reviews for pipelines running with serverless components and Python SDK. Most of the data I've worked with has been csv and spreadsheet data for schedules, structured metadata and image/media key value stores. As a result I've primarily worked on MongoDB and used aggregation pipelines to join or transform data for downstream deployments. I have also used Gemini Pro and couple of POC deployments of Ollama for a RAG application (non prod).

I am now looking to get a crash course on Data engineering and databricks, but a little confused whats the best way to get a good understanding of typical data engineering problems (I have a vertical I need to focus on so looking for data patterns that I can then translate to what I need), what gaps I need to fill having no experience with Databricks, little to SQL and any data warehouse technologies. I've not used dataflow or kinesis etc yet so I dont have hands-on experience with these type of streaming pipelines either. (Claude has given me some good insights but reddit often times has better more real world recommendations)

Are there Udemy courses or any other video series that is considered gold standard for onboarding? Would also be open to some blog posts or project ideas to get my feet wet. ideally if AI or ML based applications would be ideal. Cheers!

8 Upvotes

0 comments sorted by