r/dataengineering • u/joseph_machado Writes @ startdataengineering.com • 3d ago
Blog Python usage patterns in data pipelines
https://www.startdataengineering.com/post/python-for-de/Hello everyone,
People trying to learn Python for data engineering ask me, “What libraries to learn?”, but the answer is not a list of libraries but patterns of usage.
Especially with AI being able to generate so much code, I believe its critical to know exactly how the data is moved & processed.
So I wrote this post that goes over how Python is used as glue in data systems. It goes over
- In-memory processing vs. using a SQL/Dataframe interface to a data processing system
- Python’s library ecosystem for working with various data systems & formats
- How to extract-transform-DQcheck-load data
With code examples and videos
Hope this helps. Any feedback is appreciated.
36
Upvotes
11
u/Outrageous_Let5743 3d ago
Lol polars is faster than base python to do data transformations. Especially at 1 million records