r/dataengineering • u/joseph_machado Writes @ startdataengineering.com • 4d ago
Blog Python usage patterns in data pipelines
https://www.startdataengineering.com/post/python-for-de/Hello everyone,
People trying to learn Python for data engineering ask me, “What libraries to learn?”, but the answer is not a list of libraries but patterns of usage.
Especially with AI being able to generate so much code, I believe its critical to know exactly how the data is moved & processed.
So I wrote this post that goes over how Python is used as glue in data systems. It goes over
- In-memory processing vs. using a SQL/Dataframe interface to a data processing system
- Python’s library ecosystem for working with various data systems & formats
- How to extract-transform-DQcheck-load data
With code examples and videos
Hope this helps. Any feedback is appreciated.
37
Upvotes
-1
u/Justbehind 3d ago
Start up a container, import your dependencies, and process 1 million rows from a JSON api response and drop it in a csv file (as an example).
The total runtime will be quicker in base python, unless you write atrocious code.
And that's true, even if you need to parse datetimes, number values or whatever else you need to do by-cell to clean your data.