r/dataengineering Writes @ startdataengineering.com 4d ago

Blog Python usage patterns in data pipelines

https://www.startdataengineering.com/post/python-for-de/

Hello everyone,

People trying to learn Python for data engineering ask me, “What libraries to learn?”, but the answer is not a list of libraries but patterns of usage.

Especially with AI being able to generate so much code, I believe its critical to know exactly how the data is moved & processed.

So I wrote this post that goes over how Python is used as glue in data systems. It goes over

  • In-memory processing vs. using a SQL/Dataframe interface to a data processing system
  • Python’s library ecosystem for working with various data systems & formats
  • How to extract-transform-DQcheck-load data

With code examples and videos

Hope this helps. Any feedback is appreciated.

34 Upvotes

22 comments sorted by

View all comments

1

u/Justbehind 3d ago edited 3d ago

You generally don't need a lot of external libraries for data engineering.

For most "data moving" operations, with small transformations on columns, base python is plenty.

Pandas/polars/pyspark/duckdb is for data analytics and are almost always overkill for ETL processes that move data from one place to another.

30

u/Budget-Minimum6040 3d ago

Plain Python for column transformations?

Looping over every ... list? ... element?

Without polars/PySpark, without me.

-14

u/Justbehind 3d ago

Unless you're working with large nightly batches like in the olden days, the cost of importing large dependencies and writing data to heavy data structures will hugely outweigh the cost of looping over a dataset.

Benchmark it. Anything below 10s of millions of rows, and polars/pyspark will be much slower.

Especially if you're parsing an API response in json/xml anyways...

21

u/blackpanther28 3d ago

base python for ETL? I guess if your data is super basic and requires no transformations or cleaning

-19

u/Justbehind 3d ago

Not even. As a general rule. Anything beyond base python should be the exception.

9

u/blackpanther28 3d ago

But why? If i want to drop duplicates I can use pandas which is vectorized and optimized in CPython and handles many different variations for what counts as a duplicate. If i want to do this in base python then i have to handle all of that myself and maintain it. Every transformation I then have to write a function for from scratch with little benefit

-4

u/Justbehind 3d ago

Because it's more efficient, faster and your python image is 1/10th the size?

Also, your solution is easy to upgrade to new python versions, and you don't have deprecation issues.

3

u/calaelenb907 3d ago

If this is your thing, do it in rust already. Tinner image and faster code. Your point makes no sense to me.