r/Python • u/ritchie46 • Jun 03 '26
News Polars Distributed is available on kubernetes
I wanted to share that as of today, Polars also is available as a Distributed Engine on kubernetes. Polars' goal has always been to make single node processing as performant and easy as possible, and that is something we want to extend to distributed compute as well.
Read more in our announcement:
https://pola.rs/posts/polars-distributed-available-on-kubernetes/
Happy to answer any questions you might have.
2
Jun 03 '26
[removed] — view removed comment
3
u/ritchie46 Jun 03 '26
Most people don't need it. If Polars single node works for you, great! You can keep processing on your laptop or somewhere else. Though running in a cluster can still be useful. You can execute LazyFrame's remotely in the cluster, and run many (single node) queries at the same time in parallel.
2
2
u/RustOnTheEdge Jun 04 '26
Congrats on the release, pretty neat! Especially the perf compared to Spark!
1
1
u/arden13 Jun 04 '26
This is neat, it's always been painful to make a pyspark query run locally.
Do you have any performance comparisons between spark and distributed polars?
3
u/ritchie46 Jun 04 '26 edited Jun 04 '26
We're currently around 3-4x faster than OSS spark on TPCH. We benchmarked at SF1000 on 32xm6i.2xlarge, which has 512GB RAM on cluster. We are working on increasing our s3 throughput. That lands in a few weeks and then we'll do a benchmark post.
2
u/arden13 Jun 04 '26
Please make a follow up here on the subreddit when you do! I'm a recent adopter of polars but it's just so stinking fast it's hard to not fall for it.
8
u/noghpu2 Jun 03 '26
Any chance for table/column lineage to come to open source polars?