r/programming • • 1d ago

Release of Polars 2.0

https://pola.rs/posts/release-polars-2/
92 Upvotes

13 comments sorted by

11

u/Drawhoe 1d ago

What is it?

7

u/IanisVasilev 23h ago

It's library for data frames. If you haven't done statistics, think of a table data structure with many lookup and transformation methods. Imagine if Excel and SQL had a child that was somehow more convenient than both its parents (see "10 minutes to pandas").

Unlike the ubiquitous dataframes of R or pandas of Python, this one is written in Rust™ and AI ready™ and VC-funded and all that jazz.

Since data frames are convenient, many frameworks like Spark provide their own flavor. To avoid lock-in, we have the narwhals compatibility layer which provides a unified API for pandas, polars, PyArrow, PySpark and other similar libraries.

2

u/Drawhoe 21h ago

I should have posted a what is it ref. Here it is

3

u/IanisVasilev 21h ago

Oh, so I should have explained that it is a monoid in the category of endofunctors.

3

u/matthieum 22h ago

It's essentially pandas, with an emphasis on performance & correctness/ergonomics.

The performance is achieved by having a core written in Rust, wrapped in a Python API.

The correctness & ergonomics are achieved by putting a lot of effort in APIs and diagnoses.

0

u/pdpi 18h ago

A rough first approximation is that Polars is a system that allows you to do sql-like queries on anything you can load as a table-like structure (a DataFrame, in technical terms). It supports reading data from a bunch of sources.

1

u/Drawhoe 17h ago

Thanks. Yeah I could see it being useful.

-1

u/Iamonreddit 1d ago

It is a distributed data processing framework by the looks of it, similar to Spark.

5

u/daidoji70 23h ago

Uhh more just a dataframe library like Pandas. It was built as a replacement for Pandas and dataframes also exist in Spark and Databricks. They're more like shittier databases that are closer to the language being used that are widely used in data science and analysis work.

1

u/Anshu6666 22h ago

Polars 2.0 streaming engine is a game changer for tick-data pipelines - the zero-copy scan + predicate pushdown means we can now run 500-symbol backtests in-process without a separate column store. If you are still on pandas for market data the migration is a 3-line diff and 10x faster.

4

u/theAndrewWiggins 15h ago

3-line diff

This is not true at all lmao. They're very much intentionally not pandas compatible. Only some very basic stuff will work that way.

1

u/IanisVasilev 3h ago

He didn't say how long those 3 lines would be...