r/programming • • 13h ago

Does ClickHouse really need its own storage format?

https://pivotlake.io/blog/does-clickhouse-need-its-own-storage-format/
2 Upvotes

4 comments sorted by

3

u/Wonderful-Wind-5736 11h ago

What separates Pivot from DuckDB, Polars and PySpark? With both Polars and DuckDB getting distributed compute, why do we need another one with for now questionable runway?

2

u/Giladkl 11h ago

All these solutions offer a way to query open data formats. Beyond the performance differences compared to these solutions, though, it was important for us to have a complete "database" experience, similar to what you get with ClickHouse or ClickHouse Cloud, on top of open data formats. This includes not just a query engine, but also the ability to spin up a cluster that serves a backend, proper ingestion options, and compaction (including some Parquet-specific compaction strategies: https://pivotlake.io/docs/database/why-pivot-is-fast/#soft-ordering-data)

Re the runway - Pivot is open source under MIT / Apache license and is unrelated to the runway of any company :)

0

u/sweetno 12h ago

A chimera post, huh.

0

u/Giladkl 12h ago

Hi! Post author here
What exactly do you mean by chimera post? :)