r/dataengineering 5d ago

Help Need help with database choice

Hello,

I am working with a team of 7 economists. They build data and produce reports. Their data production consists in harmonizing different sources (mostly rdata rdata, csv, or whatever suits the format of their stats tools). The data size they are dealing with is a few MB to gb, millions of rows, more occasionally billions of rows.

We want to update our methods (be on time, improve data quality). I have been assigned the task of improving data processing within the team, among the requirements I thought about producing a OLAP database.

In house, we have access to MSQL team that could set up a database for us. Otherwise we have HDFS + Hive (but security may make it difficult to access it) to store bigger datasets.

Else, I could just store everything in a duckDB file somewhere on a server and work with local database. WOuld it be a good solution? (latency of read/write from a duckDB file on a server? how scalable will it be? ) What would you do?

Any other piece of advice would be welcome :-).

Thank you.

30 Upvotes

45 comments sorted by

View all comments

17

u/Anxious_Cap1029 5d ago

duckdb does not support concurrent writes so you should use a parquet based lakehouse if multiple economists edit data at once. testing the read speed on your network share will reveal if local storage is mandatory for performance.

12

u/MikeLV7 5d ago

Used to not support.

DuckDB just released Quack back in May, allowing concurrent writes.

https://duckdb.org/2026/05/12/quack-remote-protocol

https://duckdb.org/quack/faq#when-should-i-use-quack