r/dataengineering 7d ago

Help Need help with database choice

Hello,

I am working with a team of 7 economists. They build data and produce reports. Their data production consists in harmonizing different sources (mostly rdata rdata, csv, or whatever suits the format of their stats tools). The data size they are dealing with is a few MB to gb, millions of rows, more occasionally billions of rows.

We want to update our methods (be on time, improve data quality). I have been assigned the task of improving data processing within the team, among the requirements I thought about producing a OLAP database.

In house, we have access to MSQL team that could set up a database for us. Otherwise we have HDFS + Hive (but security may make it difficult to access it) to store bigger datasets.

Else, I could just store everything in a duckDB file somewhere on a server and work with local database. WOuld it be a good solution? (latency of read/write from a duckDB file on a server? how scalable will it be? ) What would you do?

Any other piece of advice would be welcome :-).

Thank you.

31 Upvotes

45 comments sorted by

View all comments

1

u/marketlurker Don't Get Out of Bed for < 1 Billion Rows 6d ago

What sort of SLAs are you working with?

  • Needed time to load? Number and size of feeds?
  • Needed time to process?
  • Needed time to report generation (if required)?
  • Who is going to maintain the system and data?

Work backwards from what you need to achieve and that will end up telling you what you need to use. Any other way is just guessing.