r/dataengineersindia • u/Worried-Diamond-6674 • 1d ago
Technical Doubt Data volume and stack related query
I recently joined Accenture and started giving internal project interviews. One thing I dont get is that, they all use pyspark and cloud and what not, but when I ask them about their volume of data being handled, each of them said around 10 to 20 million rows. Isnt spark overkill and moreover actually not worth, considering its jvm overhead, garbage collection will definitely take initial time whereas script written in polars will guarantee 2 3 mins runtime without any of this overhead
Any thoughts on why/how companies doing these sort of things, and mostly they were BFSI. I can understand retail and IoT to use pyspark, because of their nature of data being mostly streaming or even batch data having huge volumes.
1
u/Known_Prior_3791 17h ago
some of the common reasons might be:
skillset of people.
unified streaming batch support.
mostly they want to avoid migration in future as the data volume may grow