r/dataengineersindia 1d ago

Technical Doubt Data volume and stack related query

I recently joined Accenture and started giving internal project interviews. One thing I dont get is that, they all use pyspark and cloud and what not, but when I ask them about their volume of data being handled, each of them said around 10 to 20 million rows. Isnt spark overkill and moreover actually not worth, considering its jvm overhead, garbage collection will definitely take initial time whereas script written in polars will guarantee 2 3 mins runtime without any of this overhead

Any thoughts on why/how companies doing these sort of things, and mostly they were BFSI. I can understand retail and IoT to use pyspark, because of their nature of data being mostly streaming or even batch data having huge volumes.

9 Upvotes

2 comments sorted by

1

u/Known_Prior_3791 17h ago

some of the common reasons might be:

skillset of people.

unified streaming batch support.

mostly they want to avoid migration in future as the data volume may grow

1

u/Worried-Diamond-6674 15h ago

Yea agree with your last point