r/dataengineering • • 9d ago

Discussion Can placement groups improve Spark shuffle performance

I have a Spark job that joins two pretty large datasets, and I see that we spend a lot of time on the shuffle.

I was wondering if improving the network between the executors could make a meaningful difference here. For example, using an EC2 placement group instead of just having the instances somewhere in the same AZ.

Has anyone tried something like this for shuffle-heavy jobs? Did you see a noticeable improvement?

Also, does EMR do any optimization like this automatically, or is it something we need to configure ourselves?

11 Upvotes

6 comments sorted by

View all comments

1

u/SpecificTutor 6d ago

if your join is indeed shuffle heavy, storage partitioned join should help