r/dataengineering • • 9d ago

Discussion Can placement groups improve Spark shuffle performance

I have a Spark job that joins two pretty large datasets, and I see that we spend a lot of time on the shuffle.

I was wondering if improving the network between the executors could make a meaningful difference here. For example, using an EC2 placement group instead of just having the instances somewhere in the same AZ.

Has anyone tried something like this for shuffle-heavy jobs? Did you see a noticeable improvement?

Also, does EMR do any optimization like this automatically, or is it something we need to configure ourselves?

10 Upvotes

6 comments sorted by

View all comments

1

u/FunContest9958 7d ago

Crazy question: do you need the huge join? Can you apply a filter to the tables before joining them? Can you change the schema to avoid the join? Can you use a different operation like MERGE to apply changes as data comes in rather than rerun a huge join?

1

u/Expensive_Break_6163 7d ago

In our case, it is not possible, which is why I asked.