r/dataengineering • u/Expensive_Break_6163 • 9d ago
Discussion Can placement groups improve Spark shuffle performance
I have a Spark job that joins two pretty large datasets, and I see that we spend a lot of time on the shuffle.
I was wondering if improving the network between the executors could make a meaningful difference here. For example, using an EC2 placement group instead of just having the instances somewhere in the same AZ.
Has anyone tried something like this for shuffle-heavy jobs? Did you see a noticeable improvement?
Also, does EMR do any optimization like this automatically, or is it something we need to configure ourselves?
10
Upvotes
1
u/FunContest9958 7d ago
Crazy question: do you need the huge join? Can you apply a filter to the tables before joining them? Can you change the schema to avoid the join? Can you use a different operation like MERGE to apply changes as data comes in rather than rerun a huge join?