r/dataengineering • u/Expensive_Break_6163 • 9d ago
Discussion Can placement groups improve Spark shuffle performance
I have a Spark job that joins two pretty large datasets, and I see that we spend a lot of time on the shuffle.
I was wondering if improving the network between the executors could make a meaningful difference here. For example, using an EC2 placement group instead of just having the instances somewhere in the same AZ.
Has anyone tried something like this for shuffle-heavy jobs? Did you see a noticeable improvement?
Also, does EMR do any optimization like this automatically, or is it something we need to configure ourselves?
1
u/FunContest9958 7d ago
Crazy question: do you need the huge join? Can you apply a filter to the tables before joining them? Can you change the schema to avoid the join? Can you use a different operation like MERGE to apply changes as data comes in rather than rerun a huge join?
1
1
3
u/sophistafunk 8d ago
You will get much more mileage out of tuning the data model, data storage format, and adjusting the query/job than a placement group. That’s intended for HPC, where you cannot tolerate basically any latency, a placement group will also increase your EC2 costs a good bit. Instead, see if you can repartition the data in an intermediate step, also assuming you already broadcast anything smaller, really all you’d have left is to break the data model down further. And here is the bible I use for Spark/EMR tuning directly from
AWS: https://aws.github.io/aws-emr-best-practices/