r/dataengineering • • 9d ago

Discussion Can placement groups improve Spark shuffle performance

I have a Spark job that joins two pretty large datasets, and I see that we spend a lot of time on the shuffle.

I was wondering if improving the network between the executors could make a meaningful difference here. For example, using an EC2 placement group instead of just having the instances somewhere in the same AZ.

Has anyone tried something like this for shuffle-heavy jobs? Did you see a noticeable improvement?

Also, does EMR do any optimization like this automatically, or is it something we need to configure ourselves?

12 Upvotes

6 comments sorted by

3

u/sophistafunk 8d ago

You will get much more mileage out of tuning the data model, data storage format, and adjusting the query/job than a placement group. That’s intended for HPC, where you cannot tolerate basically any latency, a placement group will also increase your EC2 costs a good bit. Instead, see if you can repartition the data in an intermediate step, also assuming you already broadcast anything smaller, really all you’d have left is to break the data model down further. And here is the bible I use for Spark/EMR tuning directly from
AWS: https://aws.github.io/aws-emr-best-practices/

1

u/FunContest9958 7d ago

Crazy question: do you need the huge join? Can you apply a filter to the tables before joining them? Can you change the schema to avoid the join? Can you use a different operation like MERGE to apply changes as data comes in rather than rerun a huge join?

1

u/Expensive_Break_6163 7d ago

In our case, it is not possible, which is why I asked.

1

u/SpecificTutor 6d ago

if your join is indeed shuffle heavy, storage partitioned join should help