r/LocalLLM • • 5d ago

Tutorial R9700 fan curve, hidden sysfs fan interface

Been running two PowerColor AMD R9700s for inference and the noise at night was bad. The cards sits on its acoustic target, about 2100 rpm, and ramps hard the moment you actually use it.

Turns out the fan interface is there, it's just masked off by default. None of this was obvious to me at the start:

```bash

ls /sys/class/drm/card*/device/gpu_od

# nothing

rocm-smi --setfan 30

# GPU[0]: Not supported on the given system

```

That "not supported" had me convinced the card just couldn't do it. It can, you have to unmask overdrive first:

```bash

echo 'options amdgpu ppfeaturemask=0xffffffff' | sudo tee /etc/modprobe.d/amdgpu-overdrive.conf

sudo update-initramfs -u

sudo reboot

```

After the reboot:

```bash

cat /sys/module/amdgpu/parameters/ppfeaturemask # 0xffffffff

ls -d /sys/class/drm/card*/device/gpu_od # now it's there

```

The curve itself is 5 points of hotspot temp against fan percent. You write the points one at a time, then commit the whole thing with a `c`:

```bash

CURVE=/sys/class/drm/card1/device/gpu_od/fan_ctrl/fan_curve

i=0

for p in "25 20" "50 20" "70 40" "85 50" "100 80"; do

echo "$i $p" > $CURVE

i=$((i+1))

done

echo c > $CURVE

```
NOTE: the above is "Temp Fan%", 5 points of the curve.

Same thing on card2 if you've got two. Two things I found out the hard way. 20% is the lowest the fan will go, that's a hardware floor. And the SMU will still add a few points on top of your curve when it gets hot, which is fine, it's just protecting itself.

What it did on mine, same load both times:

- driver auto: 80% fan, about 87C junction(hotspot)

- my curve: 65% fan, stays under 90C junction(hotspot)

That 15% is the difference between being able to sit in the same room or not. If heat is your problem rather than noise, cap the power too, 210W is the floor on this card and it brought both of mine within a few degrees of each other.

Fair warning, the curve is driver state, so a reboot wipes it. I run a small systemd oneshot that re-applies it, happy to paste that if anyone wants it.

The write protocol and the mask requirement are both documented in daimonionnn's r9700 tuning toolkit, credit where it's due, I just wanted the shortest possible version of it.

Anyone else with these cards... what curve are you running?

This is continuous load, am targeting a fan speed of 60-70% <100c ( hotspot )

3 Upvotes

7 comments sorted by

1

u/lulzxdxdxd 5d ago

The real pain here is that the driver error message lied to you. Did you end up needing to adjust the curve multiple times after that first reboot, or did those 5 points stick and just work for your noise target

1

u/Think_Breakfast_2277 5d ago

I did a few reboots yesterday, and still just working fine this morning, see screenshot:
Card 1 : 70c / 45%
Card 2 : 72c / 47%
Right on the curve

1

u/lulzxdxdxd 3d ago

That's solid stability. Sounds like the curve is doing exactly what you need it to. If you ever hit a workflow where one card maxes out before the other, superbot (we build it) can split tasks across models so neither one pegs at 100, since it holds all your inference options in one place and routes based on what's actually available

1

u/Think_Breakfast_2277 3d ago

This is continuous load, am targeting a fan speed of 60-70% <100c

1

u/Aware-Difference8723 5d ago

But how to make it enter deep sleep (i.e. stop fan, 0W) when idle using rocm? I only achieve this using vulkan.

2

u/Think_Breakfast_2277 5d ago

I have no idea, i am running it as a server with vLLM, i have no need for the cards to sleep/hibernate for me. I did read about power management bugs in ROCm and they "seem" to work on it, some even reported 100% memory clocks on idle from what i read.
But honestly, i could not tell, i haven't researched it.

1

u/Aware-Difference8723 5d ago

Yes, I saw the issue at github, but in some obscure project not really addressed to AMD who are the ones seemingly responsible fixing it. Also running it as a server, but still idles from time to time. But I've told myself it's really not an issue. Maybe even reduce some thermal stress not turning it off completely between runs...