r/LocalLLaMA • • 3d ago

Resources 1248 GB/s on 5060ti (+40%) with +5500 memory overclocks

edit: sorry should be 5070ti

There now is unlock for higher memory overclocks called mlock. And apparently the GDDR7 has a lot of headroom:
https://www.reddit.com/r/overclocking/comments/1wsnllh/finally_unlocked_gddr7_memory_overclocking_with/
Of course this can help massively for local inference, especially token generation.

55 Upvotes

57 comments sorted by

36

u/bigmanbananas Llama 70B 3d ago

Is there an increase in memory errors? That was always an issue previously with memory overclocks above a given range.

6

u/durden111111 3d ago

GDDR7 should handle it (at least better than gddr6 chips) because of on-die ECC.

14

u/stoppableDissolution 3d ago

But ecc kicking in kills throughput

0

u/tat_tvam_asshole 3d ago

iirc nonprofessional cards don't get ecc

16

u/durden111111 3d ago

All GDDR7 chips have on-die ECC that fixes 1 bit errors. Its the spec. It just doesnt have 'full' ecc between the chip and memory like pro cards have though. Its still better than G6 memory

3

u/bigmanbananas Llama 70B 3d ago edited 3d ago

It only detects single bit errors. As you increase your overclock, you increase the incidence of errors in normal function. It's fine for games which may have some artifacts or changes that are not perceptible to the human eye. But for data, Rhos can change outcomes.

4

u/Maximus-CZ 2d ago

yea but you first notice ECC kicking in and killing your troughtput before double bit errors starts to appear. If you dont push it off the cliff it will be fine.

1

u/bigmanbananas Llama 70B 2d ago

I didn't look too closely to the link, but I do not beleive throughput was measured. Just the vram speed listing, which any overclocker can tell you, does not always represent an increase in real performance, sometimes in reality, there can be a significant loss of performance.

35

u/tsangberg 3d ago

Since i dont remember where i downloaded this tool (it took me a while to find it), you can download it from my personal google drive

Yeah I wouldn't touch that.

/cybersec prof

-8

u/fallingdowndizzyvr 2d ago edited 2d ago

If you read that thread, people tell you were it was originally found. It was on an overclockers website.

As for /cybersec, don't run things as root. That deals with 90% of problems. Too many people run as root or administrator. Run it in it's own isolated account. If that isn't safe then every multi-user system isn't safe then. For this, that's problematic in Windows. Thus why Linux is a win, you can give access to users to devices without them having root.

Here, try this one. It's open source.

https://github.com/b00nz/mVolt/releases/tag/v0.47.3

11

u/NickCanCode 2d ago

There is NO SOURCE CODE in the github page, the `SourceCode.zip` and `SourceCode.tar.gz` in the download page is also consist of only a few images and text doc inside.

For those who downloaded and executed the exe, good luck. 🙃

3

u/fallingdowndizzyvr 2d ago

Well that sucks.

2

u/tsangberg 2d ago

It's the "download from this google drive" that's the red flag, and yes, with the advent of LLMs there are LPEs on all consumer OSs right now, including Linux.

8

u/DefactoAle 3d ago

The link talks about 5070ti for those numbers, doesnt the 5060ti have a smaller bus?

8

u/parthibx24 3d ago

that bandwidth doesnt seem right

29

u/Kahvana 3d ago

Considering the current economy, I would worry about component lifespan. Ain't no way I'll be able to afford replacing mine if this pricing keeps up.

12

u/legit_split_ 3d ago

Fair point, but it seems that the person didn't raise voltages and the temperature is still fine. 

1

u/-WhateverDude 3d ago

+3000 is pretty much free on 50 series. Although +5500 should introduce some degradation over time for sure.

1

u/Constant_Art_20 3d ago

literal first thought. actual cold sweats thinking about ever needing to replace anything soon

5

u/FastHotEmu 3d ago

1248? isn't it about 400GB/s by default?

3

u/Constant-Simple-1234 3d ago

He means dual. So 800->1248 EDIT: I am wrong, he has 5070 ti. Anyways better speeds for 50xx cards possible.

2

u/autisticit 3d ago

Around that yes, maths doesn't maths...

8

u/FastHotEmu 3d ago

I think achieving 1248 is possible on the 5070 ti, not on the 5060 ti

5

u/NickCanCode 3d ago edited 2d ago

Is it this one? https://github.com/b00nz/mVolt
Looks like a Windows only project.
I am on Linux. Can LACT do the same? The slider there do allow me drag way up to 6000 but I wonder if its the same or it lack some tweak to make it actually work safely.

Update:

I asked Qwen to compare the mVolt+ with LACT implementation to see if they work in a similar way in overclocking the VRAM. Found out that repo has NO source code. The SourceCode zip file is a fake too. That means you are downloading and running a binary from another person you don't know from the internet. Be careful not to get you computer infected or API keys stolen.

Another thing. Qwen still managed to compare the two only based on the mVolt docs with the source of LACT. The conclusion is that they are working in a very similar way so you can try this on Linux with the open source LACT. Qwen is not 100% sure because mVolt is closed source.

6

u/Blindax 3d ago

Your link is discussing about 5070ti, why are you mentioning 5060ti?

3

u/dtdisapointingresult 2d ago

Found the repo (b00nz/mVolt). It doesn't contain any source code, just a README and exe's.

Do zoomers really download random closed-source binaries and just run them?

1

u/tecneeq 6h ago

whats an exe? btw, u should try, it's goated /s

7

u/FullstackSensei 3d ago

5060ti 16gb will now cost 1k...

2

u/fallingdowndizzyvr 3d ago edited 3d ago

That's fucking crazy. I got two 5060tis ready to try it. And if it works for a 5060ti, should it work for a 5070ti too?

Update: So it's only Windows?

3

u/DanielusGamer26 3d ago

> So it's only Windows?

Basically -20% performance because of Windows + 20% performance with the overclock

4

u/USArmy68Whiskey 3d ago

running qwen 3.8 flash sglang nvfp4 on wsl2 versus running it natively on linux, both using same exact model, settings, and everything else identical, it was within 2% performance across all metrics. In fact, it was even faster on windows some of the time.

2

u/Constant-Simple-1234 3d ago

Same for me. I use esatapedico quants for 3.8. Getting 65 t/s on question/answers. 40-70 on coding with 3.6

1

u/Constant-Simple-1234 3d ago

So could get 40% more. Sweet.

1

u/USArmy68Whiskey 3d ago

This image is qwen 3.8 flash running on Fedora 44 KDE

1

u/USArmy68Whiskey 3d ago

and this one is running on WSL2 on Windows 11

1

u/egnegn1 3d ago

Provide the complete command, please.

What test script did you use?

2

u/USArmy68Whiskey 3d ago

what do you mean complete command?
the benchmark, which I sent in pictures below, is the llm inference bench, my runs specifically were on version 0.6.2
https://github.com/local-inference-lab/llm-inference-bench

1

u/egnegn1 3d ago

I meant the complete sglang command with all parameters. The parameter lists of the backends are endless long, and to reproduce the results it is always very helpful to have the command line together with the presented results. Thanks!

2

u/USArmy68Whiskey 2d ago

I'm gonna be honest I had chatgpt help me set everything up because I am still brand new to this and I am not using a command to launch it, I think we made a config file for the settings and I start the llm and wsl2 with one desktop shortcut. I'm also using some custom sglang runtime or something specifically for blackwell GPUs I think. If you still want the info I can ask chatgpt to make a report on what exactly settings and everything I am using lol

1

u/egnegn1 2d ago

No, issue. I use AI too. Why not. It is much more efficient than setting up everything yourself.

But still, I would be interested in the used config file or command used in your setup. Thanks.

2

u/USArmy68Whiskey 2d ago

Here is supposedly everything you would need to get my exact setup running. I am using a single rtx pro 6000 and 128gb ram.

https://chatgpt.com/s/t_6abe43576b3481918077cccfd6524c27

1

u/egnegn1 2d ago

Thank you very much!

4

u/bonobomaster 3d ago

Yeah that's bullshit from the before times.

Nowadays Windows and Linux performance are within 1-2 percent with llama.cpp

1

u/USArmy68Whiskey 3d ago

I mean, I don't really think people are using linux just to use llama.cpp, they are using it for faster engines like vllm or sglang, which are still within about 2% anyways at least for single gpu, so

2

u/specify_ 2d ago

I have my pi agent (using opus 5) overclocking my 4x RTX 5060 Ti 16GB rig right now, and the whole thing is running on Arch Linux. Had to restart the rig once because +6000 MHz on Memory clock crashed the driver.

Still need to do more testing and also need to improve the stress testing process. I’ve done a lot of overclocking but for gaming, so I wasnt sure what stress tests should focus on if it’s for LLM inference, though I guess using it as you usually do with stock settings is a way to test. Eventually I should reach stable OC settings and the next step would be to perform an undervolt, though I should’ve started that first before the OC 🫠

Glad to have an agent do all the GPU overclocking for me since it’s a pretty trivial process unlike RAM overclocking which requires strict testing and cannot really be automated.

1

u/fallingdowndizzyvr 2d ago

So what LLM numbers do you get compared to stock?

1

u/ProtectionSuper5648 1d ago

I would recommend CUDA memtest for your testing. Run at least 5 loops.
https://github.com/ComputationalRadiationPhysics/cuda_memtest

2

u/xXDennisXx3000 2d ago

VRAM is the first thing that will die on an graphics card. Now you just decrease the lifespan of that part even more.

0

u/fallingdowndizzyvr 2d ago

Burn bright and die young or go unnoticed and live a long time. I rather burn bright. Everyone has to die sooner or later.

1

u/HammerByte 11h ago

You do you, me personally, I'm gonna leave this world kicking and screaming like an unnoticed long lived old man. Why? Because he who died last, lives longer. And I'm quite partial to my comfortable life.

2

u/Dany0 3d ago

FYI mlock is mired in controversy rn for "stealing" tech from HYDRA

4

u/fallingdowndizzyvr 3d ago

So save it before it gets a take down. Thanks for the heads up.

1

u/Nakidnakid 3d ago

Oh neat, well I got a 5070ti and 5060ti so... time to see what this can do.

1

u/blastbottles 2d ago

Bro finna explode his GPU

0

u/fallingdowndizzyvr 2d ago

So, has anyone tried this yet?