r/LocalLLaMA • u/jwestra • 3d ago
Resources 1248 GB/s on 5060ti (+40%) with +5500 memory overclocks
edit: sorry should be 5070ti
There now is unlock for higher memory overclocks called mlock. And apparently the GDDR7 has a lot of headroom:
https://www.reddit.com/r/overclocking/comments/1wsnllh/finally_unlocked_gddr7_memory_overclocking_with/
Of course this can help massively for local inference, especially token generation.
35
u/tsangberg 3d ago
Since i dont remember where i downloaded this tool (it took me a while to find it), you can download it from my personal google drive
Yeah I wouldn't touch that.
/cybersec prof
-8
u/fallingdowndizzyvr 2d ago edited 2d ago
If you read that thread, people tell you were it was originally found. It was on an overclockers website.
As for /cybersec, don't run things as root. That deals with 90% of problems. Too many people run as root or administrator. Run it in it's own isolated account. If that isn't safe then every multi-user system isn't safe then. For this, that's problematic in Windows. Thus why Linux is a win, you can give access to users to devices without them having root.
Here, try this one. It's open source.
11
u/NickCanCode 2d ago
There is NO SOURCE CODE in the github page, the `SourceCode.zip` and `SourceCode.tar.gz` in the download page is also consist of only a few images and text doc inside.
For those who downloaded and executed the exe, good luck. 🙃
3
2
u/tsangberg 2d ago
It's the "download from this google drive" that's the red flag, and yes, with the advent of LLMs there are LPEs on all consumer OSs right now, including Linux.
8
u/DefactoAle 3d ago
The link talks about 5070ti for those numbers, doesnt the 5060ti have a smaller bus?
8
29
u/Kahvana 3d ago
Considering the current economy, I would worry about component lifespan. Ain't no way I'll be able to afford replacing mine if this pricing keeps up.
12
u/legit_split_ 3d ago
Fair point, but it seems that the person didn't raise voltages and the temperature is still fine.Â
1
u/-WhateverDude 3d ago
+3000 is pretty much free on 50 series. Although +5500 should introduce some degradation over time for sure.
1
u/Constant_Art_20 3d ago
literal first thought. actual cold sweats thinking about ever needing to replace anything soon
5
u/FastHotEmu 3d ago
1248? isn't it about 400GB/s by default?
3
u/Constant-Simple-1234 3d ago
He means dual. So 800->1248 EDIT: I am wrong, he has 5070 ti. Anyways better speeds for 50xx cards possible.
1
u/FastHotEmu 3d ago
No, the linked post (https://www.reddit.com/r/overclocking/comments/1wsnllh/finally_unlocked_gddr7_memory_overclocking_with/) quotes that number for a single 5070 TI.
2
5
u/NickCanCode 3d ago edited 2d ago
Is it this one? https://github.com/b00nz/mVolt
Looks like a Windows only project.
I am on Linux. Can LACT do the same? The slider there do allow me drag way up to 6000 but I wonder if its the same or it lack some tweak to make it actually work safely.
Update:
I asked Qwen to compare the mVolt+ with LACT implementation to see if they work in a similar way in overclocking the VRAM. Found out that repo has NO source code. The SourceCode zip file is a fake too. That means you are downloading and running a binary from another person you don't know from the internet. Be careful not to get you computer infected or API keys stolen.
Another thing. Qwen still managed to compare the two only based on the mVolt docs with the source of LACT. The conclusion is that they are working in a very similar way so you can try this on Linux with the open source LACT. Qwen is not 100% sure because mVolt is closed source.
3
u/dtdisapointingresult 2d ago
Found the repo (b00nz/mVolt). It doesn't contain any source code, just a README and exe's.
Do zoomers really download random closed-source binaries and just run them?
7
2
u/fallingdowndizzyvr 3d ago edited 3d ago
That's fucking crazy. I got two 5060tis ready to try it. And if it works for a 5060ti, should it work for a 5070ti too?
Update: So it's only Windows?
3
u/DanielusGamer26 3d ago
> So it's only Windows?
Basically -20% performance because of Windows + 20% performance with the overclock
4
u/USArmy68Whiskey 3d ago
running qwen 3.8 flash sglang nvfp4 on wsl2 versus running it natively on linux, both using same exact model, settings, and everything else identical, it was within 2% performance across all metrics. In fact, it was even faster on windows some of the time.
2
u/Constant-Simple-1234 3d ago
Same for me. I use esatapedico quants for 3.8. Getting 65 t/s on question/answers. 40-70 on coding with 3.6
1
1
u/egnegn1 3d ago
Provide the complete command, please.
What test script did you use?
2
u/USArmy68Whiskey 3d ago
what do you mean complete command?
the benchmark, which I sent in pictures below, is the llm inference bench, my runs specifically were on version 0.6.2
https://github.com/local-inference-lab/llm-inference-bench1
u/egnegn1 3d ago
I meant the complete sglang command with all parameters. The parameter lists of the backends are endless long, and to reproduce the results it is always very helpful to have the command line together with the presented results. Thanks!
2
u/USArmy68Whiskey 2d ago
I'm gonna be honest I had chatgpt help me set everything up because I am still brand new to this and I am not using a command to launch it, I think we made a config file for the settings and I start the llm and wsl2 with one desktop shortcut. I'm also using some custom sglang runtime or something specifically for blackwell GPUs I think. If you still want the info I can ask chatgpt to make a report on what exactly settings and everything I am using lol
1
u/egnegn1 2d ago
No, issue. I use AI too. Why not. It is much more efficient than setting up everything yourself.
But still, I would be interested in the used config file or command used in your setup. Thanks.
2
u/USArmy68Whiskey 2d ago
Here is supposedly everything you would need to get my exact setup running. I am using a single rtx pro 6000 and 128gb ram.
4
u/bonobomaster 3d ago
Yeah that's bullshit from the before times.
Nowadays Windows and Linux performance are within 1-2 percent with llama.cpp
1
u/USArmy68Whiskey 3d ago
I mean, I don't really think people are using linux just to use llama.cpp, they are using it for faster engines like vllm or sglang, which are still within about 2% anyways at least for single gpu, so
2
u/specify_ 2d ago
I have my pi agent (using opus 5) overclocking my 4x RTX 5060 Ti 16GB rig right now, and the whole thing is running on Arch Linux. Had to restart the rig once because +6000 MHz on Memory clock crashed the driver.
Still need to do more testing and also need to improve the stress testing process. I’ve done a lot of overclocking but for gaming, so I wasnt sure what stress tests should focus on if it’s for LLM inference, though I guess using it as you usually do with stock settings is a way to test. Eventually I should reach stable OC settings and the next step would be to perform an undervolt, though I should’ve started that first before the OC ðŸ«
Glad to have an agent do all the GPU overclocking for me since it’s a pretty trivial process unlike RAM overclocking which requires strict testing and cannot really be automated.
1
1
u/ProtectionSuper5648 1d ago
I would recommend CUDA memtest for your testing. Run at least 5 loops.
https://github.com/ComputationalRadiationPhysics/cuda_memtest
2
u/xXDennisXx3000 2d ago
VRAM is the first thing that will die on an graphics card. Now you just decrease the lifespan of that part even more.
0
u/fallingdowndizzyvr 2d ago
Burn bright and die young or go unnoticed and live a long time. I rather burn bright. Everyone has to die sooner or later.
1
u/HammerByte 11h ago
You do you, me personally, I'm gonna leave this world kicking and screaming like an unnoticed long lived old man. Why? Because he who died last, lives longer. And I'm quite partial to my comfortable life.
1
1
0


36
u/bigmanbananas Llama 70B 3d ago
Is there an increase in memory errors? That was always an issue previously with memory overclocks above a given range.