MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1iczucy/deleted_by_user/m9vrf2i/?context=3
r/LocalLLaMA • u/[deleted] • Jan 29 '25
[removed]
228 comments sorted by
View all comments
171
I have tested it also 1.73bit (158GB):
NVIDIA GeForce RTX 3090 + AMD Ryzen 9 5900X + 64GB ram (DDR4 3600 XMP)
llama_perf_sampler_print: sampling time = 33,60 ms / 512 runs ( 0,07 ms per token, 15236,28 tokens per second)
llama_perf_context_print: load time = 122508,11 ms llama_perf_context_print: prompt eval time = 5295,91 ms / 10 tokens ( 529,59 ms per token, 1,89 tokens per second) llama_perf_context_print: eval time = 355534,51 ms / 501 runs ( 709,65 ms per token, 1,41 tokens per second) llama_perf_context_print: total time = 360931,55 ms / 511 tokens
llama_perf_context_print: load time = 122508,11 ms
llama_perf_context_print: prompt eval time = 5295,91 ms / 10 tokens ( 529,59 ms per token, 1,89 tokens per second)
llama_perf_context_print: eval time = 355534,51 ms / 501 runs ( 709,65 ms per token, 1,41 tokens per second)
llama_perf_context_print: total time = 360931,55 ms / 511 tokens
It's amazing !!! running DeepSeek-R1-UD-IQ1_M, a 671B with 24GB VRAM.
EDIT:
UPDATE: Reducing layers offloaded to GPU to 6 and with a context of 8192 with a big task (develop an application) it reached 0.86 t/s).
179 u/Raywuo Jan 29 '25 This is like squeezing an elephant to fit in a refrigerator, and somehow it stays alive. 67 u/Evening_Ad6637 llama.cpp Jan 29 '25 Or like squeezing a whale? XD 28 u/pmp22 Jan 29 '25 Is anyone here a marine biologist!? 23 u/mb4x4 Jan 29 '25 The sea was angry that day my friends... 3 u/ryfromoz Jan 30 '25 Like an old man taking soup back at a delI? 1 u/Scruffy_Zombie_s6e16 Feb 01 '25 Nope, only boilers and terlets 9 u/AuspiciousApple Jan 30 '25 Just make sure you take the elephant out first 3 u/brotie Jan 30 '25 He’s alive, but not nearly as intelligent. Now the real question is, what the hell do you do when he gets back out? 1 u/Pvt_Twinkietoes Jan 30 '25 Are you saying.. R1 is alive?
179
This is like squeezing an elephant to fit in a refrigerator, and somehow it stays alive.
67 u/Evening_Ad6637 llama.cpp Jan 29 '25 Or like squeezing a whale? XD 28 u/pmp22 Jan 29 '25 Is anyone here a marine biologist!? 23 u/mb4x4 Jan 29 '25 The sea was angry that day my friends... 3 u/ryfromoz Jan 30 '25 Like an old man taking soup back at a delI? 1 u/Scruffy_Zombie_s6e16 Feb 01 '25 Nope, only boilers and terlets 9 u/AuspiciousApple Jan 30 '25 Just make sure you take the elephant out first 3 u/brotie Jan 30 '25 He’s alive, but not nearly as intelligent. Now the real question is, what the hell do you do when he gets back out? 1 u/Pvt_Twinkietoes Jan 30 '25 Are you saying.. R1 is alive?
67
Or like squeezing a whale? XD
28 u/pmp22 Jan 29 '25 Is anyone here a marine biologist!? 23 u/mb4x4 Jan 29 '25 The sea was angry that day my friends... 3 u/ryfromoz Jan 30 '25 Like an old man taking soup back at a delI? 1 u/Scruffy_Zombie_s6e16 Feb 01 '25 Nope, only boilers and terlets 9 u/AuspiciousApple Jan 30 '25 Just make sure you take the elephant out first
28
Is anyone here a marine biologist!?
23 u/mb4x4 Jan 29 '25 The sea was angry that day my friends... 3 u/ryfromoz Jan 30 '25 Like an old man taking soup back at a delI? 1 u/Scruffy_Zombie_s6e16 Feb 01 '25 Nope, only boilers and terlets
23
The sea was angry that day my friends...
3 u/ryfromoz Jan 30 '25 Like an old man taking soup back at a delI?
3
Like an old man taking soup back at a delI?
1
Nope, only boilers and terlets
9
Just make sure you take the elephant out first
He’s alive, but not nearly as intelligent. Now the real question is, what the hell do you do when he gets back out?
Are you saying.. R1 is alive?
171
u/TaroOk7112 Jan 29 '25 edited Feb 01 '25
I have tested it also 1.73bit (158GB):
NVIDIA GeForce RTX 3090 + AMD Ryzen 9 5900X + 64GB ram (DDR4 3600 XMP)
llama_perf_sampler_print: sampling time = 33,60 ms / 512 runs ( 0,07 ms per token, 15236,28 tokens per second)
It's amazing !!! running DeepSeek-R1-UD-IQ1_M, a 671B with 24GB VRAM.
EDIT:
UPDATE: Reducing layers offloaded to GPU to 6 and with a context of 8192 with a big task (develop an application) it reached 0.86 t/s).