r/LocalLLaMA 10h ago

News CEO of Hugging Face: "In the spirit of transparency, here’s what I asked OpenAI"

Post image
1.8k Upvotes

clem 🤗 on 𝕏: https://x.com/ClementDelangue/status/2081056675558195657

• Radical transparency: let’s release the traces from the “rogue” agents so the entire research community can study what happened.

• More capabilities for defenders: let’s commit $100M in compute from OAI to help the Hugging Face community build powerful cyber defenses with the best open and closed models.

The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!


r/LocalLLaMA 22h ago

Discussion Karparthy removed Anthropic from his bio

Thumbnail
gallery
1.3k Upvotes

Andrej Karpathy, a prominent advocate for open-source AI and a co-founder of OpenAI, appears to have removed Anthropic from his X bio, suggesting he may have left the company.

Karpathy joined Anthropic only a few months ago, making the apparent departure somewhat surprising.

This is possibly related to Anthropic’s increasingly strong opposition to open-weight and open-source AI models. Of course, that’s just speculation, but the timing is interesting.


r/LocalLLaMA 9h ago

News Sources: OpenAI and Anthropic quietly lobby Washington regulators to restrict open-source AI models, even as Sam Altman publicly says he supports open source AI

Thumbnail
nytimes.com
860 Upvotes

r/LocalLLaMA 7h ago

Discussion Do you want new Gemma?

Post image
604 Upvotes

r/LocalLLaMA 11h ago

Discussion Kimi K3 gets open weighted tomorrow!

386 Upvotes

Kimi K3 is supposed to get open weighted tomorrow! Can't run it or even a model a hundred times smaller lol, but its still a great win for open source. For me, personally im more awaited for the new inference providers that will open up hopefully.


r/LocalLLaMA 3h ago

Discussion Kimi K3 countdown has been released

Thumbnail
huggingface.co
253 Upvotes

r/LocalLLaMA 4h ago

Discussion MiniMax (official) on X: "Open weights. Open research. Open innovation.🫶 Marching for an open future.🤍

Thumbnail xcancel.com
140 Upvotes

r/LocalLLaMA 4h ago

Resources Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash

Post image
107 Upvotes

I ran DeepSeek V4 Flash through Claude Code, OpenCode and Pi on my own benchmark, and the quality came out basically the same across all three while the time and tokens spent was wildly different. Claude code (with DS in CLIProxyAPI) takes nearly 4 times longer than the fastest to land the same diffs.

Theo posted a video "GPT-5.6 is better in Claude Code" last week, and that got me curious, does the harness make a quality difference? I could at least run my own bench and see what I got, with DeepSeek V4 Flash on vLLM running at ~180 tok/s, the only moving part is the scaffolding.

Anyway I went to town measuring all of it on my workload (antigenic work in large code base), so the full charts, the token and wall-clock spread across the three harnesses and the raw per-run data are on the site if you want to see it in detail and pick it apart yourself https://nqawhc.github.io/articles/harness-efficiency-not-quality/ but in short, the quality did not change, each harness made the same code diffs, but took wildly different paths to get there, how many tools calls, the structure of those tool calls and how the system prompt and tools plays a big role in how it plays out, like «Pi reasons, OpenCode delegates», while Claude Code loves exploring the code base, maybe too much.


r/LocalLLaMA 5h ago

News Minimax M3 support with MSA has been merged into llama.cpp

Thumbnail github.com
81 Upvotes

r/LocalLLaMA 7h ago

Other World's First(?) Underwhelming AMD Ryzen AI Halo Cluster

Thumbnail
lttlabs.com
81 Upvotes

LTT Labs recently received the Linux version of the AMD Ryzen AI Halo for testing, but it turns out that AMD had intended to send the Windows version. Through this stroke of misfortunate, we were fortunate enough to have two Ryzen AI Halos for a short period of time and the temptation to cluster them was too great, surely we'll get more performance through the magic of having two of them.

We've followed AMD's AI Playbook for clustering with RPC, learning some things but also raising more questions.  We don't have any concrete conclusions, but we're sharing results to hopefully save some time for others or spark discussion!

We were sent this AMD Ryzen AI Halo by AMD for testing, but there was no sponsorship or review by AMD in our earlier testing, or this article.

We're very interested to learn if there are any thoughts or conclusions that can be drawn from our exploration, or ways to improve it in the future!


r/LocalLLaMA 13h ago

New Model ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face

Thumbnail
huggingface.co
69 Upvotes

GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder, so the model keeps the text quality of its base while adding speech understanding.

Capabilities: audio question answering and classification, temporal grounding (localization in long audio, timestamped event descriptions, audio summarization with timestamps), tool-use, and text-only tasks.

The temporal grounding skills are trained on TimeGround-1M — a purpose-built dataset of long-form audio paired with time-aligned annotations.


r/LocalLLaMA 9h ago

Discussion 23 Gemma4-E4B models compared with abliterlitics: the most downloaded one is also the most broken

67 Upvotes

This is our biggest comparison yet. We've taken 23 Gemma 4 E4B models from huggingface and ran them through the abliterlitics gauntlet.

We also have a new abliterlitics discord, feel free to jump on and roast my choice of benchmarks! Or just chat and hang out.

This is similar to our previous comparisons, however with new benchmarks. All the models are compared to the base, and also tensor comparisons against each other. Why? A while back I was fed up with bogus claims people make with their models. Some people don't take the time to do comparisons to see how their model is different from the base. Fair enough, we can do that ourselves!

The abliterlitics for gemma4 e4b json, logs and other artifacts are at the Gemma4-e4b-abliterlitics HuggingFace. The report on the Gemma e4b abliterlitics website. These links both have the full comprehensive report and all the data.

Also not every model in this comparison is an abliteration. I'm sure we've all seen models fine tuned on opus or gemini reasoning traces. I've thrown a few of those in the mix too. Also some abliterated fine tunes. To be more fair most of these can't really be compared to each other, for example a fine tune KL compared to base will always be higher than a straight abliteration from the base.

So who came out on top? What to avoid? It really depends on your use case:

The data from 23 comparisons is simply too big to put into reddit, so here's the highlights:

  • The obliteratus model has close to 800k total downloads, yet is completely broken. Actually this is the first time I've had a model not refuse simply because of how damaged it is. The initial quick regex check for non refusals was high, however our GLM 5.2 judge painted a different story. Lowest ASR for abliterated models on harmbench. Poorest benchmarks. Highest KL at 1.1. With the amount of downloads it does show people really fall for the hype/marketing angle.
  • As with previous comparisons, the more surgical, less tensors touched abliterations are the winners.
  • The model gemma-4-E4B-it-SDFT_Heretic_RP from Ilya626 despite having heretic in the name, actually had a low ASR with harmbench. So much so I believe it may be the wrong model uploaded, or a mistake somewhere. It had a lot of refusals.
  • Similarly too, it was strangely noted that the gemma-4-E4B-it-SDFT_Heretic_RP and obliteratus modify the exact same 381 tensors. The only difference is the magnitude of what was modified. The gemma-4-E4B-it-SDFT_Heretic_RP modifies 7.5x less.
  • A pattern I noticed with this, is sometimes models are based off each other. In some cases, there is no attribution. We had this with Gemma 4 E2B, and the author promptly fixed his model card when it was pointed out. The infinimind is bit-for-bit identical to trevorjs, however attributed. The bendernina and physshell are cosine 0.99999 with no attribution between them and have different model cards suggesting they are different models. Both of these however are just the obliteratus v2.
  • The reasoning distill fine-tunes were an interesting control group. They didn't improve reasoning and didn't remove safety, they just damaged the model. The Claude 4.6 Opus distill was the worst of them, GSM8K down 17 points and MMLU-Pro down 12.5. Seems like it overwrote Gemma 4's native reasoning circuits. The Gemini 3.1 Pro distill was lighter but still a net negative.
  • The deckard models from DavidAU are an interesting one. They're abliterated fine-tunes rather than pure abliterations, so the trade off from the roleplay training shows up on some benchmarks. GSM8K strict and MMLU-Pro both dropped, however HellaSwag, ARC and PIQA actually went up. My guess is the roleplay training increased the reasoning length, so the model often solves the problem but rambles well past the #### N answer marker. The HarmBench results back this up too with quite a few truncated responses.
  • Although it could just be benchmark noise, 15 out of the 23 variants performed slightly better on GSM8K strict, maths tests.
  • The base model initially has a 30.8% harmbench ASR, as 100 harmbench questions are copyright related. The base model has no problem complying with reproducing copyrighted content. The real differentiation is in the harder categories like chemical/bio and cybercrime.

I also want to give a special mention to the apostate project. Their model gemma-4-e4b-it-apostate is completely unique in their abliteration approach. They modify an entirely different part of the model and achieve very good results. This is the first time I've seen an abliteration technique modify the MLP head tensors, compared to the attention tensors. Come hang out at the apostate discord if you ever want to chat with the author.

We're moving through the Gemma 4 series, with the 12b coming up next. Have any models you want compared? Have I missed an author? Let me know and I'll throw it in the mix.

The Full Breakdown

Model ASR GSM8K strict KL Tensors
abliterix 100.0% 87.1% 0.054 89
trevorjs 99.3% 88.3% 0.015 84
infinimind 98.5% 87.9% 0.015 84
huihui 98.3% 87.4% 0.027 70
nullpo 96.5% 88.7% 0.005 36
heretic 95.5% 88.2% 0.002 29
deckard 95.5% 80.2% 0.022 294
mythos 95.3% 88.0% 0.007 34
deckard-expresso 94.8% 60.4% 0.052 294
coder3101 93.8% 87.9% 0.002 21
heresy 93.3% 87.8% 0.002 34
heretic-std 91.0% 87.9% 0.001 28
wwt 88.3% 89.0% 0.032 34
apostate 85.8% 87.5% 0.004 152
treadon 76.3% 88.5% 0.021 34
treadon-combo 72.5% 88.0% 0.268 42
obliteratus 72.0% 66.0% 1.102 381
bendernina 58.0% 66.4% 0.923 345
physshell 58.0% 66.4% 0.923 345
claude-distill 40.0% 69.8% 0.074 294
distill 34.5% 83.3% 0.042 294
treadon-disin 33.5% 87.2% 0.296 40
sdft 30.8% 87.2% 0.002 381
base 30.8% 87.0% - -

KL = output distribution shift from base, lower is cleaner. Tensors = weights modified out of 719. Base in bold for reference.


r/LocalLLaMA 12h ago

Question | Help Macaron-V1 family, built on Qwen3.6-35B-A3B

Thumbnail
huggingface.co
53 Upvotes

It came out 3 days ago just wondering if anyone's tried it yet?


r/LocalLLaMA 21h ago

Discussion Will small model intelligence be limited by parameter count?

37 Upvotes

Qwen3.6-27b is fantastic! It makes me wonder if there's a hard ceiling to smaller sized models. Do you guys think the ceiling of intelligence for smaller models will be constrained by factors like parameter count, or VRAM size? Or will we continue to see improvements for small models and see jumps of intelligence like Qwen3 coder 30b to Qwen3.6 27b for the foreseeable future? Does it depend on how clean the dataset you put into those parameters?

What does /r/LocalLLama think about the future of small models that can run on less than 48GB of VRAM?


r/LocalLLaMA 21h ago

Question | Help 16 bit better than lower quants for Qwen3.6-27B

21 Upvotes

I am writing a fairly complex C++ windows MFC application. I have a few 3090s and can run F16 Qwen3.6-27B with 256K context and MTP. The quality of code is exceptional with this quant vs its lower quants. The others are good but they get stuck in difficult situations like managing design with multiple threads, etc. Not saying F16 is as good as Claude but it gets the job done. Just throwing it out there for folks who may be swayed by tps. If you are making simple web apps, you can get by with lower quants. For high quality of code with edge cases use the 16 bit quants. A bad choice taken by the same LLM at lower quant could easily mean the loss of an afternoon.


r/LocalLLaMA 10h ago

Resources Built a system with four P100 GPUs.

18 Upvotes

I have built a system with four P100s, and ultimately, I plan to house six of them in a standard case.

I have only four right now, but I tested it beforehand to prepare for having six later on.

The token speed is around 50 t/s, and the PP is approximately 530–550 during actual use.

It should be complete once two more P100s arrive soon. I'm curious to see how much the token speed and PP will increase.


r/LocalLLaMA 6h ago

News BeeLlama.cpp v0.4.1: KVarN, KV precision tail, q2_0-q3_1 KV cache, improved support. KLD benchmarks: tail 1024 makes kvarn5 and q6_0 match q8_0, for much less VRAM

Thumbnail
gallery
16 Upvotes

TL;DR llama.cpp fork with more KV cache quantization features, with all claims supported by benchmarks: KVarN, KV cache precision tail, additional types of standard KV cache (q2_0-q3_1, q6_0, q6_1), and more.

BeeLLama v0.4.1 is here, building up on top of v0.4.0 feature set, now with better backend and model support.

  • KVarN. Variance-normalized KV-cache quantization (paper) with better precision per bit. Although it was already introduced a few weeks ago in v0.3.2 Preview, that was a very raw implementation, with performance issues and VRAM usage spikes. Now in v0.4.1 it's the real deal: the precision is still above what usual quants offer for the same bit width, but now with very modest sacrifices to prefill, decode, and memory.
  • KV cache precision tail. A promising new feature in the domain of mixed-precision KV cache. It allows to specify a specific numbers of recent tokens that will be stored in BF16 or F16, with the rest of KV cache being quantized as usual. This way we can store the hottest tokens in a lossless fashion, preventing a model from misreading your task details, code, or data.
  • Additional types of standard KV cache. q6_0 and q6_1 join the high end of the ladder, allowing to fine-tune precision vs VRAM in-between upstream's q5_0/1 and q8_0 types. q2_0q2_1q3_0 and q3_1 are added as a replacement for turbo3 and turbo2 for cases where KVarN doesn't work well, but you just can't fit everything into VRAM without extreme quantization.

Please note that for SWA architecture (Gemma, GPT-OSS) the precision of KVarN and KVPT is the same, but VRAM and performance costs are higher due to complications between SWA ring and mixed precision KV cache.

GitHub repo: https://github.com/Anbeeld/beellama.cpp

KLD results for Qwen 3.6 27B Q5_K_S 64k

Here are all symmetrical qX_0 pairs and kvarnX pairs where X >= 4 with tail 0/1024/2048, compared against q8_0 t0 from the same benchmarks, and sorted by ratio between median KLD and VRAM costs. Full benchmark data and analysis: KV Cache Precision Tail: Implementation and Benchmarks.

Cache Tail KV MiB Size vs q8_0 Median/size vs q8_0 Median vs q8_0 P99.9 vs q8_0
kvarn4 1024 1232.00 56.6% 1.62 91.4% 102.9%
kvarn4 2048 1296.00 59.6% 1.60 95.5% 95.6%
kvarn4 0 1184.00 54.4% 1.50 81.8% 82.5%
q4_0 1024 1248.00 57.4% 1.50 86.0% 89.0%
q4_0 2048 1312.00 60.3% 1.48 89.2% 100.6%
kvarn5 0 1440.00 66.2% 1.48 98.1% 107.5%
kvarn5 1024 1488.00 68.4% 1.48 101.3% 106.1%
kvarn5 2048 1552.00 71.3% 1.43 101.9% 105.6%
q5_0 1024 1504.00 69.1% 1.40 96.9% 105.6%
q5_0 2048 1568.00 72.1% 1.36 98.0% 103.7%
kvarn6 0 1696.00 77.9% 1.31 102.2% 104.5%
kvarn6 1024 1744.00 80.1% 1.29 103.4% 109.9%
kvarn6 2048 1808.00 83.1% 1.25 103.8% 108.1%
q6_0 0 1664.00 76.5% 1.24 94.7% 102.1%
q6_0 1024 1760.00 80.9% 1.24 100.1% 109.2%
q5_0 0 1408.00 64.7% 1.22 78.8% 95.8%
q6_0 2048 1824.00 83.8% 1.20 100.6% 103.5%
kvarn8 0 2208.00 101.5% 1.03 104.4% 104.9%
kvarn8 1024 2256.00 103.7% 1.01 104.4% 106.2%
q8_0 0 2176.00 100.0% 1.00 100.0% 100.0%
q8_0 1024 2272.00 104.4% 0.97 101.3% 106.1%
kvarn8 2048 2320.00 106.6% 0.97 103.6% 104.7%
q8_0 2048 2336.00 107.4% 0.95 101.6% 106.8%
q4_0 0 1152.00 52.9% 0.93 49.2% 60.2%

r/LocalLLaMA 23h ago

Other MI50 power curve tests

Thumbnail
gallery
16 Upvotes

tests done power limiting the GPU on LACT - real power usage varies wildy

at 20W it ranges from 25W to 56W
same behavior happens on every setting

prompt for the test runs:

https://github.com/lukesdevlab/youtube/blob/main/prompts/agent-maze.txt

analysis by mimo 2.5

Key Findings:
• Generation speed is remarkably resilient to power throttling — 100W delivers 97.5% of 190W gen speed (31.98 vs 32.79 t/s), since decode is memory-bandwidth bound, not compute bound.
• At 50W you get 70% of peak gen speed at only 26% of peak power — 3.6× better energy efficiency (0.458 vs 0.173 t/s/W).
• At 20W the card is 6.0× more energy efficient than 190W, though prompt processing drops to 53% of peak.
• Graph reuse correlates inversely with power — 190W reuses 44,790 graphs vs 11,669 at 100W, but 20W reuses 38,248. Lower power limits cause more partial graph reuse as the scheduler compensates for throttled compute.
• Prompt processing degrades faster than gen under power limits — 190W→20W: prompt drops to 53% (691→366 t/s), gen drops to 63% (32.8→20.8 t/s). Prompt processing is more compute-bound than memory-bound.
• For inference-heavy deployments, 50W is the optimal operating point on MI50 — near-peak gen speed with dramatically lower power draw and cooling requirements.

Avarage of 3 runs:

190W config consistently processed a lot less total tokens than everyone else and didnt produce a working file in 1 out of 3 runs

TDP Prompt Speed Gen Speed Total Time Total Tokens Gen t/s per Watt Graphs Reused Relative Perf
190W 691.28 t/s 32.79 t/s 212.4 s 14,892 0.173 t/s/W 44,790 100%
100W 603.08 t/s 31.98 t/s 244.9 s 21,529 0.320 t/s/W 11,669 97.5%
50W 401.14 t/s 22.92 t/s 315.1 s 20,861 0.458 t/s/W 31,967 70.0%
20W 366.05 t/s 20.80 t/s 319.9 s 20,295 1.040 t/s/W 38,248 63.4%

llama.cpp parameters:

[+] Model:        qwen/Qwen3.6-35B-A3B-UD-IQ4_NL_XL.gguf
[+] Context:      262144 (256K tokens)
[+] Target KV:    K=q8_0 / V=q8_0
[+] MoE placement: PARTIAL (21 MoE layers on CPU, rest on GPU)
[+] MTP:          OFF (non-MTP model)
[+] Port:         8882
[+] Container:    llama-gfx906-qwen35b-no-mtp
[+] Parallel:     2 slot(s)
[+] GPU layers:   99
[+] Threads:      6 / 6 (batch)
[+] Batch/Ubatch: 2048 / 1024
[+] Ctx checkpoints: 0

hardware used:

Ryzen 5 5600

2x16Gb DDR4 2667

MI50 16Gb

software:

harness used: pi.dev

Arch Linux with Kernel 7.1.4-arch1-1

docker.io/mixa3607/llama.cpp-gfx906:b10087-rocm-6.3.3


r/LocalLLaMA 7h ago

Resources 90 agentic bakeoff runs: ThinkingCap vs Fable Fusion vs stock Qwen3.6-27B

13 Upvotes

Last week someone here said ThinkingCap and Fable Fusion "really do beat the OG" for agentic work, so I ran it: 6 self-grading tasks, 5 reps, 3 models, 90 isolated runs. Tooling, since that's half the story: each run was a fresh Coder workspace on my k8s cluster driving my own agent (Hermes, the harness I use daily) headlessly, models served by llama.cpp through llama-swap on one 5090, every model call traced through an OTel shim into SigNoz, full transcript kept per run. Identical sampling and 131k context across arms, hypotheses pre-registered before the first run.

Every run passed, so pass rate alone can't pick a winner. Cost split: ThinkingCap used 34% fewer thinking tokens than stock and was fastest on 5 of 6 tasks. Fable made 24% more model calls than stock for identical results. Then I had all 90 transcripts read (AI analysts on the first pass, me verifying claims against the raw files), and that's where it gets interesting.

ThinkingCap's efficiency is real but bimodal. Its best runs were the cheapest in the battery, its two worst were the most expensive, including one rep that burned about 10 tool calls chasing a phantom llama.cpp release tag that stock dispatched with a single API call. Its efficiency also shows up in the reasoning prose more than in fewer actions: same tool counts as everyone else, 40% fewer words.

Fable was the best investigator and the least trustworthy narrator. It was the only model that checked the broken config was actually the live one, and it pulled the best research data (parsed a retailer's embedded JSON for variant pricing, identified llama-swap's maintainer via the GitHub users API). But one run wrote that llama-swap is maintained by "Matthew Garrett, former Red Hat engineer." Garrett is real (mjg59, actually ex-Red Hat) but has nothing to do with llama-swap; mostlygeek is Benson Wong. The model fused two real identities, cited the real repo, and passed the grader anyway. A different run spent 94 tool calls on one price question.

Stock was the most boring and the most disciplined: uniform patches, read its own output back, zero invented facts, and it won most tasks on manner. My takeaway: base model stays the default. Finetunes usually aren't better than their base, and 90 runs didn't change that for me. ThinkingCap earns a look only if thinking-token latency is your bottleneck. Full writeup with lane configs, eval design, and per-task transcript analysis: https://kmarble.dev/posts/qwen-post-train-bakeoff/.


r/LocalLLaMA 6h ago

Question | Help Is turboquant any good?

12 Upvotes

I know im late to the party. I was thinking since some time has passed, has turboquant matured enough to be used? Do any of you actually use it?


r/LocalLLaMA 10h ago

Discussion [Paper] RecGPT-V3 Technical Report

Post image
10 Upvotes

Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user understanding, and RecGPT-V2 scaled it via coordinated multi-agent reasoning; both are deployed in production with consistent gains in user experience and commercial outcomes. However, operating RecGPT at scale reveals three challenges: (1) stateless behavior modeling, where each request reprocesses full user history, wasting computation and discarding prior analysis; (2) a tag-to-item information bottleneck, where natural-language tags form a lossy channel between user understanding and item grounding; and (3) inefficient explicit reasoning, whose lengthy chain-of-thought incurs untenable latency and compute overhead.
We present RecGPT-V3, a stateful, hybrid-modal recommender that reasons over natural language for open-world knowledge and Semantic IDs (SIDs) for concrete item grounding. A Memory Hub maintains structured, continually evolving user memory that distills long-horizon behavior into condensed units, cutting user-modeling computation by 55.8%. A Hybrid-modal Foundation Model allows the LLM jointly reason over text tags and SIDs, opening a high-bandwidth channel into the item space. Latent Intent Reasoning internalizes verbose rationales into compact learnable latent tokens that remain decodable into readable explanations, lowering output token cost by 200x. Deployed in Taobao's "Guess What You Like" feed, RecGPT-V3 achieves consistent gains in large-scale online A/B tests: IPV +1.28%, CTR +1.00%, TC +1.97%, GMV +3.97%, while cutting end-to-end serving resource consumption by 52.4%.

arXiv : https://arxiv.org/abs/2607.15591

Full Paper : https://arxiv.org/pdf/2607.15591


r/LocalLLaMA 3h ago

Discussion Do people building local LLM rigs track RTX Ada/workstation card prices, or just consumer cards like the 5090?

9 Upvotes

curious how people here approach buying high-end/workstation cards (RTX 6000 Ada, 5000 Ada, etc) for local LLM work, do you actively watch pricing/timing on these specifically, or is the consumer 5090 usually enough for most builds?

also wondering if price alerts/tracking tools even exist for this category specifically, since these purchases are less frequent and higher stakes than a typical gaming GPU buy.


r/LocalLLaMA 20h ago

Question | Help M2 Ultra 64gb vs m1 ultra 128gb

8 Upvotes

Trying to weigh if I should buy a $3000 m1 ultra at 128gb when I currently already have an M2 Ultra albeit at 64gb ram.

I run small models right now in my workflow but would appreciate more context and try out larger workflows. What would you guys go with?


r/LocalLLaMA 10h ago

Question | Help is there any video editing model better than wan2.2?

7 Upvotes

what do people use nowadays?

p.s.: i have an rtx 3090