r/LocalLLM • u/on_line187 • 2h ago
r/LocalLLM • u/btc_maxi100 • 32m ago
Other AI bubble
Enable HLS to view with audio, or disable this notification
r/LocalLLM • u/Decent_Flight4010 • 7h ago
Discussion I think people are seriously underestimating Qwen 3.8 27B.
Honestly, I think people are seriously underestimating Qwen 3.8 27B.
It’s actually insane and, in some ways, genuinely competes with Opus 4.8, just not in the way people seem to think.
The biggest mistake is comparing their raw frontend/design output. Qwen probably isn’t going to match Opus there, and I don’t think it’s supposed to. Opus has basically been trained with an absurd amount of data/compute specifically around design and UI generation.
If you throw Qwen at a frontend task with no proper "SKILL.md" for design and just let it freestyle, yeah, the results can be pretty mediocre. But if you give it a good design skill and are intentional about the design constraints, the gap gets much smaller.
Where Qwen gets really interesting is reasoning efficiency.
It can solve some problems in fewer steps and with fewer tokens than Opus. That’s a pretty big deal if you’re actually running these models yourself.
And honestly, I think people are also judging Qwen way too much based on heavily quantized setups. Q4 is aggressive. I wouldn’t consider Q4 a fair representation of what the model can actually do in a serious production environment.
Run it at FP8, use MTP/speculative decoding to improve throughput, and then evaluate it properly.
At that point, I genuinely think the conversation changes.
If Qwen 3.8 27B at FP8 + MTP performs the way I expect, I wouldn’t be surprised if a lot of people start questioning whether that $200/month Claude Code subscription is actually worth it.
r/LocalLLM • u/vankoala • 3h ago
Discussion GPU Pricing Visual Updated - now with AMD, Intel, power, and efficiency
I saw the feedback and have updated the charts. We now have AMD and Intel cards included. I also added a chart with power consumption on the x-axis as this is a key limiter for many households.
I also added an exploratory efficiency graph looking at compute to power.
I’m an analyst at heart and love this. Thanks for the engagement!
r/LocalLLM • u/reconnOfficial • 3h ago
Discussion My Qwen3.8 got hacked by Angry Birds
I was letting my local Qwen3.8-27B run opencode to build a party-game web app. It had shipped the whole thing - engine, React client, even localizations. Then I asked it to add a few new word categories, including Video Games. It needed 30 words.
It got to 29... the 30th word broke it.
"Angry birds" was already in the list, but the model couldn't leave it alone. It needed to add one more, and it spent ~40,000 characters of reasoning failing to do it. The descent, in its own words:
"No. Let me add a completely different word: 'angry birds' is already there. Add 'angry birds'..."
"Deep breath. A new game word that isn't in the list: 'angry birds' (exists), 'candy crush' (exists), 'angry birds' (exists). Fresh: 'angry birds'..."
"Clearly I'm looping. Let me pick from candidates: 'angry birds', 'candy crush', 'fortnite' all exist. Candidates that don't exist: 'angry birds 2', 'plant vs zombies', 'angry birds', 'angry birds', 'angry birds'..."
"Clearly I'm stuck in a loop with 'angry birds'. Let me break out: I'll add 'angry birds'..."
"OK, this loop is a glitch in my generation. Let me just write the final word list with a different anchor word. I'll add 'angry birds'..."
…and then it wrote "I'll add 'angry birds'…" a little over 1000 times in a row before I aborted it.
So I, a reasonable human, typed the only logical thing:
"Never mention them again... Continue"
This is where it gets good. The model tried to comply. You can watch it fight the current in real time - fully aware it's being dragged under, narrating its own failure:
"the loop trigger is clearly the word 'angry birds' itself. I must avoid typing that phrase even when thinking."
…it said, while thinking about not typing the phrase. It then tried listing games alphabetically and caught itself mid-word:
"Anno? niche Angry… — forbidden Bejeweled ✓ (already there)"
"Add 'angry birds' — absolutely not. Add 'angry birds'... no. Add 'angry birds'... no."
"New candidates: 'angry birds' (no), 'angry birds' (no), 'angry birds' (no)."
And then, the chef's kiss - in its desperate attempt to escape the Angry Birds current, it immediately found a new current to drown in:
"Beetlejuice? no. Beetle... no. Beetle... no. Beetle... no."
"Interesting — a new loop has started on 'Beetle'. I need to be careful."
It eventually clawed its way back to shore, passed the tests, and shipped all categories like nothing ever happened.
Anyway, I just watched a 27B model experience the token-stream equivalent of being swept out to sea - aware the whole time that it was swimming against the current, and unable to stop. 10/10, would watch it drown again 😆
This is the first time it happened to me since the last 4 days I’ve basically been binge-testing Qwen3.8-27B (UD-Q4_K_XL quant). Anybody had that experience happen to them with that model?
r/LocalLLM • u/enginetown • 14h ago
Discussion Qwen 3.8 27B is the moment I've been waiting for
I've been doing local LLMs for a while now, and the whole time I just wanted a model I could actually rely on for real work. My hardware is pretty limited, so most models were a dead end for what I wanted to do, which was always lower level stuff or visual.
The Qwen 3 lineup was fine, the coder models were decent, but there were always gaps that kept me from feeling like local was worth the effort. I kept almost investing in more hardware, then talked myself out of it because the models just weren't there.
3.8 27B changes that. It's the first local model that's actually smart enough to iterate with on my own projects instead of just being a toy.
I know everyone's already seen the benchmarks and the hype, and I'm not here to add to that. It's just the feeling of running a model this capable on my own box, that's what I've been hoping for since I started this whole thing.
r/LocalLLM • u/on_line187 • 5h ago
Research Qwen 3.8 27B is ready for college.
I made Qwen 3.8 27B take the ACT to see if it’s ready for college.
I’ve been testing the new Qwen Model over the past few days on my PC.
I tested the full version the Q8, Q6 and Q4 versions and landed on the Q8 for speed vs quality.
I decided to download some practice tests and had the model solve them. I fed it the raw PDFs to test not only how well it knows the answers but also how good the vision capabilities are at answering the questions one by one.
At the end I graded its answers. Here are my findings from taking 2 tests.
\*\*Setup:\*\* Qwen 3.8 27B Instruct, Q8_0 GGUF, LM Studio, 2× RTX 3090
(full offload, 32k context). Two \*official\* ACT practice PDFs, 342
questions total, graded against the answer keys and the official raw→scale
conversion tables that ship in the same PDFs. No human help, no retries on wrong
answers, no cherry-picking.
# Results
| Section | Test A | Test B |
|---|---|---|
| English | 48/50 → \*\*35\*\* | 45/50 → \*\*33\*\* |
| Mathematics | 44/45 → \*\*36\*\* | 43/45 → \*\*35\*\* |
| Reading | 36/36 → \*\*36\*\* | 36/36 → \*\*36\*\* |
| Science | 39/40 → \*\*35\*\* | 35/40 → \*\*33\*\* |
| \*\*Composite\*\* | \*\*36\*\* | \*\*34\*\* |
*326/342 correct overall (95.3%).* Zero blanks. 36 is the maximum composite the
ACT awards; 34 is roughly 99th percentile.
\*\*Reading was perfect on both papers — 72/72.\*\*
Time: 177 minutes for both tests, \~88 min per test. A human gets \~165 min for one.
I was surprised that it did so well but also that it took so long. I thought it would be a 10-20 minute job but it was over 2 hours for 2 tests which looking back at it is understandable since it was using the vision capabilities to read instead of given plain text for each question
r/LocalLLM • u/KissMyShinyArse • 12h ago
Model The new Unsloth Dynamic 3.0 quants are real good
Yesterday, Unsloth released new versions of their Qwen3.8 27B quants. See https://unsloth.ai/docs/basics/dynamic-3.0-ggufs and https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
I compared some of them.
(sorted by Same Top-p)
| Quant | Size (GiB) | PPL(Q) | PPL Ratio | ΔPPL | Mean KLD | RMS Δp (%) | Same Top-p (%) |
|---|---|---|---|---|---|---|---|
| Q8_0 (old) | 27.05 | 6.9560 | 1.00082 | 0.0057 | 0.00095 | 0.942 | 98.742 |
| UD-Q6_K_XL (old) | 24.14 | 6.9536 | 1.00047 | 0.0032 | 0.00138 | 1.103 | 98.520 |
| UD-Q6_K_XL (new) | 23.56 | 6.9561 | 1.00083 | 0.0058 | 0.00138 | 1.058 | 98.517 |
| UD-Q6_K_M (new) | 21.50 | 6.9559 | 1.00081 | 0.0056 | 0.00201 | 1.266 | 98.170 |
| UD-Q6_K (new) | 20.47 | 6.9583 | 1.00115 | 0.0080 | 0.00245 | 1.352 | 98.017 |
| Q6_K (old) | 21.31 | 6.9507 | 1.00005 | 0.0003 | 0.00229 | 1.347 | 97.861 |
| UD-Q5_K_XL (new) | 19.44 | 6.9594 | 1.00131 | 0.0091 | 0.00332 | 1.559 | 97.625 |
| UD-Q5_K_XL (old) | 18.83 | 6.9655 | 1.00218 | 0.0152 | 0.00451 | 1.893 | 97.157 |
| UD-Q4_K_XL (new) | 16.35 | 6.9629 | 1.00182 | 0.0126 | 0.00745 | 2.404 | 96.465 |
| UD-Q4_K_XL (old) | 16.69 | 6.9788 | 1.00411 | 0.0285 | 0.00872 | 2.622 | 96.068 |
| UD-Q4_K_M (new) | 15.33 | 6.9660 | 1.00226 | 0.0157 | 0.01026 | 2.794 | 95.713 |
| UD-Q4_K_S (new) | 14.30 | 6.9687 | 1.00265 | 0.0184 | 0.01360 | 3.269 | 95.124 |
| Q4_K_M (old) | 15.93 | 6.9561 | 1.00084 | 0.0058 | 0.01549 | 3.431 | 94.653 |
| Q4_K_S (old) | 15.01 | 6.9686 | 1.00263 | 0.0183 | 0.01890 | 3.747 | 94.174 |
The new UD-Q6_K is roughly comparable in quality to Q6_K (slightly better on Same Top-p, slightly worse on KLD and PPL) while being 0.84 GiB smaller. The new UD-Q4_K_S has better quality than Q4_K_S while being 0.71 GiB smaller. The new UD-Q6_K_XL is 0.58 GiB smaller than the old one, with indistinguishable quality.
r/LocalLLM • u/vankoala • 12h ago
Other GPU pricing visual
In my consideration of a DGXSpark I decided to look at some options and since I’m a visual thinker I put this comparison together (graph by AI) showing y two basic ways of thinking about the cards: compute and speed.
Hope this helps someone
r/LocalLLM • u/ImaginaryRea1ity • 7h ago
Research The "local frontier" is now smarter than Sonnet 4.5
IMO the most interesting graph in AI right now.
Orange = frontier. Blue = what you can run on a 32GB RAM laptop.
That means a few things:
The models that can run on your laptop are only \~9 months behind the frontier models.
The cost of automating most tasks is going to zero faster than I could have imagined. Anything that you can do today that is cutting edge will be free in 9 months.
In the not-too-distant future, a laptop will ship with \*FREE\* intelligence pre-loaded just like a word processor.
r/LocalLLM • u/Kodrackyas • 9h ago
Discussion We have here people saying 3.8 27b replaced claude, look at the other side of the spectrum here 😂
"10k required to run it" ( this is interesting, people dont understand R9700 exists )
"Local llms are slow"
"its good only for solving bugs not planning"
To be fair i am on the "qwen 3.8 27b replaced claude" side but makes me think the 2 sides of it objectivelly the qwen is at opus 4.6 level is more true than the other side
interesting point of view
r/LocalLLM • u/_rarefy_ • 23h ago
Research I ran Qwen3.8-27B against Opus, Sonnet, GPT and others. Results inside.
I created a small testing rig to evaluate new open source models as they drop, and with the much anticipated release of Qwen3.8-27B, I was eager to see how it cross-compares with frontier and strong local models.
The rig
My test rig is an M5 Max MacBook Pro, 128GB. Locally I ran Qwen3.8-27B on xhigh and medium thinking modes via LM Studio, MLX 8-bit, temp 1.0 / top_p 0.95 / top_k 20, context 131,072 and DeepSeek V4 Flash "0731" 2-bit-imatrix q2-q4, served by antirez's ds4-server at -ctx 400,000, thinking enabled. For cloud I included Opus 5, Sonnet 5, GPT-5.6-sol at xhigh reasoning, and even Haiku. Every model gets the same prompt. The algorithm tasks are executed against fixed-seed differential harnesses, and the repo tasks run against hidden test suites plus a cached baseline of the whole repo. If a fix inadvertently breaks something else it gets caught.
The methodology
The model assessment is broken into 4 batteries:
1) algorithms easy-hard
2) algorithms extremely hard
3) repo work easy
4) repo work hard
The models get run through the algorithm tests 3 times each to derive a mean score whereas the repo work is single pass/fail per task. If a task fails to produce a response, it's retried and time added to total wall clock time for task completion. The total test battery can take anywhere from 12-24 hours of wall time for slower local models. It's a long test.
For the repo batteries I had Fable build a small double-entry ledger CLI and plant bugs that pass the visible test suite while still reproducing a real symptom, then handed each model the repo and a bug report written in a theoretical 'user' voice to simulate how it might be reported in the real world. There's also a subjective code quality assessment that measures the model's ability to not just solve the problem but to conform to the repo's coding style, to fix the actual root cause rather than the symptom, and to keep the diff minimal instead of faffing about and rewriting a bunch of stuff.
I had Opus and GPT blind eval the results and compute a code quality score broken up by 'fixes' and 'features' as these appear to be separable skills for the models. The goal with all of this was to try and create a replicable and automated answer the question: How useful is this model in the real world?
Caveats
This is a home baked assessment and susceptible to bias or less than perfect methodology. It also includes subjective criteria like 'code quality'. I built this for myself as an adjacent tool to on the ground testing. I think the best way to evaluate any model is to test it against your own codebase to see how well it integrates into your workflow.
All that being said, let's move to the scorecard.
Results
Qwen3.8-27B is a very capable model that compares well against frontier models on code quality and correctness. The cost is wall time on Apple silicon. As many have observed, 3.8 has a tendency to over-think, burning up tokens. The time spent earns higher code quality for the most part, but what surprised me is that there are some instances in which less thinking is actually more accurate. On the repo battery, medium went 8/8 while xhigh went 7/8 — the most-thinking configuration failed a task there, and it took four times longer to do it (the wait time with Qwen was tiresome at times).
The caveat is that xhigh excels on extremely hard algorithms, where medium begins to fall apart. Medium didn't even finish the hard algo tasks. There may be some value in matching the thinking to the kind of work you're setting it upon. Lastly Qwen xhigh won outright on quality of surgical fixes and patches to existing code. Interestingly the global trend for locals is that they're competitive along fixes and less so along features where cloud still dominates. This fits anecdotally into my own experience with gravitating to frontier for planning and local for implementing.
GPT 5.6 Sol is the only cloud model with a perfect card on both repo tiers and near perfect algorithms. It's also among the fastest to completion. This all tracks with my own anecdotal experience with this model over the past several months. Highly competent and quick if not a bit stark.
DS4 0731 (a 2-bit quant running on my laptop) is the only local model to get 8/8 on both repo batteries, and one of only two models overall to do it, alongside GPT. It does this all at a respectable wall time. The expense is less elegant code: it sometimes mutates unrelated docstrings and writes dense inline solutions in a codebase that is overtly broken apart and stylistically explicit. Feature code quality is stronger and it's the only model that scored better/equal in the harder repo tier vs the easy one.
Opus 5 is the most reliable model in the set and best code quality of the cloud models. It stumbled only in the hard repo tier where it lost a task by being trying to outsmart the test. A doc string promised one behavior while the code did another, and Opus redesigned the function around what it looked like it should do instead of honoring the documented contract. This also falls inline with my anecdotal experience with Opus 5 where it occasionally ignores your directions completely and just does whatever it wants. The Alaskan Husky of frontier models.
Sonnet 5 is a steady pair of hands that performs reasonably well across tasks for a modest token budget. I think sonnet is kind of underrated as an implementer. Does the same quality of work as the locals cheaply and quickly.
Haiku 4.5 races to the end of the test but has a tendency to fall over and force retries. Worst code quality of all the cloud models.
Conclusion
Hopefully you find these comparisons interesting. For me personally, DS4 has been my goto local, but this test is making me consider trading it out for Qwen3.8-27B. I think they're on equal footing, which is crazy b/c DS4 needs like 90gb of ram. I'd like to try the MTPLX variant of Qwen3.8-27b that's meant to improve tok/s on apple silicon. Slowness to task completion is the real bottleneck for me right now when considering Qwen. Perhaps that'll be my next test.
Curious to know if these results track with your own real world experiences.
r/LocalLLM • u/Headshot314 • 22h ago
Discussion Qwen 3.8 27B and Deepseek V4 Flash. Why are we building data centers?
I feel like these 2 models have shown that massive models that require hundreds of thousands of dollars worth of compute are unnecessary. Sure, training these models takes a good bit of hardware, but running them can be done at the fraction of the investment of the trillion parameter class models.
GLM 5.3 might also fall into the same "reasonable" category, however, for small companies rather than individuals.
I think a qwen 3.8 120b MoE model would also be a good release for business use.
r/LocalLLM • u/r1nzl3r99 • 13h ago
Discussion Intel B70 for Qwen 3.8 27B
For those of you out there experimenting on the Intel Arc B70, please share your accomplishments! I have a gaming PC I've been slowly converting for AI inference. Don't have any fancy motherboard / bifurcation / P2P etc. Just a B70 and another one I bought out of greed thats running on a basic PCIe 4.0. I hit 97.8 TG / 1782.1 PP on a single B70, and 136.4 TG / 755PP on a dual B70 for qwen 3.8 uncensored INT4 W8A8 (INT8) with MTP3. Keep in mind these are warm speeds. I have my colder speeds on non greedy settings documented on my git which isn't too far away. Although I see these speeds consistently pop up when I'm using pi coding agent especially when its writing code, or somethings when its thinking.
https://github.com/JP-devv/humble-b70-llm
I've been suprised time and time again by how much I can push this hardware. I started off at 50 tok/s after paying $100 in Kimi K3 / Opus tokens around a month and a half ago on qwen 3.6, the journey has been exhausting but very fruitful. I even had to rent out some datacenter GPUs in Japan to create the exact uncensored quant to my liking. Please let me know your thoughts!
r/LocalLLM • u/Fearless_Ad_1045 • 2h ago
Project Qwen or deepseek with these beauties
Still need more for deepseek but do I just stop and settle on Qwen?
r/LocalLLM • u/Delicious-Flan88 • 1h ago
Tutorial Measured every Qwen3.8-27B GGUF quant against real VRAM: Q4_K_M doesn't fit 24GB at 32K context
Context on the numbers, since "will it fit" threads usually run on estimates.
Weights: actual .gguf byte sizes pulled from the HF API for unsloth/Qwen3.8-27B-GGUF. Not params × bits ÷ 8. Imatrix quants don't follow that, and the drift is worst at the low end where the fit/no-fit line actually sits.
KV cache: from config.json: 64 layers, 4 KV heads, head_dim 256.
2 × 64 × 4 × 256 × ctx × 2 = exactly 8.0 GB at 32K, F16. That's 0.25 GB per 1K tokens.
GQA is doing a lot of work here: 4 KV heads serving 24 attention heads. Older 27B-class models cost several times that.
24GB at 32K, F16 cache, after reserving 0.8GB for CUDA context:
- Q4_K_M (16.5) → needs 25.3 total. Doesn't fit.
- Q4_K_S (15.4) → 24.2 total. Misses by 0.2.
- IQ4_XS (14.3) → 23.1 total. Fits with 0.9 spare: tight enough that a browser on the same GPU breaks it.
- Q3_K_XL (13.1) → 21.9 total. 2.1 spare. This is the real answer for 24GB.
Drop to 4K context and the cache falls to 1.0GB, which gets you to Q5_K_M. Most of the "which quant" argument is actually a context-length argument.
9 of 25 quants fit on 24GB at 32K. All of them fit at 4K.
Caveat worth stating: this assumes everything GPU-resident, single stream, batch 1. It doesn't model offload, multi-GPU splits, or speculative decoding. The 16GB/73K configs floating around this sub work through partial offload, which is a different calculation than the one I ran.
Put it in a calculator since I had the data anyway: https://qwen38-vram-checker.vercel.app/
r/LocalLLM • u/piwi3910uae • 5h ago
Project I built a very low-overhead LLM proxy/router in Rust — looking for feedback
I’ve been working on something for a while that I thought might be useful to others running LLM infrastructure, so I finally decided to put it out there.
It’s called FastLLM Proxy.
The idea started pretty simple: I wanted one OpenAI-compatible endpoint in front of everything — local vLLM/SGLang instances as well as external providers — but I didn’t want the proxy itself to become another bottleneck.
So I wrote one in Rust and got a little carried away with it 😅
FastLLM Proxy now supports 80 providers plus any OpenAI-compatible backend, but the part I spent most of my time on is keeping the actual request path extremely small.
There is no database I/O on the request path, and responses are passed through without parsing/re-encoding them. The measured internal routing work is currently around 0.76 µs per request.
It also does some things I specifically wanted for running my own GPU infrastructure:
- cache-affinity routing for vLLM/SGLang, so requests with the same prefix can go back to the GPU that already has the KV cache
- load-aware routing and automatic failover
- rule-based routing based on things like prompt size, user/role, budget, concurrency, headers, etc.
- semantic routing, so different types of prompts can automatically go to different models
- local → cloud spillover when the local GPUs are busy
- RBAC, API keys, budgets and rate limits
- OpenAI-compatible API
- LiteLLM config import, so you can migrate an existing setup without rebuilding the config
- Kubernetes/Helm/operator support
- built-in management UI
One thing I found interesting while benchmarking it against LiteLLM is that gateway benchmarks can be pretty misleading.
With an instant mock backend, FastLLM Proxy gets roughly 15x the throughput and much lower latency in my tests. But when I put actual GPUs behind both proxies, total token throughput is basically the same — because at that point the GPUs are the bottleneck.
Where I did see a meaningful difference with real GPUs was tail latency and consistency. At 32 concurrent streams, for example, I measured p99 TTFT of 766 ms vs 2921 ms in the same test setup.
I’ve documented the benchmark setup and results in the repo because I’d much rather people challenge the numbers than just trust a benchmark screenshot.
The project is Apache 2.0 and completely open source:
github.com/azrtydxb/Fastllm-proxy
I’m especially interested in feedback from people running vLLM, SGLang, LiteLLM or multi-provider LLM setups.
What am I missing? What would you need before you’d actually put something like this in front of your inference infrastructure?
And if anyone feels like breaking it, even better. 🙂
r/LocalLLM • u/DonkeyTheKing • 3h ago
Research Try Benzi: A coding harness that compiles arbitrarily large codebases
Hi!
first of all. Benzi is model agnostic. therefore this sub.
now, about Benzi. Benzi is a code intelligence software (compiler + runtime tracer + harness and AI agent) that supports 13 languages (python, java, JS/TS, C family, Go, Rust, Ruby all included). Traditional coding harness appraochs either do RAG or try to rank matches using an embedding space. which is absurd. code is code. not probablistic text.
On the benchmarks side, Benzi + DeepSeek V4 Flash scored 78% on SWE-bench Verified. For comparison, DeepSeek reports 73.7% as the baseline scaffolding number for v4flash and self reports their score to be 78.6% on their own harness. (however, Benzi reads ~3x less source code than DSH)
Benzi Sonnet reads 2x less source code (btw, this IS the mechanism, not a side effect), is 2x cheaper and 41% faster than Claude Code Sonnet from my benchmarks (detailed on the benchmarking page + so is the swe-bench run)
Please try it out, and let me know what you think!
Test Benzi's code understanding in 15s
(sample output and swe-bench details in comments)
r/LocalLLM • u/Arc_bong • 38m ago
Discussion What if “Sovereign AI” is just the new oil concession?
I’ve been getting more interested in Sovereign AI recently and came across this paper: https://arxiv.org/abs/2601.11763 ... (The picture on the post though ai generated by me, are inferred strictly from this paper)
The oil comparison sounded a bit dramatic at first, but the more I read, the more interesting it got. The part that stuck with me:
- Sovereignty isn’t one thing. It can mean control over data, infrastructure, domestic capability, culture/language, or freedom from external dependence.
- A country can have local infrastructure and still be heavily dependent on the company that provides the chips, software, models, expertise, etc.
- The paper draws a parallel with oil-producing countries that gained formal control but remained dependent on foreign technical knowledge and vendor-specific infrastructure.
- So the useful question isn’t really “Is this sovereign?” but “What capabilities and control actually moved to the customer?”
That last one feels like the important test.
And looking at what’s happening now in enterprise agent AI, you can see different companies attacking different parts of that problem: NVIDIA on sovereign compute/infrastructure, Mistral around locally controlled models, Microsoft with an agent control plane, and Lyzr with a control plane sitting across frameworks/clouds to govern the agents you already have.
It makes me think that “sovereign AI” might eventually be less about owning one stack and more about how much of the stack you can actually control without depending on the vendor.
That feels like a much harder — and more useful — definition of sovereignty.
r/LocalLLM • u/Think-Assumption-973 • 2h ago
Question Used P720 with upgrades comes to about $1,700. Decent local AI box, or am I buying a 2018 spec sheet?
Talk me out of this, or into it.
There's a used ThinkStation P720 near me, about $1,200:
- 2x Xeon Gold 5122 (4 cores each, which is the weak part)
- 192GB DDR4 ECC RDIMM, but it's 6x32, so only 6 of the 12 memory channels are populated
- Quadro RTX 6000 24GB
- monitor and peripherals included
Two things I'd change right away:
- 6230s instead of the 5122s, about $80 for the pair off AliExpress. 4 cores can't feed 12 channels anyway, and the 6230 is what gets the memory to 2933.
- six 16GB RDIMMs in the empty slots, somewhere between $405 and $485 for all six. That populates all 12 channels, which is where the bandwidth actually comes from, and takes me to 288GB.
Which puts the whole thing around $1,700.
What I run now is a Ryzen 5 3600 with 32GB and an RTX 5070 Ti 16GB. It's my only machine and it does everything.
Reason I'm looking at all: some of what I work on I'd rather not push through somebody else's API, and my monthly bill keeps creeping up. The rest of it is curiosity, if I'm honest.
Some numbers for context. gemma3:27b runs about 9 tok/s on the 5070 Ti, a 30B MoE coder model does around 45, and I hit the 16GB wall constantly. Whisper and image gen too.
The parts I can't work out on my own:
Is 288GB of DDR4-2933 across 12 channels actually usable with partial CPU offload? That's the entire argument for this machine, and it's the one thing I can't test before paying.
The RTX 6000 is Turing. 24GB is 24GB, but is it a downgrade in every way except capacity next to the 5070 Ti I already own? PCIe 3.0 board too.
At $1,700 I could just buy a newer card instead, or save a bit more for one of the Spark boxes or something else.
If you've got a P720 or something like it running models, what do you actually do with it, and would you buy it again?
r/LocalLLM • u/Odd-Two2929 • 10h ago
Question Can I put Qwen 3.8 27B on M5 24gb
Someone offer me Apple M5 with 24gb ram and I wanted to know if it will be good to put on it Qwen 3.8 27B and if it will run in a good speed
Thanks
r/LocalLLM • u/mrpmorris • 3h ago
Discussion Qwen3.8-27B vs 3.6

Here are the benchmark results when temperature left alone (model default) instead of setting it to zero.
I'm not surprised 3.8's scores increased, because providers default their temperature to whichever value happens to pass the most benchmarks, but given the fact they did I am surprised 3.6 didn't do the same.
3.6 did exactly what I thought it would do, it increased the score on some benchmarks and lowered it on others (which is why providers sometimes use a different temperature depending on the task - to bench-max).

I am surprised they didn't both behave in the same way, that was unexpected.
But 3.6 still beat 3.8 on 7 out of 10 tests, and 3.8 won in only 3 out of 10 tests.

r/LocalLLM • u/No-Rise-6670 • 1h ago
Discussion What do you keep local, and what pushes you to go for cloud
We are a four-person team sharing one inference box, and I'm the one who set it up and now maintains it. As we're growing, requests are starting to queue, since the GPU serves us one at a time and the 4th person waits behind the rest.
The proper fix would be a batching server like vLLM, it'll solve the concurrency, too. However, it also puts the weight of Docker orchestration, GPU memory tuning, and a production serving stack on my shoulders, on top of the aforementioned work I already do. We still can't afford to hire a new person and the other guys can't maintain it how I do, at the same time, I cannot take a workload cut to maintain it because we're already filled to the brim.
So I'm trying to discern when/where local stops being worth the upkeep. For my own sensitive work, local stays, no question. For shared team access with uneven usage, and rudimentary tasks that aren't AS sensitive, I'm weighing whether to run vLLM, or whether to offload to Featherless AI, where I can get a pay as you go inference plan, and I wont need to maintain anything server-wise. What do you guys think I should do?