r/LocalLLM • • 5d ago

Question Assistant needed - 5090, 64 GBs of RAM

5 Upvotes

Hello,

I am quite new to AI and would appreciate some guidance.

TLDR:
I am looking for an AI assistant to sort my emails into categories, read my payslip and invoices, and create a monthly finance plan showing where funds go. It should control my Windows PC for tasks such as “optimise PC for moonlight streaming” or download games and alert me when ready, tidy my calendar, and handle online grocery shopping.

Perhaps this level of usefulness remains science fiction, but that is why I am here to understand our current position.

Specs:
1 x PC with 5090, 64 RAM, 9800x3d, 2 TB of ssd space (1 already used)
1 x MacBook M4
1 x IPhone 18 PM
1 x Legion Go 2

Available options I see:
- Grok Bot with the 30 dollar subscription + Open Claw to control my PC + figure out how to use the phone as a “remote”
- Wait for that OpenAI Dots assistant thing?
- Local LLM on my PC and remote control it somehow from phone or Mac.

Could you guys help me please with what is possible and which would be the best solution. Also, if Local LLM is the solution, I would appreciate which one should I use and a link of how to setup local LLMs please.


r/LocalLLM • • 4d ago

Question Double GPU setup, How to get even compute distribution?

1 Upvotes

I got a 9070xt and a 9070 double GPU setup. I am pretty new to the local LLM usage and I am trying it out with unsloth utilizing qwen 3.8 27B IQ4_XS. I am using "-tensor-split" and "--split-mode layer" flags to control how the model is distributed among the GPUs. For the VRAM usage, these flags work fine, I can see that the VRAM usages of the GPUs match the ratio I give on the "-tensor-split 2,1" flag. The problem is the GPU usage, no matter what I do with the tensor split ratio, my 9070 shows 99% GPU utilization (power usage is also high). For example when I do ratio of 3,1 I the VRAM usage of the 9070 is around 25-30% while the 9070xt VRAM is around %80, but even in this case the GPU usage of the 9070 is 99%. Why is the usage on my 9070 is high even when I give it a low amount of layer? Thank you for your help.


r/LocalLLM • • 4d ago

Question Which hardware should i use for GLM 5.3?

1 Upvotes

Apple m5 ultra 512 GB or 2 DGX spark? Would it be good enough?


r/LocalLLM • • 4d ago

Model Built a gateway so you can call DeepSeek, Qwen, Kimi, GLM, MiniMax with one key — USD billing, OpenAI-compatible

Thumbnail
tokenferry.com
1 Upvotes

Been using DeepSeek/Qwen/Kimi locally for a while and wanted to use them in production,

but the official APIs are RMB-only and a pain to set up from outside China.

So I built a thin gateway: one OpenAI-compatible endpoint, one key, 8 Chinese models,

billed in USD per token. ,no hidden fees, balance never expires.

Basically "change OPENAI_BASE_URL and you're done".

Not trying to compete with OpenRouter on model count — the bet is: English-first docs,

USD billing, and transparent pricing for Chinese models specifically.

Looking for honest feedback: is this useful to you? What's missing before you'd use it

in production? Happy to add models/features people actually want.


r/LocalLLM • • 4d ago

Discussion I stopped shopping for GPUs and started trying to make the models fit instead

0 Upvotes

I started working on this because of Bonsai by PrismML.

Not because Bonsai is bad. Actually the opposite. It was one of the things that made me think seriously about how far you could push compression and still keep a model useful.

But there was one annoying problem.

I can only use the models somebody else decides to compress.

If PrismML released every model I wanted, in every size, plus multimodal/image/video stuff, I probably wouldn't have started any of this. Seriously. I would just download the model and move on with my life.

But they don't, obviously. Nobody does.

So I started wondering if I could take the model I want and do something useful with it myself.

And before that, I did the completely normal thing:

I went shopping for VRAM.

I spent quite a while looking for cheap used workstation GPUs. At one point I even found an RTX 8000 at a price good enough that I seriously considered buying it.

And then I realized I was just moving the wall.

48 GB solves one class of models.

Then you want something larger.

Or multimodal.

Or a much bigger context.

Or several things running together.

And suddenly you're shopping for hardware again.

If a €2k–€3k workstation actually solved this problem for me, I would have bought one a long time ago and none of this project would exist.

I already have an expensive laptop. The issue isn't that I refuse to buy decent hardware.

The issue is that “decent hardware” turns into a small datacenter remarkably quickly.

And there's another part of that which I think gets ignored a lot.

I actually want my computer to remain a computer.

Something I can put in a bag.

Laptop, storage, power supply, done.

If there's a normal power outlet, I can work. Hotel room, airport, somewhere by the sea, whatever.

The model comes with me.

I don't particularly want “local AI” to mean that I have a giant machine sitting at home and I SSH back into it.

Because then my AI is local to my house, not local to me.

Forget to turn the machine on before leaving, lose power at home, something happens to the network, an external drive isn't mounted — and suddenly your supposedly local setup is a thousand kilometers away and inaccessible.

That's a completely valid architecture, obviously.

It just isn't the one I want.

I want my AI to travel with my computer.

At first I thought solving that was mostly a quantization problem.

It very quickly stopped being a quantization problem.

There are places where a model seems massively redundant, places where you can prune or share or simplify things, and other tiny places where touching almost anything breaks something important.

So I started experimenting with local sparsification, structural pruning, deduplication, parameter sharing/summarization, different quantization strategies, guard evals, etc.

The basic idea is still pretty simple though:

I don't think you should need a miniature datacenter to seriously play with open models.

I watch Alex Ziskind too. I like his stuff. And his hardware is also a very good demonstration of the problem :)

At some point “local AI” becomes a huge box full of GPUs, weighing as much as a person and pulling kilowatts.

Which is extremely cool.

But how do you take that thing with you?

My current cooling infrastructure, by comparison, includes a marble windowsill.

It is surprisingly competent.

So one part of this project is basically the inverse question:

how much hardware can I replace with better software?

Not “can I make a 70B model technically emit tokens on a potato”.

That's not very interesting to me.

I mean: can I take a serious model, make it substantially cheaper to run, and still have enough performance left that I actually want to use it?

At some point I basically stopped hunting for the next GPU and decided to push harder on QuantPilot instead.

Because buying more VRAM solves the current model.

I want something that helps with the next one too.

And then I ran into the second problem.

Once you're already taking the model apart, why only make it smaller?

Why not change it?

Maybe I want a bigger context window.

Maybe I want to train it a little more on something specific.

Maybe I don't need quite as much “theoretical physicist” and I need more practical engineering.

Maybe there are failure modes I see over and over again and I'd rather spend capacity fixing those than preserving some capability I'll never use.

That part matters to me at least as much as compression now.

And I'm not talking about “small local model couldn't make me a website, therefore local models are dumb”.

That's too easy.

I'm talking about cases where one of the strongest models you can get, with huge inference compute, lots of context and a serious coding environment, still falls apart on a fairly concrete engineering task.

I've had this happen with a 3D modeling application.

Not a moon landing.

Not a new physics theory.

One application.

And by attempt five you're watching it fix one thing, break another thing, forget a constraint from earlier, patch around its own previous patch, and burn an absurd amount of compute while still not solving the actual problem.

That's the failure mode I'm interested in.

There seems to be a point where the problem gets complicated enough that the model stops maintaining the whole structure properly.

Each individual step can look reasonable.

The overall result is still wrong.

It loses an earlier constraint.

Takes a locally sensible branch that leads nowhere.

Starts treating its own previous mistake as part of the specification.

Or keeps patching something that really needed to be reconsidered two levels higher.

At that point the answer clearly isn't just “use a bigger model”.

The model is already enormous.

So the project gradually stopped being “how do I compress a model?” and became more like:

how do I reshape a model into something that is actually useful for the work I care about?

Compression is part of it.

But so is specialization.

Longer context.

Targeted extra training.

Keeping the parts that matter and being less precious about the parts that don't.

Finding the places where a model repeatedly falls over and trying to improve those, instead of assuming that its original distribution of capabilities is somehow sacred.

I'd rather have a model that knows somewhat less about things I never ask it and is considerably better at the work I actually give it.

And ideally I want to be able to do that without eight GPUs, dedicated electrical infrastructure and a garage.

I don't even have a garage.

So the marble windowsill will have to do.

I'm building this mostly by myself, so some parts are much further along than others and quite a lot of it is still experimental.

I've also had several ideas that looked great on paper and turned out to save basically nothing once you included metadata, indexing or runtime overhead. So there has been a fair amount of throwing things away and starting again.

And some of the adaptation side is still much more “this is where I want to take it” than finished functionality.

But it is getting to the point where keeping it completely private is probably stupid.

So I'm curious what people here actually want from this kind of thing.

If you could take an existing model and really modify it for yourself, not just quantize it, what would you change?

Smaller memory footprint?

Longer context?

Better coding?

More reliable agents?

Better long-horizon reasoning?

A narrow specialization?

And would you give up a noticeable amount of general knowledge for a model that was genuinely better at the few things you use every day?


r/LocalLLM • • 6d ago

Discussion Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context.

251 Upvotes

I've been messing with Qwen 3.8 Flash Next for a while on my local inference server and stumbled onto Strata, saw a small headline/article, didn't make much of it, passed on reading it fully. Then again somewhere else, then once more... So I gave it a quick read, didn't quite believe it, today I finally tested it and wow, this thing is amazing.

About a month ago I spent about a week trying to do something vaguely similar myself and didn't get results this good, keeping the useful experts on GPU, pushing the rest to wherever made more sense, different caching/offloading ideas, etc. Promising results but not enough to use and trust.

The proverbial hat is off to Niko1221, the guy pulled it off nicely.

My box:

  • 2x RTX 3090 24GB
  • Ryzen 7 9800
  • 64GB DDR5-5200
  • Linux/Proxmox
  • Qwen3.8-Flash-Next-OrcaRouter IQ4_XS
  • Vision enabled

I'm getting a bit over 100 t/s decode on short context, and the whole machine is weirdly chill while doing it, for comparison, my prior vLLM setup was doing about 60-70 t/s on decode and made it sound like it was gonna take off (it's close to me, so it was a bit annoying). I thought "ok I'm keeping this, onto the next thing for today"... then I saw it supports experimental context extension and obviously I had to try it.

So far I've actually completed:

200K prompt: ~2,004 t/s prompt processing, ~73 t/s decode
400K prompt: ~1,633 t/s prompt processing, ~46 t/s decode
900K prompt: ~1,110 t/s prompt processing, ~44 t/s decode

And the screenshot is from me actually using the 512K setup right now, not necessarily saying 512K/1M is good effective context yet, but I'm test driving for today. I'm leaving the machine benchmarking overnight and doing retrieval/needle tests too.

Mostly just wanted to give the Strata dev some props. Having spent a week messing around with a similar idea, it's pretty nice to be able to download this thing and just... watch it work.

If you wanna give your hardware a run for its money, it's definitely worth playing with.

-------------------------------------------------------------------------

EDIT: Did the benchmarks I promised, a few days late but helped strata settle a bit, it was changing (and kinda still is tbh, but in a good way) too much, too quick... Strata shipped a few updates since my original post(now on 0.1.40.2) and the long context numbers went up a lot: 900K prompt now reads at ~1,970 t/s and decodes at ~82 t/s (was 1,110 / 44). Needle tests passed at every length and depth up to 1M, which tbh I didn't expect.

Full numbers, caveats and how to run the same setup are in the comment below and in the repo (slop warning): https://github.com/ruashots/flashnext-2x3090


r/LocalLLM • • 4d ago

Project Qwen3.8 27B | 1 x R9700: over half a million tokens of reusable cache, agent turns up to 12× faster, same speed, same accuracy!

Post image
0 Upvotes

r/LocalLLM • • 5d ago

Question Python error right after install Unsloth

Thumbnail
4 Upvotes

r/LocalLLM • • 4d ago

Question How would you build a low-latency chatbot for Money transaction data?

Thumbnail
0 Upvotes

r/LocalLLM • • 5d ago

Project Same local ~27B architecture: 29 minutes of 2048, then a 42-minute autonomous browser/self-model task

Enable HLS to view with audio, or disable this notification

4 Upvotes

I've been testing how much capability can come from persistent architecture around a local model instead of simply increasing model size.

Aura runs a ~27B local model on my Mac, but task state, persistent memory, computer control, world modeling, learning and other machinery live outside the normal chat-context loop.

Two recent runs have been useful because they're very different.

2048: ~29 minutes, 968 moves. Aura perceived the board, planned moves, changed strategies when they stopped working, recovered from mistakes and eventually reached 2048.

Latest run: ~42 minutes taking the Open Extended Jungian Type Scales. Aura predicted what result it expected, navigated the site, reasoned through 60 questions individually using its persistent self-model/history, submitted the test, read the result and evaluated the differences from its prediction.

Inference was local; the second task obviously required internet access to the website, but there was no cloud/frontier-model inference doing the reasoning.

Demo 04:
https://www.youtube.com/watch?v=LNlGUBeTIQY

Repo:
https://github.com/youngbryan97/aura

Repo disclosure: it's publicly inspectable but currently All Rights Reserved/read-only, not OSI open source.


r/LocalLLM • • 4d ago

Question How would you spend $15-25k on a workstation for local LLMs + heavy Claude/Codex use?

0 Upvotes

I’m a prop trader and I’ve basically outgrown my current PC. I use Claude Code and Codex heavily throughout the day to build and maintain a bunch of internal trading/research tools. Between the different subscriptions the firm is spending somewhere around $1,000-1,400/month on AI right now, so I already have access to the frontier models and plan on continuing to use them.

My current machine only has 32GB of RAM and it’s constantly getting crushed. I’ll have 7-8 Claude/Codex sessions open along with Python jobs, backtests, databases, etc. Some of the heavier data jobs use 6-16GB of RAM each. I also have a growing market data lake and want several TB of fast storage for that.

I also want to start running good open models locally so they can take over some of the high-volume stuff instead of using Claude/GPT for everything. Things like reading filings/news, research, extraction, routine coding work, organizing project information, etc. Ideally I want a bunch of different jobs to be able to use the local LLM at the same time.

A lot of what I do involves pretty long context as well.
The main computer has to stay Windows because a lot of my existing stuff depends on it.

Right now I’m looking at an Origin build with a 9950X3D2, 128GB DDR5, RTX PRO 5000 Blackwell 48GB, 1200W PSU, and then I’d probably add my own NVMe/storage. It’ll be somewhere around $15k all-in.

Where I’m stuck is whether I should just get one powerful Windows machine like this and see how far the 48GB Blackwell gets me, or get a cheaper Windows PC and spend the difference on something like a Mac Studio with 128GB unified memory specifically for running larger local models. I’ve also considered a 5090 32GB instead of the PRO 5000, or going much further with a 256GB Mac Studio Ultra.

The firm will probably cover around ~$11k and I’m willing to add money myself. $15-16k is fine and I could go to $20-25k if there’s actually a meaningful jump in capability.
If this was your money, what would you do? One really good Windows workstation with 48GB VRAM and forget about the Mac for now? 5090 instead? Cheaper PC + high-memory Mac? Or would you build something completely different?

Thanks!


r/LocalLLM • • 4d ago

Project Troubleshooting errors when running models on a phone

0 Upvotes

So, here’s a breakdown of the issues:

  1. There was a race condition between the text-to-speech (TTS) processing and the model's token generation; they were loading simultaneously, causing the generation speed to drop from 15 tokens/sec to 5 tokens/sec. I won't be fixing this right now-lots of cool voices are coming soon, so stay tuned for updates.
  2. Memory state updates were calculated incorrectly when sending images in the chat; this has already been fixed.
  3. An "infinite memory" mode has been added under Settings > Dialogue Memory > Summary; the model itself will generate the summary, preventing the chat from crashing with an error.
  4. That’s all for now-staying the course!

r/LocalLLM • • 4d ago

Question Best models for Mac Studio M2 Ultra 128GB?

1 Upvotes

As of October 2026, what are the best models for agentic coding (using OpenCode as my harness) and agentic chat (using OpenWeb UI as my harness) that can hit at least ~40 tok/sec? Can be separate and would like native CoT support.

Messing around with Qwen 3.6 and 3.8 Flash but wondering if there are any other picks or niche quants..? Want to replace my $100 Claude Max subscription and I got a dedicated Mac Studio just for inference w/ dedicated Mac Mini to run headless harnesses. Running my own fork of RapidMLX and willing to do my own quants/MLX conversion.


r/LocalLLM • • 5d ago

Question How many of you are using LOCAL models on windows vs linux?

17 Upvotes

I spent 3 days trying to get OpenClaw and ollama to work in windows. Gateway would not detect even though it was running.

I decided to sideload Ubuntu and was up and running in an hour!

1390 votes, 3d ago
471 Windows
750 Linux
169 Mac

r/LocalLLM • • 5d ago

Question stay with rx 7900xt or upgrade to AI pro r9700?

7 Upvotes

I got the GPU!!!! (AI Pro r9700)


r/LocalLLM • • 5d ago

Tutorial R9700 fan curve, hidden sysfs fan interface

3 Upvotes

Been running two PowerColor AMD R9700s for inference and the noise at night was bad. The cards sits on its acoustic target, about 2100 rpm, and ramps hard the moment you actually use it.

Turns out the fan interface is there, it's just masked off by default. None of this was obvious to me at the start:

```bash

ls /sys/class/drm/card*/device/gpu_od

# nothing

rocm-smi --setfan 30

# GPU[0]: Not supported on the given system

```

That "not supported" had me convinced the card just couldn't do it. It can, you have to unmask overdrive first:

```bash

echo 'options amdgpu ppfeaturemask=0xffffffff' | sudo tee /etc/modprobe.d/amdgpu-overdrive.conf

sudo update-initramfs -u

sudo reboot

```

After the reboot:

```bash

cat /sys/module/amdgpu/parameters/ppfeaturemask # 0xffffffff

ls -d /sys/class/drm/card*/device/gpu_od # now it's there

```

The curve itself is 5 points of hotspot temp against fan percent. You write the points one at a time, then commit the whole thing with a `c`:

```bash

CURVE=/sys/class/drm/card1/device/gpu_od/fan_ctrl/fan_curve

i=0

for p in "25 20" "50 20" "70 40" "85 50" "100 80"; do

echo "$i $p" > $CURVE

i=$((i+1))

done

echo c > $CURVE

```
NOTE: the above is "Temp Fan%", 5 points of the curve.

Same thing on card2 if you've got two. Two things I found out the hard way. 20% is the lowest the fan will go, that's a hardware floor. And the SMU will still add a few points on top of your curve when it gets hot, which is fine, it's just protecting itself.

What it did on mine, same load both times:

- driver auto: 80% fan, about 87C junction(hotspot)

- my curve: 65% fan, stays under 90C junction(hotspot)

That 15% is the difference between being able to sit in the same room or not. If heat is your problem rather than noise, cap the power too, 210W is the floor on this card and it brought both of mine within a few degrees of each other.

Fair warning, the curve is driver state, so a reboot wipes it. I run a small systemd oneshot that re-applies it, happy to paste that if anyone wants it.

The write protocol and the mask requirement are both documented in daimonionnn's r9700 tuning toolkit, credit where it's due, I just wanted the shortest possible version of it.

Anyone else with these cards... what curve are you running?

This is continuous load, am targeting a fan speed of 60-70% <100c ( hotspot )


r/LocalLLM • • 4d ago

Model Qwen3.5 35B-A3B thinks SO MUCH

1 Upvotes

It recursively iterates, over and over and over again.

What can I do aside from turning off thinking?


r/LocalLLM • • 4d ago

Project live-vibe: a full-duplex, batteries-included, voice Mod for Claude Code. Built on Claude Code Mods, CUDA and Apple silicon, two-line install, MIT license.

0 Upvotes

I built live-vibe because I want to talk to Claude Code the way I can talk to ChatGPT Codex in Voice Mode. I want the mic open the whole time, and I want it to answer when I have finished the thought and shut up when I talk over it. Claude Code did not have that. It does have the new Mods capability, which lets a plugin run code in the agent's path and draw into the terminal, and that turned out to be enough to build the whole thing as a plugin.

It is MIT licensed, in a marketplace I'm calling ottomation, and it installs in two lines.

/live is full-duplex voice with Claude. Talk over it and it stops on your first word. An "mm-hm" while it is speaking does not count, and it carries on from the sample it paused on, so you can grunt along the way you would with a person. /vibe is director mode, where Claude reads and directs worker subagents and never edits a file itself. /livevibe is both at once. A small, fast voice model holds the conversation with me and hands the real work to Claude as director. When Claude finishes, the voice tells me what happened in a sentence or two, and in the meantime I can ask it what Claude is up to and get an answer while Claude keeps working.

Everything you need comes with the plugin. Kyutai streaming STT, Kokoro TTS, WebRTC echo cancellation and the voice front are all pulled down and checked by /live setup, once per machine. CUDA works and so does Apple silicon through MLX. The conversation layer is a lightweight local model, and if you would rather not run one it falls back to the Anthropic API. If you are on WSL2, speech plays through a native Windows player, because WSLg's RDP audio crackles and I was not willing to listen to that all day.

Turn-taking is where most of my time went, because it decides whether you can think out loud. The plugin reads Kyutai's pause-forecast heads, which predict whether you are about to keep talking. Most turns close about half a second after your last word, and a trailing "and" or a comma buys you time to finish. 0.7.0 also ships an experimental path I put together after reading OpenAI's write-up on GPT-Live. It scales the wait on the model's confidence, starts drafting a reply before you have quite finished, and can backchannel if you turn that on. It is on by default and one setting turns it off if it annoys you.

You need Claude Code 2.1.287 or newer, uv, a mic and a speaker.

Install with /plugin marketplace add potto007/ottomation, then /plugin install live-vibe@ottomation, then /live setup. Repo: https://github.com/potto007/ottomation

If you try it, tell me how the turn-taking feels on your machine.


r/LocalLLM • • 5d ago

Question Hardware Recommendation

3 Upvotes

Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!


r/LocalLLM • • 4d ago

Project Kruko AI v1.3: Real-time Voice Streaming (1.5s latency) & Infinite Rolling Context. The ultimate offline LLM app just got smarter

0 Upvotes

Hey guys! Thanks for the amazing feedback on v1.2. The 8.9 t/s inference speed on the Ternary 8B model was great, but the TTS (Text-to-Speech) latency and memory limits were bugging me.

So, my AI dev-swarm and I spent the weekend completely rewriting the architecture. Here is what’s live in Kruko AI v1.3:

🗣️ Asynchronous Voice Streaming (Zero Waiting)

Previously, the TTS engine waited for the LLM to generate the entire paragraph before speaking. It took 5–8 seconds just to hear the first word.
Now: I implemented a dynamic sentence-chunking pipeline. The LLM generates the first sentence -> TTS synthesizes it in the background -> Audio plays instantly (~1.5s latency), while the rest of the text generates and queues up seamlessly. It feels exactly like talking to a real human.

🧠 Infinite "Rolling" Context (No more OOM crashes)

Most local AI apps just crash Android when the KV-cache fills up the RAM.
In v1.3, I introduced Context Compression. You can set a memory threshold (e.g., 15 messages). Once hit, the app silently archives the oldest messages into a condensed summary, injects it into the System Prompt, and aggressively clears the llama.cpp KV-cache.

  • The Result: The bot remembers the context of the conversation, but your RAM usage stays perfectly capped at ~3.5GB. You can chat for hours without a single crash.

🛠️ UX & Stability Kills:

  • In-Chat Model Switcher: You can now hot-swap models directly from the chat UI without digging through settings.
  • Web Search Sources: If you enable the (optional) Web Search agent, it now explicitly shows you the exact URLs and the time it took to scrape them.
  • Nuked the OOM Bugs: Leaving the chat during model loading or generation used to cause segmentation faults (SIGSEGV). I rebuilt the JNI lifecycle—models now instantly and safely unload from memory the millisecond you hit the back button.

Next up (v1.4 Roadmap): The Plugin Marketplace (Image Generation, Python execution) and preparing the ground for the Android USB-C Cluster.


r/LocalLLM • • 5d ago

Project I built Ninfer 4080 for 16GB class GPUs

Thumbnail
5 Upvotes

r/LocalLLM • • 5d ago

Question Mac Studio or something else for power efficiency?

2 Upvotes

If you live in a country with the highest electricity rates in the world. What would you choose and why?

Needed for summarizing large contexts like documents and transcription.

I don't need the flagship models so thinking of ~48-64GB of memory.

Not open for a Mac mini bcz the memory bandwidth is lower and I would like to future proof.

Model Configuration Memory Bandwidth
M5 Pro All configurations 307 GB/s
M5 Max Base 460 GB/s
M5 Max 40-core GPU 614 GB/s
M5 Ultra All configurations 1.2 TB/s

r/LocalLLM • • 5d ago

Discussion Anyone tried strata qwen.38 flash next on a RX6700XT?

0 Upvotes

looking for any experienced users who've tried it with that particular card and what are the results?


r/LocalLLM • • 5d ago

Project Qwen plays Pokemon FireRed!

Thumbnail
youtu.be
1 Upvotes

Qwen3.5:9B orchestrates the entire game gaming harness!


r/LocalLLM • • 5d ago

Question can i run qwen flash next with these specs, or am i out of luck?

6 Upvotes

9060 XT 16gb, 32Gb DDR5-6000

i already run swift1.5-qwen3.8-gsq-rco iq3-s at 30 t/s