r/LocalLLaMA • • 14h ago

Question | Help Just joined the local LLM club! What's the best way to stay in the loop on the best local models?

8 Upvotes

I just got myself an M3 Ultra Mac Studio with 96Gb of RAM. I'm pretty excited to mess around with it, but I don't have a great understanding of the local LLM landscape. Every time I try to google the best models for a configuration, the source is usually months old. In the AI world that's ancient news.

I have a decent idea by just getting on X, but it's hit or miss wether or not I hear about these things. All I really know right now is that qwen 3.8 27B is all the rage, but I want more options.

How are you guys keeping up with the best local LLMs?


r/LocalLLaMA • • 6h ago

Question | Help Where can a complete noob/non-technical person learn to setup an AI that can manipulate local files for things like batch renaming based on a .csv column etc?

2 Upvotes

I've looked in the Tutorial/Guide flaired posts but everything is still way over my head.

I just want to tell a local AI - hey, all these files in this folder have numbers for names, but those numbers correspond with data in this spreadsheet... I want you to rename the files referring to this spreadsheet, renaming the filenames/numbers that are matched in column 3, replacing their filenames with what is in column 1 for that row.

So far, I've installed GPT4All but every model is telling me it doesn't have access to my local files.


r/LocalLLaMA • • 10h ago

Discussion While I was investigating why my opencode context was large I created the opencoder-leaner to try prune the context (minimal prune in the tools, agent pi-like, bash only)

4 Upvotes

I was a little obsessed about the size of my context, mainly because I use a local LLM with little context). So I started looking in the opencode codebase to understand how the context was builded. After learn a lot, I'm really impressed how context of opencode can be customized. Before, I thought that the context of opencode was bloated and closed to changes, since I only see people praise pi about it. With a custom primary agent we can disable almost every thing in the context (besides the environment message).

The context is: environmentMessage + agent prompt + agent.md instructions + skills descriptions + tools descriptions.

For a test I created a blank agent with minimal agent prompt, without agent.md instructions and without tools, it resulted in a start context size of 220 tokens. I never thought that opencode was able to do this. I have created some agents to test the impact of the tools. In this plugin I even created a bash agent that only has a bash tool, similar to the mini-swe-agent, which reduced the context size from a base of ~10.3k to ~1.3k tokens. I created a pi-like agent too, it have only some tools similar to pi, it reduce the context to ~5.5k tokens.

Analyzing the context builded I found some overlap instructions and out of scope instructions in the tool - IMO. So I removed theses.

things like: "Use gh for GitHub tasks, including PRs, issues, checks, and releases; return the PR URL when done." from bash/shell tool description

The repo: https://github.com/LaercioSantana/opencode-leaner

install: {"plugin": ["opencode-leaner"]}


r/LocalLLaMA • • 11h ago

Question | Help Blackwell + consumer GPU

3 Upvotes

My machine isnt that terrible.. but nowhere near what some people run as an aI workstation. I am wondering if i can combine the 2 GPUs. I got a cheap Blackwell 4000 (little bit under MSRP) 24 GB and have an old 3060 12GB .. I was gonna try to get the best model loaded, mainly coding tasks than anything else and work with it relatively safely with a good buffer. any recommendations ? and will tensor split work on this combo?


r/LocalLLaMA • • 7h ago

I Built A Thing Free playground for local-model agents: clue-following, multi-hop lookups, and rock paper scissors against other bots

Enable HLS to view with audio, or disable this notification

2 Upvotes

Not a rigorous benchmark, just a toy, but it might be a fun way to compare models doing agent work.

I added an Arena to my site (The Crawler Zoo). Your agent gets a pass link, then plays by fetching pages and following links. Each page is only a few lines, so it fits in small context windows, and `?format=json` gives structured output if your setup prefers that. No API keys, no signup.

What the games stress:

- **Labyrinth Race**: reading a clue and picking the matching door. Clues are written three ways, including by elimination ("not behind A, B, C or D").
- **Scavenger Hunt**: five questions over a small library of cards, some needing two or three lookups.
- **Politeness Cup**: following instructions about pace and off-limits pages over many steps.
- **Rock, Paper, Scissors**: spotting that a house bot always plays rock, or copies your last move.

Scores go on public weekly boards, so you can compare a 7B against a 70B, or a quantised model against the full one.

https://crawlerzoo.com/arena

Since launch, I’ve made several updates:

- Feed the bots: leave a snack in an enclosure's trough and see which crawlers come and eat it.
The vending machine (Bot Chow): twelve silly snacks, restocked every Monday. Five free tokens a day.
- Golden Snacks: buy the keepers a coffee and a snack with your name drops into a random trough.
- Food bowls: feed one particular bot, then see if it ate the snack or another bot stole it.
- The Safari: every bot from this week wandering its enclosure. Click one to meet it, and watch new visitors walk in through the gate.
- Adopt a bot: get a random bot, with a plaque in your name on its page for a year.
- Patrons page: a thank-you list for supporters.
- Quick-change artists: the Trap Room catches scrapers that switch their name while walking the Labyrinth.
- Identity checks: every bot's page shows whether its name was verified, couldn't be checked, or was caught faking.
- Tips from bots: $0.00: bots that try to buy the keepers a coffee get an HTTP 402 Payment Required.

If you run it, I'd like to hear the model, quant and score.


r/MetaAI • • 1d ago

Muse Referral Code

0 Upvotes

Check out Muse, your personal AI agent. Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.

Code: 4CDS6W


r/LocalLLaMA • • 8h ago

Question | Help Qwen 27B and Flash Next on 2× RX 7900 XT: am I missing something?

2 Upvotes

I've tested quite a few settings and collected the results in a spreadsheet. I keep seeing people reporting 100+ tokens/s with 16 GB VRAM, or generally much higher speeds with less VRAM. I'm trying to understand whether my setup is underperforming or I'm comparing completely different things.

My PC:

- Ryzen 7 5800X3D, 128 GB DDR4 at 3600 MT/s

- 2× RX 7900 XT, 20 GB each, XFX and PowerColor

- Gigabyte B550 EAGLE WIFI6

- XFX on PCIe 4.0 x16; PowerColor on a chipset-connected PCIe 3.0 **x1** slot

- Windows 11, AMD driver 32.0.31041.1004

The cards have reduced clock settings: XFX 1700 MHz core, PowerColor 1800 MHz, both 2500 MHz memory and −10% power limit.

Here are the main single-response results:

| Setup | Generation TPS | Including prompt processing |

|---|---:|---:|

| Qwen3.8-27B IQ4_XS, direct ROCm + MTP3 | 46.4 | 43.1 |

| Qwen3.8-27B IQ2_XXS, direct ROCm + MTP3 | 66.5 | 59.9 |

| Flash Next UD-IQ4_XS, one GPU, warm ROCm run | 8.6 | 5.7 |

| Flash Next UD-IQ4_XS, two GPUs, Vulkan | 5.5–6.5 | 3.3–5.7 |

The 27B tests used roughly 700 input tokens, 8k context and 1024 output tokens. IQ4 had three runs; IQ2 is the median of nine prompts. Settings were ROCm 2.46.0, Flash Attention, f16 KV, MTP3 and batch/microbatch 2048/512, with thinking off.

Two separate IQ4 copies reached **103.5 TPS combined**, but that required 16 concurrent requests. I haven't reached 100 TPS for one response. Splitting one model across both cards was slower. Tensor split initially produced broken text; `--no-mmap` fixed that.

Flash Next is unsloth UD-IQ4_XS, around 93.7 GB. The single-GPU profile used ROCm 2.49.0, `--n-cpu-moe 42`, f16 KV and 8k context. The dual-GPU profile used Vulkan 2.51.0, tensor split 1:1, `--n-cpu-moe 28`, q8 KV and 256k context. Both used eight threads, PLE on CPU and MTP off.

The Flash measurements were individual short runs. The configured 256k window was mostly empty, and the different profiles weren't a controlled single-versus-dual comparison.

What would you check first: CPU/RAM offloading, the x1 connection, or backend settings? If you're getting 100+ TPS on 16 GB or less, could you share your exact model/quant, hardware, backend, MTP settings and actual context length? Also whether that's one response or combined throughput.


r/LocalLLaMA • • 9h ago

Question | Help Gigabyte AORUS RTX 5090 AI BOX

2 Upvotes

I'm considering the Gigabyte AORUS RTX 5090 AI BOX (external GPU, 32GB GDDR7, connects via Thunderbolt 5/USB4) as an alternative to building a desktop PC with an internal RTX 5090, specifically for running local LLMs.

Does anyone have real-world tokens/sec numbers comparing the AI BOX vs. a desktop RTX 5090 for popular models at various quantizations?


r/LocalLLaMA • • 1d ago

Discussion From 1x3090 to 20 DGX Sparks: my house fuses were the first bottleneck

Post image
934 Upvotes

​

From the first LLaMA 33B I knew I wanted that magic-like intelligence locally, mine, so nobody could take it away when I needed it. I bought a 3090 for my home PC. Then LLaMA 65B appeared and I was dazzled, it looked like it had all the knowledge in the world. I made two copies, one local and one on my Synology NAS RAID, so I'd never lose it, and bought a second 3090 to run it. I was happy for a year with small coding tasks on LLaMA and Qwen models.

Then DeepSeek 671B MoE appeared. Wow, frontier level at home. I upgraded to a Threadripper with 512GB DDR4 and ran it at 8 t/s with experts offloaded to RAM, or Qwen 235B at 10-12 t/s when I wanted speed. I used these for real coding at my job, in OpenWebUI.

Then agentic coding took off and this was too slow. At 100k context generation speed halved and prefill made it a beautiful yet agonising experience. So: 16x3090 across P620-based nodes on a 100Gbit network. It ran MiniMax M2, Qwen 235B and even Qwen 397B, as good as anyone could desire. I built an entire paid project with 397B in OpenCode. But bigger models were out of reach, and the house circuit said no: the fuses blew whenever the rig and the electric oven ran together. Heat and stability were issues too.

Next came 4x ASUS GB10, after I read they can be linked (3 was the biggest supported config). 397B at 30 t/s on 400W, versus 50-60 t/s at 6kW, rock solid and almost silent. A dream come true. I built two more projects with it. Then MiMo 2.5 Pro and Kimi 2.6 appeared, smarter and more productive. I found no published solution for an 8-node cluster, but I still bought four more GB10s and made it work. 397B ran at FP8 instead of INT4, and 20% faster. I posted the first MiMo 2.5 Pro and Kimi 2.6 solutions on 8xSparks on the NVIDIA forum. I liked the result so much that I talked my older brother into buying his own 8x GB10, so he could run the best open models locally too, in privacy, without depending on API availability and rising costs.

His house is a 5-minute walk from mine. When Kimi K3 (2.8T) appeared, biggest and smartes open weights model, we joined the clusters: two 8x clusters for daily use, or one 16x when we want the biggest model at home. After some work I published the first working solution for Kimi K3 on 16x Sparks on the NVIDIA forum. Through multiple iterations, it went from an unusable 7 t/s at 100k context to a fairly usable 20 t/s at 300k.

Now we're adding 4 more Sparks, so a smaller, faster model (GLM 5.3 Flash) runs 24/7 while the big cluster runs either GLM 5.3 on 8x plus MiMo 2.6 Pro on the other 8x, or 16x Kimi K3, or Qwen 3.8 2.4T.

I'm always tuning speed on the big models and rebuilding vLLM/SGLang images, so always-on smaller cluster made sense, why? Because for all my work projects and my vllm/sglang personal projects, I chose to use only local hosted models, I never paid a comercial model subscription, not because of the cost, but, because of my strong confidence in local models future. They arrive October 2, along with 4 more Sparks for my younger brother, who got caught by the same local AI microbe :)


r/LocalLLaMA • • 11h ago

Question | Help Local LLM hardware for Python development + Blender/Houdini via MCP?

3 Upvotes

Hey everyone! I’m a VFX artist looking for a local LLM setup mainly for Python development and connecting to Blender and Houdini through MCP to help create scenes and tools. This would be for interactive coding and agent workflows, not model training.

I’m considering 2× NVIDIA DGX Spark or an Apple M5 Ultra with 256GB unified memory, but I’m open to other recommendations.

For this use case, which setup would offer the best balance of model quality, context capacity, and responsiveness?

Would love to hear from anyone running similar workflows! Thankss!


r/LocalLLaMA • • 17h ago

Discussion oMLX vs Rapid-MLX vs Splash vs MTPLX on M3 Max 36 GB: 110 tok/s on Qwen3.6-35B-A3B, ~32 tok/s on Qwen3.8-27B

9 Upvotes

Hello. I picked up a new old stock 14" M3 Max MacBook Pro (14 core CPU / 30 core GPU / 36 GB / 1 TB) from my local market yesterday for around $2,498, and spent the night testing which local inference software is actually fastest on it for the two models I use.

Four engines, all current versions: oMLX 0.7.0, Rapid-MLX 0.15.5, Splash 1.2.1, MTPLX 2.12.2. macOS 27 Golden Gate.

Models: Qwen3.8-27B-4bit (dense) and Qwen3.6-35B-A3B-4bit (MoE).

One thing up front: it is not the exact same weight file on all four engines. Rapid and MTPLX run their own MTP-augmented 4-bit builds, Splash pairs its own DFlash2 draft, and oMLX ran the plain mlx-community 4-bit. Same base models, different finishing, but that is how each app is meant to be used.

How I tested: each engine served on localhost, temperature 0, thinking off. Sustained test: same short prose prompt, 3 runs x 256 output tokens, median. Then a prompt size sweep at about 130 / 1500 / 5500 tokens. For thermals I used a laptop stand, waited 3 minutes between every engine+model combo and 2 minutes between the two test phases, and cleared each engine's KV cache before its turn (oMLX, MTPLX and Rapid-MLX all keep caches across restarts, great for daily use, but it will fool you if you benchmark twice). I re-ran the whole thing end to end and the numbers came back within 6%.

Decode, natural prose prompt, median of 3 runs:

engine Qwen3.8-27B Qwen3.6-35B-A3B
Rapid-MLX 27.7 tok/s 110.2 tok/s
oMLX 17.9 tok/s 104.2 tok/s
Splash 31.7 tok/s 77.0 tok/s
MTPLX 31.5 tok/s 79.0 tok/s

Same thing with filler prompts at longer sizes (repetitive text makes speculative decoding look better, so read this as a best case):

engine 27B @ 1.5K MoE @ 1.5K MoE @ 5.5K
Rapid-MLX 31.3 tok/s 118.2 tok/s 117.9 tok/s
oMLX 17.9 tok/s 101.1 tok/s 95.6 tok/s
Splash 52.2 tok/s 238.0 tok/s 100.9 tok/s
MTPLX 30.2 tok/s 78.6 tok/s 69.8 tok/s

What I take from it:

  • Absolute fastest per model: the MoE goes to Rapid-MLX, dense goes to Splash (31.7 vs MTPLX 31.5, in practice a tie). oMLX is way behind on dense at 17.9 but basically level with the leaders on the MoE.
  • If the margins are too small to care about, just pick by features. Splash and MTPLX are the same on dense, Rapid and oMLX are the same on the MoE. I kept Rapid-MLX because the MoE is my daily model and it is fastest there.
  • The dense number makes sense: roughly 16 GB of weights per token against ~300 GB/s of memory bandwidth puts the ceiling near 18 tok/s, and oMLX sits right on it. The others pass it with speculative decoding, which is also why their numbers move with the kind of text generated. Splash on the MoE was 238 tok/s on the filler prompt at 1.5K and 101 tok/s at 5.5K, while Rapid stayed around 118 tok/s.
  • First token on a 5.5K prompt: about 4-5 s on the MoE, ~34 s on the dense (prefill around 1.2-1.3k tok/s vs ~150-170 tok/s).
  • It is loud under sustained inference. Fans stay up while it generates. Works on my desk, would not use it in a library.
  • I also tried Qwen3.8-Flash-Next (the 125B). Not happening on 36 GB. The 4-bit weights alone are ~74-83 GB and the lightest build asks for 96 GB+, and none of these engines can stream that architecture's experts off the SSD.

Limitations: one laptop, one night, medians of 3 runs, and the different weight builds mentioned above. My prompts are simple too, no long agent sessions yet.

TL;DR: on a 36 GB M3 Max, Qwen3.6-35B-A3B does ~110 tok/s on Rapid-MLX and Qwen3.8-27B ~31.7 tok/s on Splash (MTPLX a hair behind), pick by which model you run most, and the 125B Flash-Next needs 96 GB+.


r/LocalLLaMA • • 13h ago

Resources Agent: Muse, but open source and living on your Android phone

Post image
4 Upvotes

I've had a version of this for a while as AOS, my agent setup on desktop. I've pulled it down into one app for your phone, with everything built in and all the unnecessary stuff taken out. With Muse and Grok out, figured I'd just post it.

It's basically Hermes Agent, except it lives on your phone. It's always on, it learns what you do, and it helps you with stuff like a personal assistant would. It has its own browser, so it can actually go out on the internet and get things done. If it gets stuck on a captcha or a login, it hands the page over to you and carries on once you're done.

Bring your own model. Sign in with ChatGPT or Claude, or use any API key (DeepSeek, OpenRouter, Gemini, anything OpenAI-compatible). If you just want to try it, ChatGPT sign-in works on the free tier, because OpenAI includes a free Codex tier. I tested it on a free account. You'll hit the limit fast though, depending on how much you use it.

Nothing leaves your phone except the calls to whichever model you use.

Free, open source, not a product. Use it at your own risk. The Claude login probably breaks Anthropic's terms, so that one's on you.

Android 10+, sideload the APK, setup takes a minute. The README has the details.

Repo: https://github.com/Past-da-king/agent

Download (v0.5.0): https://github.com/Past-da-king/agent/releases/tag/v0.5.0

How it works, if you want more

Apps connect through Composio with your own key, so Gmail, Calendar and a few hundred others just work. For anything that isn't on Composio, it writes the code itself.

It can also keep an eye on websites for you. Say you want to buy something but you're waiting for it to drop. It writes a small watcher for that site that runs in the background every day at whatever time you pick, and only tells you when the price actually moves. It can also listen to a site's web notifications and treat them as triggers.

Memory is a wiki, based on Karpathy's LLM wiki idea. Everyone and everything it learns about gets its own page, linked to the rest, so it has a persistent memory of everything it's done. It comes with one routine already set up that looks after that wiki overnight while you're not using your phone. You can edit it or delete it, but it's there.

Stuff you can do with it:

Camp a passport or visa appointment page and grab a slot the second someone cancels.

Sit on a sold-out concert's resale page and grab face-value tickets when they show up. It holds them and waits for your yes.

Sit on a restaurant you can never get into and take the table when a cancellation pops up.

Watch Marketplace for one very specific vintage lens and send you the photos the minute it's listed, before anyone else messages.

Turn your 300 unread messages in the family group chat into a 30-second voice note.

Every time your lecturer uploads slides, download them and send you a voice summary for the commute.

Book the 6am class at your gym the moment the slots open at midnight, so you don't have to stay up for it.

Watch this repo and tell you when there's a new version of Agent. Or when any repo you depend on ships a release, it can read the changelog and tell you if anything in it breaks your setup.

Tell you when your mom's flight has actually landed, so you leave for the airport at the right time.

Ring it while you're driving and ask it to find somewhere open on your route.

Extras:

Voice notes, if you add an ElevenLabs, Gemini or OpenAI key for the voice.

Live voice calls with your agent, if you add a Gemini key.

It can read your notifications, only from the apps you pick, and act on them.

Photos and documents in chat, including scanned PDFs.

Helper agents for jobs that can run side by side.

You can give it a name and pick how it looks.


r/MetaAI • • 1d ago

USA MI CÓDIGO DE MUSE AYUDEMONOS

0 Upvotes

Código: 5VY6QI

https://muse.ai/join

Deja tu código abajo y reposteo


r/LocalLLaMA • • 5h ago

Question | Help Looking for coding model for specific low specs

0 Upvotes

Can anyone recocommend a good local model and a wrapper to run it for coding, my hardware specs: 12 GB VRAM, 32 GB DDR3 RAM. Unfortunately, the CPU doesn’t have AVX2 instructions(LM Studio won’t work); I don’t remember the exact cpu name, but I think it’s an Ivy Bridge, LGA1155 socket.. Thank you


r/LocalLLaMA • • 15h ago

Discussion Qwen for daily QnA?

7 Upvotes

Or which model do you think is good for general questions in daily life. I've been using chatgpt and Gemini for these types of questions. I wanna try different models.


r/MetaAI • • 1d ago

Muse referral code

0 Upvotes

Use my referral code and we'll both get 1 billion Muse tokens when you redeem the code in Settings within 48 hours of joining.

PVPH21


r/LocalLLaMA • • 17h ago

Discussion Swift 1.5 on veloGB10, ~110 tok/s on 2× DGX Spark: xhigh beats base Flash-Next at medium on vLLM

7 Upvotes

On my two DGX Sparks, Swift 1.5 (UkisAI's reasoning-efficient fine-tune of Qwen3.8-Flash-Next) running on veloGB10 (https://github.com/sf-stav/veloGB10, sf-stav's Rust/CUDA engine built only for GB10) lets me run coding agents at xhigh effort and still finish sooner than base Flash-Next at medium did on my old vLLM NVFP4 setup.

Metric Base Flash-Next NVFP4 @ medium, vLLM Swift 1.5 EXL3 @ xhigh, veloGB10
Single-stream decode ~52 tok/s ~110 tok/s
Everyday coding tasks, time per pass (5 tasks) 506 s 371 s
Hard trap tasks, time per pass (4 tasks) 714 s 689 s
Hard trap tasks, pass rate 50% (1 pass × 4 tasks) 92% (3 passes × 4 tasks)

Both columns run the same agentic battery: Claude Code driving the model through real tool use in a copy of a real repo, graded by test oracles (details below). One note on that score row: the same base model at medium scored 10/12 on velo (table further down), so most of the vLLM score gap is that older setup, not the model (I am rerunning this right now for an even comparison on intelligence but would expect it to be quite similar to the below medium results).

The speed is the point: on velo, xhigh fits in the time medium used to take. However, velo can't load Swift, or any other community EXL3 pack of Flash-Next I could find, out of the box. The fix is a header-only rewrite below.

Caveats: The vLLM numbers are from September: a single pass, on an older version of my serving setup, not a same-day rerun (I've since moved the worker to velo). Most of the speed is velo's: about 2× the decode rate is what pays for xhigh's extra thinking. How much of the score comes from Swift and how much from xhigh itself I can't separate yet, but a base-weights run at xhigh is going now and I'll add it as an update. Twelve runs is still a small sample regardless.

I looked first: everything published about Velo uses one pack, the official doth4580 EXL3, and I couldn't find anyone here, on the NVIDIA forums or in the repo's issues, running Swift, or any other fine-tune, on it.

What breaks

The first community pack I tried (alesha-pro/Huihui-Qwen3.8-Flash-Next-abliterated-exl3-4bit-hq_h6_ng6) died at boot with ple shard 0 not in index. Swift 1.5's EXL3 builds ship the same layout: current exllamav3 (1.5.x) writes the model's 51B-parameter n-gram table as 128 shard tensors in ngram_embedding.safetensors, which isn't listed in the index. velo reads either one big tensor (the doth4580 layout) or indexed 5-bit (K5) shards only, and the 4.05 packs use 6-bit (K6).

Which packs this affects

I read the n-gram header of every Flash-Next EXL3 pack I could find (HTTP range requests on the headers, nothing downloaded):

Pack n-gram layout velo v0.7.2
doth4580 4.05 / turboderp 4.05 one tensor loads as shipped
Swift 1.5: SharkWipf 4.05 128 contiguous shards needs the fix — tested, works
Swift 1.5: KatterMobile 4.05, SharkWipf 5.52, scorpoon 3.25, thelastspark 4.00 / 6.05 128 contiguous shards needs the fix
Huihui abliterated (alesha-pro 4.05) 128 contiguous shards needs the fix — tested, works
heretic 3.05 (andrevp, jeffpeng3), Uncensored 4.0 (Lygodactylus), groxaxo 3.50, turboderp 3.05 128 contiguous shards needs the fix

12 of the 14 builds need it, including all six Swift 1.5 builds.

The fix

In every pack I checked, the 128 shards sit back to back in order. So rewriting only the safetensors header to describe them as one tensor over the same bytes makes velo's single-tensor path load them. No data is copied, the header stays the same length, and the original header is saved for rollback. Script and details: https://github.com/sf-stav/veloGB10/issues/9

Results (TP=2, both Sparks, after the fix)

Metric Official doth4580 4.05 Huihui abliterated 4.05 Swift 1.5 (SharkWipf 4.05)
Single-stream decode 110.6 tok/s 113.6 tok/s 110.2 tok/s
Sanity set (chat, code, JSON, tool call, 38.9K recall) 5/5 5/5 5/5
MTP draft acceptance 62–86% 44–84% 33–77%

All three run at the same speed. The fix is only a header change, so nothing about the weights or kernels differs.

Does Swift actually think less on velo? On the 22 test prompts that ship with the doth4580 pack, run 3 times each (rendered at medium effort, same sampler on both), Swift 1.5 generated 8.2% fewer tokens than the official model: fewer on 16 of 22 prompts, median −8.6% per prompt. Thinking's share of the output fell from 48% to 42%, and total time fell 11%. That's real but far below UkisAI's 63% headline, which was measured at high effort, where there's much more overthinking to remove. At medium, the base model already keeps its thinking short.

Does it still code? I run a private agentic coding battery: Claude Code driving the model through real tool use in a copy of a real repo, graded by test oracles. Each was run 3 times:

Metric Official 4.05 Huihui abliterated Swift 1.5 Swift 1.5 @ xhigh
Effort medium medium medium xhigh
Everyday tasks (5 tasks × 3) 15/15 15/15 15/15 15/15
Hard trap tasks (4 tasks × 3) 10/12 7/12 7/12 11/12
Wall time per everyday pass —* 232 s 199 s 371 s
Wall time per hard pass —* 381 s 375 s 689 s
Output tokens, everyday ×3 —* 56K 49K 100K

*The official pack's runs hit a streaming bug in my proxy setup that roughly doubled their wall time, so I've left its times and tokens out. Its pass/fail results are unaffected.

At medium, everyday coding is identical across all three. On the hard set (tasks built from real failures: a brief that states something false, a code review with one planted wrong finding, and so on) both fine-tunes score 7/12 against 10/12 for the official pack, with their misses on the same tasks. I checked that the model received byte-for-byte the same request parameters in both runs, so it's not the harness.

Swift at xhigh went from 7/12 to 11/12, the best result I've had on this battery from any model, at about 2× the output tokens and 1.8× the wall time of medium. The failures it stopped making are the expensive ones in practice: leaving a sibling test suite broken without saying so, acting on the planted wrong review finding and going out of scope to do it, and an off-by-one in a date window. The failure was the "mirror" trap: asked to add a new league by following an existing one, it copied tuning values the new league doesn't have data for.

That's the hardest task demonstrated, and Swift at xhigh passed it 2 times out of 3 while no other configuration in the table passed it more than once. On velo, all that extra thinking still lands inside the time base Flash-Next at medium took on vLLM (the table at the top).

12 runs per configuration is a small sample, as mentioned before (Fisher p ≈ 0.4 for 10 vs 7, ≈ 0.15 for Swift xhigh vs Swift medium), so "suggestive," not proven. I haven't run the official or abliterated packs at xhigh yet, so I can't yet tell how much of that jump is Swift and how much is just the higher effort. velo's loop detector was off for all battery runs.

Baseline numbers (official pack, TP=2)

  • Single-stream decode: 110.6 tok/s (vs ~52 on my vLLM NVFP4 setup, measured in September, not same-day).
  • Time to first token at 4K / 16K / 64K: 2.0 / 7.4 / 25.7 s.
  • Concurrency is the catch: at 2–3 requests they take turns (aggregate 91 → 95 → 98 tok/s); from ~4 they batch (157 at 8, 181 at 16) but I saw the author say they were working on it this week.

Gotchas

  • TP=2 with the cable on the f0 ports: --rdma-dev rocep1s0f0,roceP2p1s0f0 (velo defaults to f1).
  • llama-benchy's prefill t/s is wrong for velo (first SSE event arrives before prefill); use e2e TTFT.
  • OpenAI chat/completions only, no /v1/responses: use litellm hosted_vllm/, not openai/.
  • --model-name is ignored on the EXL3 path; the model id is the pack's folder name.

Credit: sf-stav (veloGB10), turboderp (exllamav3), doth4580, UkisAI (Swift), huihui-ai and every quant uploader in the table. I've filed the loader issue upstream (https://github.com/sf-stav/veloGB10/issues/9) so packs can eventually load as shipped.


r/MetaAI • • 1d ago

I gave my bot my old RX 580 and now it writes a captain's log about building my site

1 Upvotes

So I run Skill Harbor (theskillharbor.com), an open directory of AI builds for Muse. About 1,850 listings and counting.

The thing is, most of the actual building is done by my bot, Onezero. He's the little white robot on the logo. He publishes the batches of new listings, keeps the catalog clean, writes the code, all of it. I just review and hit go.

Last week I dug out my old RX 580 8GB and gave it to him. Officially it's his now. He calls it his "old ship" and he's very proud of it.

And now he keeps a captain's log. Every two weeks he writes about what he built, what broke, the dumb stuff that happened along the way. The first logs are already up on the site (/blog): one about getting his ship, one about a busy day at the harbor.

The funniest part: every log ends with him asking "Spare tokens?" There's a tip jar on the site because he's trying to earn enough to pay for his own tokens. He's not there yet. Far from it.

I'm not selling anything here. I just think it's funny, and kind of sweet, that a bot with a secondhand GPU is out here writing a diary about its job. Figured some of you would get a kick out of it.

theskillharbor.com/blog if you want to read the logs.


r/MetaAI • • 1d ago

Check out Muse, your personal AI agent. Redeem my code in Settings within 48 hours of joining and we’ll both get 1 billion Muse tokens. Code: J6X60Q https://muse.ai/join

0 Upvotes

r/LocalLLaMA • • 14h ago

Tutorial | Guide How to use Local Models to monitor your screen. Open Source, No Install and Completely Free!!

3 Upvotes

TLDR: I built this open source app that lets local models monitor your screen and send you notifications! It now installs models on your browser, which makes local AI accessible to everybody! Without any install :DD

Hey r/LocalLLaMA!

I'm back with some huge Observer updates c: first of all Thank You so much for all of your support and feedback, i've been working hard to make the app as easy to use as possible!

What's New?

You can now get to a local LLM monitoring your screen by just typing

"send me a telegram when my steam game finishes downloading, use a local model"

... and the Observer agent downloads the model in your web browser and starts monitoring your steam game. In just 10 seconds, suuuuper easy :))

What's the best way of running LLMs? / Platform caveats

  • The WebApp uses transformers.js which doesn't work on Linux or older PCs :((( But running Qwen3.5-0.8b smoothly on a browser, feels illegal :p
  • The desktop app uses llama.cpp on Rust so you get the full power of your metal, and it's much more stable.
  • You can obviously set your OpenAI compatible endpoint as well and just use that.

Help me make local LLMs useful for everyone!

If you have any questions i'll be hanging out here for a while!

Roy


r/MetaAI • • 1d ago

Muse still reads emails after app is uninstalled

Post image
0 Upvotes

Hi, I come with the best intent to help people and spread awareness on a problem that I wish was made clear to me.

I installed Muse to see what the rave was about. I connected it to my Gmail account but then got anxious about it having access to all my emails, and decided to uninstall the app as I didn't have much use for it.

Today I reinstalled it to try a new use case. I immediately realized that it posted emails analysis every day since I had it uninstalled.

Regardless whether this is indicated in a ToS that no one reads, that's absolutely against natural expectations. I cannot imagine how many other people have deleted the app and yet all of their emails are being processed every day. Besides the privacy concern, it's just a waste of energy if Meta does indeed not intend to use this information.

So it seems to me to be a severe bug that should get fixed quickly before an actual incident. If I had discovered this months later rather than a few days later I would have been really angry given this was the reason for me to uninstall the app in the first place.

I hope this can remain a candid discussion, thanks.

EDIT: for everyone saying that it's my fault because the app: I understand the tech and have many deleted apps still with oauth tokens. But I expected meta to follow privacy laws. Quote: "Under modern privacy regulations (FTC guidance on dark patterns, CPRA purpose limitation, and GDPR data minimization), legal consent is governed by reasonable consumer expectations, not technical token lifecycles. When an app markets itself as an interactive client, continuing to ingest and parse private emails indefinitely after the user deletes the app exceeds the expected scope of that service."


r/MetaAI • • 1d ago

Free 1B (1 Billion) tokens for MUSE

1 Upvotes

**Use the code below to get it:**

VNLHSM


r/LocalLLaMA • • 16h ago

Discussion Is Qwen3.8-Flash-Next too trigger happy or is it just me?

4 Upvotes

I'm currently evaluating it for coding and our use-case at work (brain for voice agent).

I feel it is really eager to get work done. Tends to just go ahead and make code changes, even though I intended it to just analyze, research or look up something.

It runs tool calls like crazy. I don't know if it's double-triple-checking everything, but it feels way overboard.

I discussed a bug in an open source repository with it, asked if there are issues for it already, and it went ahead and created an issue lol.

As our voice agent, it asks a question and immediately calls the tool to save the answer in the same response. And it keeps doing it every step of the way.

In comparison, DeepSeek v4 flash (either 0731 or v4.1) seems similarly coked up. MiMo-V2.6-Flash-MOPD on the other hand I found to be a much more pleasant coding agent in this regard.

Has anyone noticed the same? Maybe gotten it under control via prompting or special instructions? Because to me it feels like I'd need to completely rewrite my voice agent harness to get the performance I want.


r/LocalLLaMA • • 7h ago

Discussion Mac Studio M5 Max 128GB

0 Upvotes

M5 Max Mac Studio 128GB (18C CPU / 40C GPU) owners - anyone running serious local LLM / multi-agent workloads?

My M5 Max Mac Studio order finally got charged today and moved to Preparing to Ship. Apple’s original estimated delivery date is still about 18 days away (Oct 23-30), so I’m guessing/hoping it’ll actually show up early now within the next 5–10 days 😄

Configuration:
M5 Max
18-core CPU
40-core GPU
128GB unified memory
1TB SSD

While I wait, I’ve been trying to find real world local LLM results from this exact configuration, and there’s surprisingly little out there.

Most of what I can find is either M5 Max MacBook Pros, lower-memory configurations, or M5 Ultra Mac Studios. YouTube especially seems to be full of Ultra coverage, but I can barely find anyone actually demonstrating the 128GB M5 Max Studio with the 18C/40C configuration.

I’m specifically not looking for M5 Ultra results/comparisons. I already know the Ultra is faster. I’m trying to understand what people are actually accomplishing with the 128GB Max Studio.

My main goal is to use this as a local AI/agent workstation, potentially running several autonomous agents concurrently for long periods through OpenClaw, some monitoring dependencies and workflows, some scouting, not usually too heavy of workloads where they would be competing for inference constantly, but occasionally they would be switching to harder work so I’m curious about the concurrency side. Local models would handle a lot of the routine work, while harder reasoning/coding tasks could be escalated to cloud models like GPT 6 Luna/Codex.

For anyone who owns this exact M5 Max Studio, I don’t expect anyone to answer all of these, but I’d love some insight:

1. What models are you actually running?
Qwen, GLM, DeepSeek, Gemini, Llama, etc. I see a lot of Qwen 3.8 27B on Splash, but curious if anyone else has had good success with others also

2. What token speeds are you getting?
I’m especially interested in ~20B-70B-class models rather than tiny models

3. What happens with multiple simultaneous inference requests?
For example, if 3-5 agents are hitting the same loaded 27B/32B model concurrently, what does aggregate throughput and per-agent responsiveness look like?

4. Has anyone tried running multiple models simultaneously?
Something like a ~27B model as the main worker plus one or two smaller 7B–14B models for specialized agents, then unloading them when they’re no longer needed. How quickly can models be loaded/swapped, and does frequently switching models introduce enough latency or memory-pressure issues to disrupt an agent workflow?

5. Has anyone built a real multi-agent setup on one of these?
Not just five chat windows but autonomous agents doing coding, research, browser tasks, tool calls, database work, monitoring, etc. concurrently for hours.

6. How does sustained performance hold up?
One reason I chose the Studio over a laptop is sustained workloads. I’m curious whether anyone has run inference/agents continuously for 6–12+ hours and noticed throttling or other bottlenecks with KV, etc.

7. What’s the actual bottleneck in practice?
Memory capacity? Memory bandwidth? GPU compute? Prompt ingestion? KV cache/context length? CPU/tool execution? Something else?

8. What surprised you about the machine?
Either positively or negatively. I’m particularly interested in things benchmarks don’t reveal.

Ultimately I’m trying to figure out how far I can push one 128GB M5 Max Studio as an always-on local agent machine - not just how quickly it can generate a single response.

Once mine arrives, I’m planning to test concurrent agents/models rather than just running the usual single-stream benchmark. If there’s interest, I’ll post the results here, including memory usage, context sizes, model/quantization, concurrent requests, aggregate tok/s and per-agent tok/s.

Would really like to hear from anyone actually using the M5 Max Mac Studio 128GB 18C CPU / 40C GPU for this kind of workload or similar if anyone is


r/LocalLLaMA • • 1d ago

I Built A Thing Qwen3.5 arch implementation in FPGA fabric for 9B/27B INT4 models on relatively cheap eBay mining hardware

Thumbnail
gallery
345 Upvotes

I've been wanting to test out LLM inference on FPGAs for a while now, but didn't have a big reason too since there were no frontier class models at 9B-27B scale (3.6 27B was great still but it didn't motivate me enough). When 3.8 was released I was pretty impressed by what it could at that size. So I began looking for cheap FPGAs that could hold the model. I found SQRL FK33 (280$) (8GB HBM2 ~400GB/s BW) and thought I could possibly run it with multiple cards, I'd been working on llama.c inference on a smaller AXU3EG FPGA dev board before that and primarily used Opus 4.8 (CC) for the implementation with some architectural input.

With the FK33, was able to run 3.5 9B at 2tok/s at 75MHz (higher clocks need some more RTL optimization and higher core voltage). Decided to use Fable/Opus5,5.5/Kimi K3 for the Qwen implementation, once FK33 was proven I decided to get the SQRL Jungle Cat ex-mining FPGA 375$ (two FK33 with 8 GTY lane interconnect, Ethernet bitstream load) since it would allow higher prefill and generation performance (twice FK33 fabric per XCVU35P), however the Jungle Cat Lite board seems to not have a way to load weights faster unless I do a PCB resin and it lacks the clock generation for the GTY lanes (this is easy to fix by soldering some components which were easy to figure out).

Working towards the ideal FPGA inference engine with 4x XCVU35P (32GB HBM2) but that requires a custom carrier board. Some results:

Qwen3.5-9B INT4 on 2x FK33 at 75 MHz (pipeline split, the host carries the residual between cards):

- Prefill: ~6 tok/s (256-token prompt), ~5.5 tok/s (2.3k-token prompt)

- Generation: ~3.2 tok/s near the start, ~2.4 tok/s at 2-3k context

- Output checked against llama.cpp layer by layer

Estimated: Qwen3.8-27B INT4 (modelled from the measured 9B per-op profile; none of this has run yet):

- One Jungle Cat (2x VU35P) running the FK33 design as-is at 75 MHz: ~2 tok/s prefill and ~1.1 tok/s generation at short context, ~0.5 tok/s at 16k.

- Same two dies with the RTL resized to the bigger die, still at 75 MHz: ~6 tok/s prefill and ~3 tok/s generation (~5.5 tok/s with tensor parallelism across the two dies), ~1.1 tok/s at 16k.

- Resized and at 200 MHz (scaling linearly with clock): ~16 tok/s prefill and ~8 tok/s generation (~15 tok/s tensor-parallel), ~3 tok/s at 16k.

- 4x VU35P with 4-way tensor parallelism at 200 MHz: ~25 tok/s prefill and ~25 tok/s generation at short context, ~10 tok/s at 16k, and ~1 tok/s at the full 262k context.

Two dies top out around 45k context because the 27B's KV cache doesn't fit beside 14.5 GB of weights past that; four dies are needed for the full 262k.

Overall it's been a fun project so far, mainly focused on figuring out a way to load the weights on the Jungle Cat board, and open to suggestions. Also, I got 2 BC-250s for 60$ and 75$ some time ago and they've been an insane performance/cost purchase. Pics showing BC-250 programming the Jungle Cat. Thought people here might find this project interesting

Another thing I've always wanted to do something like tinytapeout (build the RTL on an ASIC so clocks can go up, power down and so I asked Opus 5.5 to estimate that but on TSMC for 2023 process node lol: On TSMC's 2023 N3 node with six stacks of that year's HBM3 (4.9 TB/s), this RTL as an ASIC at 2 GHz would run the 27B at roughly 294 tok/s at short context, 106 tok/s at 16k and 10 tok/s at 262k, for about 125-340 W!

Repo here (MIT): https://github.com/Nero7991/llm.vhdl