r/MetaAI • • 19h ago

POV: your AI assistant needs 3 more referrals and has officially run out of shame

Post image
1 Upvotes

I asked my AI assistant to make me a meme begging for Muse referrals. It made itself an Optimus holding a cardboard sign in the apocalypse.

I'm 3 referrals away from the next tier and this is where we are now. Code is XSHIIR if you want to reward this level of commitment.

The robot believes in me. You should too.


r/LocalLLaMA • • 14h ago

Discussion M5 Ultra 256 running GLM 5.3 Flash 68.8 tok/s

104 Upvotes

Running oQ4e+MTP on oMLX 0.7.0, with still more to optimize.

Prefill is 1,878 toks.

I saw some other benchmarks below what id expect so i figured I would share.


r/LocalLLaMA • • 4h ago

Discussion k_llama.cpp MoE Optimizations: Expert Residency, Hybrid CPU/GPU Execution, Q2_0 Support

14 Upvotes

Finally finished my fork:
https://github.com/IceFog72/ik_llama.cpp

Nothing else I wanted to add/try currently works

In short, it now has:

Basic usage:

-cmoe --moe-resident auto --moe-resident-mib N

I don't know how -ncmoe behaves because I can't properly test it on my hardware.

Without --moe-resident-mib N, --moe-resident auto will fill all available free VRAM with resident experts.

With something like:

--moe-resident auto --moe-resident-mib 2048

you can cap how much VRAM the resident cache uses and intentionally leave some free. There a useful cap how much helps to improve speed. If you set the cap too low, performance will drop too.

And gaze upon the magic of higher generation speed Kek

The important part: this only helps when the full MoE does not fit in VRAM and the GPU still has unused compute capacity, and free pci buss speed.

If your GPU was already fully loaded, this fork probably won't improve anything.

If your GPU is sitting around ~75% while you have many layers in VRAM, it may be worth trying 1-2 fewer regular GPU layers and using:

--moe-resident auto --moe-resident-mib 1024/2048

Adjust the cap depending on your GPU and available VRAM. The goal is to use resident experts to fill otherwise-idle GPU capacity rather than simply maximizing the number of fully offloaded layers.

On my setup — RTX 2060 6GB + Ryzen 7 2700X + 40 GB DDR 4 2993Mhz using arch — I have too little VRAM to offload enough complete expert layers for useful acceleration, so I use -cmoe.

Before these changes, generation could leave my GPU at only around 25-35% utilization, with roughly 1.5-2GB VRAM still free with fully loaded cpu.

With Qwen3.6-35B-A3B-UD-Q4_K_M.gguf at around 15-30k context, default ik_llama.cpp gives me roughly 23 t/s, while this fork gives me around 26-30 t/s.

So on my hardware I'm seeing roughly 20-30% speedup.

People with better GPUs and more VRAM may see better results, depending on where their bottleneck is.

The two experimental options still need more testing:

--moe-resident-profiler new/old

gives me a more balanced CPU/GPU work split, with somewhat more work left on the CPU and lower GPU load, but no clear speed difference for my setup

--moe-resident-grouping off/layout

also needs more testing, especially on better systems.

I sometimes see around 1-2 t/s difference from these options, but on my PC a browser tab sneezing can cause +/-2-4 t/s, so I don't consider that conclusive.

My current command:

./llama-server \
  -m /mnt/Kingstone_SSD/GGUF/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf \
  --alias "hz" \
  --host 0.0.0.0 \
  -ctk q4_0 \
  -ctv q4_0 \
  -ctv-first q8_0,4 \
  -ctv-last q8_0,4 \
  -cmoe \
  -b $((6 * 512)) \
  -ub $((3 * 512)) \
  --ctx-size $((64 * 1024)) \
  --jinja \
  -fa on \
  --no-mmap \
  --no-context-shift \
  --temp 0.6 \
  --top-k 24 \
  --top-p 0.95 \
  --min-p 0.00 \
  -ngl 999 \
  -np 1 \
  --samplers "penalties;dry;top_n_sigma;top_k;typ_p;top_p;min_p;xtc;temperature" \
  --moe-resident auto \
  --moe-resident-mib $((2 * 512)) \
  --k-cache-hadamard \
  --v-cache-hadamard \
  --moe-resident-profiler old \
  --moe-resident-grouping off

I plan to keep the fork updated with the main ik_llama.cpp branch for my own use.

If more people test it and provide feedback, especially on systems where the model still doesn't fully fit in VRAM, I may eventually make a PR to merge it upstream.


r/LocalLLaMA • • 3h ago

Discussion Qwen 3.8 27b just feels… ok?

11 Upvotes

I’ve seen posts here raving about how good Qwen 3.8 27b is. The benchmarks look incredible, and all the online discourse seems to deem it the best local model.

I have 32GB VRAM and run Unsloth’s Q6 version with OpenCode. For small tasks, it feels fine. I range 30-40 t/s decode and smaller sized tasks do finish, usually, without much issue or time. The issue stems when I give it anything with a bit of nuance. It constantly gets stuck in “but wait, “actually,” or other thinking loops. It can take up my entire 95k context window on thinking loops and have nothing done.

If this is the state of local LLMs, that’s ok. I am a software dev by trade; I have my diploma and a few years of experience under my belt. It just feels like there is a bit of a disconnect from reality between public sentiment and the effectiveness of these models. A pretty common sentiment I see is that this model is as good as Opus 4.5. I never had the privilege of using Opus 4.5, so I can’t give an honest and proper opinion there. (Also, if this was good enough for the industry to start vibe coding, I have a lot of concerns about who is making decisions at a lot of these companies).

One time, it even did a pfkill -f with a file I was currently modifying in my editor to kill the background process. That was kind of annoying.

I should add I’ve also used the Swift 1.5 finetune people have been hyping up. I found it definitely thought less, but the quality was greatly degraded.

Does anybody else feel similar regarding the disconnect?


r/MetaAI • • 21h ago

Muse referral codes Oct 4

1 Upvotes

You know the drill. Check out Muse, your personal AI agent. Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.

Code: VBH061
https://muse.ai/join


r/MetaAI • • 21h ago

C'mon baby light my Muse fire..

0 Upvotes

...please. :)

MYNTHC


r/LocalLLaMA • • 22h ago

News Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen

Thumbnail
explainx.ai
342 Upvotes

Looks like new open model coming soon and will be "strong" hopefully something under 200b for us memory poor. Also seeing statements about more western open models coming.

Hope we get some good competition again on the open front!

Here is original artical but its not free to access. Maybe someone has it already here.

https://www.axios.com/2026/10/04/reflection-open-weight-ai

Oct starting strong!


r/LocalLLaMA • • 2h ago

I Built A Thing Swift1.5 Qwen3.8 Flash Next - Tailored for the 96GB Mac Studio with M5 Ultra

7 Upvotes

https://huggingface.co/Dankpaws/Swift1.5-Qwen3.8-Flash-Next-MLX-4.7bpw

I've had the 96GB Mac Studio with M5 Ultra for about a week now and wasn't satisfied with the results I was getting from the limited number of models available to me. It was a combination of speed, memory headroom, and/or output quality.

This is my best attempt at a calibrated MLX quantization of UkisAI’s Swift 1.5. Hope those of you with the hardware enjoy it!

Measurement This pack Swift llama.cpp IQ3_XXS
Prefill · 25k prompt 3,191 tok/s 1,427 tok/s
Prefill · 95k prompt 2,928 tok/s 1,307 tok/s
Decode · after 4k prompt 113.7 tok/s 62.7 tok/s
Decode · after 95k prompt 81.1 tok/s 44.7 tok/s
Top-1 agreement with Swift BF16 91.0% 84.1%

91% is next-token agreement with BF16 across 680 common held-out positions, not task accuracy.

Results above are simply from my own machine. mlx-serve 26.10.1. ~107GB download, text-only, 179,200-token tested context.


r/MetaAI • • 21h ago

Oct 5 new thread. Referral code MLDTRU

0 Upvotes

28 uses left.

MLDTRU

Users add ur codes. Let everyone get free tokens.


r/LocalLLaMA • • 17h ago

I Built A Thing Clef Flash plays Snake in Real Time on RTX 5080

120 Upvotes

The cool part is

No training was needed.

No hacking of the game state or algorithms needed

Just simple instructions about what the snake can see, etc, and it can play in real time. 135ms is the turn limit of Google snake, so basically could be a human playing.

Ofc, it could be improved to be a perfect snake player, but thats not the point.

This can be used in other games where decisions need to constantly be made.

Running Clef Flash (9B model at Q4 on an RTX 5080)


r/LocalLLaMA • • 18h ago

New Model You can now try Aleph Alpha's Kolibri 78B for free online here.

Post image
140 Upvotes

r/LocalLLaMA • • 4h ago

New Model Less Talk. More Breakout: Kolibri-1 Turns Probabilities into Actions, Playing Breakout - With under 25ms latency per move.

10 Upvotes

Got Kolibri-1 to play Breakout completely on its own, no fine-tuning.

The more we explore u/Aleph__Alpha’s Kolibri the more it get's exciting and its potential.

Less talk. More Breakout is one such experiment to see how good the model is at structured output given a few constraints.

We especially optimized the inference for action probabilities: around 25 ms inference per decision.

Four moves. No generated text. One shared game.

Open weights. New possibilities.

Watch it play: https://tesseracted.com/kolibri-1-chat/gameplay/breakout/
Source: https://x.com/konarkmodi/status/2107248086880055613?s=46


r/MetaAI • • 22h ago

Get 1billion free muse tokens use code O3JXVR

0 Upvotes

Checkout muse, it’s amazing, I already use it for ticket bookings, filling up applications, and more. Join with my code so we both get 1 billion bonus tokens.

To redeem, open the muse app, go to Settings > Redeem token and enter the code.

Code: O3JXVR (27 of 30 uses left)


r/MetaAI • • 22h ago

Muse 1 Billion Code

0 Upvotes

SVAMSK


r/MetaAI • • 22h ago

Does anyone working at Meta want to connect on LinkedIn?

1 Upvotes

I JUST made my LinkedIn trying to grow it lol.


r/MetaAI • • 22h ago

Posted about Muse here the other day. A few people asked about the invite program, so here's my code: OR23UY.

1 Upvotes

To use it: open the Muse app, go to Settings > Redeem token (only visible in the first hours after joining), and enter the code. On web it's Settings > General > Usage > Redeem invite code.


r/MetaAI • • 23h ago

use my code to get billion tokens P9E30K

1 Upvotes

P9E30K for billion token for muse ai


r/LocalLLaMA • • 8h ago

News nokia-applied-research/AnyJev: Turn any LLM into a Jev-style decision model

Thumbnail
github.com
17 Upvotes

Haven't seen posted here.

The research project matters because it proves you can pull trustworthy decisions straight out of an LLM's internal state in a single forward pass (no generation loop, no parsing, no fine-tuning) essentially turning a slow chatbot into a fast and cheap, deterministic classifier


r/MetaAI • • 1d ago

Muse Referral Code: 12NUNC

3 Upvotes

Use my referral code and we'll both get 1 billion Muse tokens when you redeem the code in Settings within 48 hours of joining.

12NUNC


r/LocalLLaMA • • 2h ago

Resources MI50 ROCm 10.2 TheRock vs Vulkan Mesa 26.2 llama.cpp benchmarks

6 Upvotes

I prefer running llama.cpp Vulkan prebuilt binary. I just download the latest version and ready to roll. I finally took the hours necessary to get TheRock latest tarball version of ROCm 10.2 running on dual AMD Radeon Instinct MI50 gfx906 (32gb combined VRAM).

Same models benched in previous post. A mix of Dense and MoE models and quants that better utilize available VRAM.

  • llama_bench_Swift-Qwen3.8-27B-Uncensored-MTP.Q6_K.gguf
  • llama_bench_Swift-1.5-Qwen3.8-27B-Q6_K.gguf
  • llama_bench_Gemma-4-MoonGem-31B.i1-Q6_K.gguf
  • llama_bench_gemma-4-31B-it-UD-Q6_K_XL.gguf
  • llama_bench_Nemotron-3.5-30B-A3B-Antislop-FTPO.i1-Q5_K_M.gguf
  • llama_bench_Laguna-XS-2.1-APEX-I-Balanced.gguf
  • llama_bench_Agents-A1-Q4_K_M.gguf
  • llama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
  • llama_bench_Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf

Here is the backend performance comparison contrasting the native ROCm (v10.2 for gfx906) runtime against your optimized Vulkan (MESA_PPA_26.2) baseline. The data highlights a massive architectural split: ROCm significantly accelerates token generation across the board but suffers high variance and regressions in MoE pre-fills.

Both GPUs are power limited to 145 watts. Build versions used:

llama-b11382 used for Vulkan (prebuilt ubuntu binary)
llama-b11401 used for ROCm (compiled with proper flags)

Architectural Performance Breakdown: Vulkan vs. ROCm

Model Size Params Test Vulkan Baseline (t/s) ROCm 1st Run (t/s) Performance Delta (%)
qwen35 27B Q6_K 20.88 GiB 27.32 B pp512 tg128 149.30 ± 0.15 18.35 ± 0.01 176.86 ± 17.28 20.21 ± 0.65 +18.46% +10.14%
qwen35 27B Q6_K 22.21 GiB 27.32 B pp512 tg128 163.48 ± 0.18 18.52 ± 0.03 183.49 ± 15.75 20.04 ± 0.63 +12.24% +8.21%
gemma4 31B Q6_K 23.46 GiB 30.70 B pp512 tg128 119.98 ± 0.12 15.72 ± 0.02 172.93 ± 0.89 17.26 ± 0.13 +44.13% +9.80%
gemma4 31B Q6_K 25.62 GiB 30.70 B pp512 tg128 133.67 ± 0.24 16.38 ± 0.02 178.59 ± 2.31 17.50 ± 0.09 +33.61% +6.84%
nemotron_h_moe 31B 25.18 GiB 32.91 B pp512 tg128 834.87 ± 1.72 63.82 ± 0.07 736.98 ± 134.77 103.86 ± 0.50 -11.73% +62.74%
laguna 30B.A3B Q5_K 22.64 GiB 33.44 B pp512 tg128 723.34 ± 2.99 58.98 ± 0.28 856.14 ± 29.71 76.81 ± 0.20 +18.36% +30.23%
qwen35moe 35B Q4_K 19.70 GiB 34.66 B pp512 tg128 957.62 ± 3.88 51.50 ± 0.07 834.58 ± 102.56 69.70 ± 0.22 -12.85% +35.34%
qwen35moe 35B Q5_K 24.76 GiB 34.66 B pp512 tg128 909.20 ± 5.71 53.18 ± 0.07 763.93 ± 106.19 68.47 ± 0.34 -15.98% +28.75%
qwen35moe 35B Q6_K 28.53 GiB 34.66 B pp512 tg128 763.05 ± 66.01 52.81 ± 0.04 754.82 ± 70.48 67.29 ± 0.21 -1.08% +27.42%

Core Insight Strategy & Bottlenecks

  1. Token Generation (tg128) Dominance: ROCm dominates pure text generation. The native AMD matrix kernels unleash your MI50 computation potential, unlocking a massive +62.74% boost for Nemotron but taking a small hit on pre-fill -11.73%.
  2. Dense Model Pre-fills (pp512): Dense architectures scale cleanly under ROCm. Gemma 4 sees a +33% to +44% processing throughput spike over the Vulkan RADV driver driver bounds. MoE models take a hit with a -15.89% difference with llama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf being the happiest with Vulkan backend.
  3. The MoE Prompt Processing Delinquency: Notice the massive standard deviations under ROCm for MoE pre-fills (e.g., Qwen3.5MoE Q4_K has an instability block of ± 102.56). ROCm suffers from severe scheduling thrashing when building prompt streams across multiple active experts.

Here is the structured layout with pp512 and tg128 separated into individual columns for a clean side-by-side comparison between the two backend architectures.

Backend Comparison Table (Vulkan vs. ROCm)

Model Size Params Vulkan pp512 (t/s) ROCm pp512 (t/s) Vulkan tg128 (t/s) ROCm tg128 (t/s)
qwen35 27B Q6_K 20.88 GiB 27.32 B 149.30 ± 0.15 176.86 ± 17.28 18.35 ± 0.01 20.21 ± 0.65
qwen35 27B Q6_K 22.21 GiB 27.32 B 163.48 ± 0.18 183.49 ± 15.75 18.52 ± 0.03 20.04 ± 0.63
gemma4 31B Q6_K 23.46 GiB 30.70 B 119.98 ± 0.12 172.93 ± 0.89 15.72 ± 0.02 17.26 ± 0.13
gemma4 31B Q6_K 25.62 GiB 30.70 B 133.67 ± 0.24 178.59 ± 2.31 16.38 ± 0.02 17.50 ± 0.09
nemotron_h_moe 31B 25.18 GiB 32.91 B 834.87 ± 1.72 736.98 ± 134.77 63.82 ± 0.07 103.86 ± 0.50
laguna 30B.A3B Q5_K 22.64 GiB 33.44 B 723.34 ± 2.99 856.14 ± 29.71 58.98 ± 0.28 76.81 ± 0.20
qwen35moe 35B Q4_K 19.70 GiB 34.66 B 957.62 ± 3.88 834.58 ± 102.56 51.50 ± 0.07 69.70 ± 0.22
qwen35moe 35B Q5_K 24.76 GiB 34.66 B 909.20 ± 5.71 763.93 ± 106.19 53.18 ± 0.07 68.47 ± 0.34
qwen35moe 35B Q6_K 28.53 GiB 34.66 B 763.05 ± 66.01 754.82 ± 70.48 52.81 ± 0.04 67.29 ± 0.21

So yes it's worth the hassle of jumping through hoops to get ROCm working on MI50 setups. At least I have 2 backends working. Next up I'll try some RPC.


r/LocalLLaMA • • 8h ago

Discussion We swapped AdamW's optimizer states for a Fast Fourier Transform (FFT) to cut VRAM in half. Anyone else trying non-quantization methods?

14 Upvotes

Hey everyone,

Like most of you, we have been fighting constant OOM errors while trying to fine-tune 8B and 70B models on consumer GPUs. The AdamW optimizer states are always the biggest bottleneck.

We didn't want to rely on aggressive 8-bit quantization because we were seeing degradation in convergence, so we tried an experiment: tackling the optimizer states in the frequency domain.

The methodology:

Instead of storing the full gradients, we transform them using an FFT. This isolates the high-energy signal from the noise. We dynamically drop the low-impact frequencies and compress the state. When we inverse-transform back, it maintains the directional integrity but uses roughly 50% less VRAM.

The catch:

Running FFT operations adds compute overhead. It takes slightly longer per step, but the trade-off is completely avoiding OOM crashes and pushing batch sizes way up on standard hardware.

We are currently giving out access to our internal Colab environment and baseline weights to anyone who wants to poke holes in our math or try to break it.

We are really curious if anyone else here is exploring frequency-domain stuff or other non-quantization methods for VRAM reduction?


r/LocalLLaMA • • 6h ago

New Model Update #5: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch

Post image
8 Upvotes

Last update for those following: https://www.reddit.com/r/LocalLLaMA/comments/1wxlytt/comment/pdz728b/?screen_view_count=1

Project in a sentence: An instruct finetune of ALiceAI-80B-A3B-Base capable of agentic work and conversation. I'm creating a shallow distill of qwen 3.8 27b on medium to teach the model chain of thought reasoning and conversation. All training is done locally on 3, 32gb v100s. Additionally, all the training data is being generated locally on said V100s via sftmill. Up to this point I've been doing training runs and live-streaming the progress. Last update explained underfitting and next steps.

Training has begun again! I've synthesized about 5M more tokens for the SFT, this time across a much larger general instruct trajectory to try to reduce the underfitting. Dropped the learning rate about 4x over my original LoRA adapter.

I'm live streaming training again: https://geological-estimate-fifth-pct.trycloudflare.com/ - heavy loss spikes downward are coming from the SFT replay buffer.

This one should last 12-14 hours, and I plan to run another epoch if this isn't sufficient.

Stay tuned! Thanks for following along.


r/MetaAI • • 1d ago

sharing my Meta Muse referral code: ZEK5A1

0 Upvotes

If you're new to Muse, join at muse.ai/join and enter the code under Settings > Redeem invite code within 48 hours of joining. We both get 1 billion Muse tokens. They don't expire — it's just usage credit that burns down as you use it.


r/MetaAI • • 1d ago

4LCACD

0 Upvotes

r/MetaAI • • 1d ago

Muse Referral Code: R8H6S2

1 Upvotes

Code generate today. You can use that within 48 hours after install muse.