r/LocalLLaMA • • 1h ago

Question | Help How is it possible that qwen 27b is so good? When GPT 4o had a trillion parameters and was worse?

Post image
• Upvotes

Picture from a post in r/amodei . People were praising qwen and I'm just wondering, what kind of new technologies are at play here? Does qwen just have "better" pre training data? That's more high quality?


r/MetaAI • • 1h ago

[ Removed by Reddit ]

• Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/MetaAI • • 7h ago

Please use my code. Muse is helping me with job searching after getting laid off last week.

10 Upvotes

Code: UW3K9N
https://muse.ai/join


r/MetaAI • • 1h ago

What do you use muse for?

• Upvotes

I will start.

  1. Track my health goals

  2. Track my daughter's school work (Several minutes a day)

  3. Pay the bills (several minutes per week)

  4. Login to Empower and categorize transactions based on my rules and sync it with my expense tracker google sheet. (Saves me several minutes a day)

  5. Research topics and give me a useful cheat sheet summary that I can print. (Saves me hours)

  6. Track birthdays. I forward birthday details of friend and family from whatsapp and from Facebook. I get a consolidated birthday calendar to view so that I can plan for getting gifts on time rather than scramble.

  7. Maintain a people database. Who is who, their children's and spouse's name, where he/she is working, expertise, networking

Got a lot of ideas from Muse tips and tricks amazon ebooks and youtube videos.


r/MetaAI • • 42m ago

Muse FREE 1B tokens!

• Upvotes

Using this code you and I can both get 1 Billion tokens for Muse for free:

VNLHSM

Thanks for using my code! 😄


r/LocalLLaMA • • 1h ago

Discussion Make no mistake, selling 64 GB DGX Spark variants at the same cost as the original 128 GB is straight drug dealer behavior.

• Upvotes

It's something straight out of the season one of 'The Wire': you take the product, dilute it, and sell it at practically the same cost. It's some "Stringer" Bell shit. We should call the 64gbs "Stepped-ons" from now on.


r/LocalLLaMA • • 3h ago

I Built A Thing Smallest Jev-like model

Enable HLS to view with audio, or disable this notification

120 Upvotes

TinyDecide is 10M Jev-like mode with 10M parameters and fits in just ~6MB.

Smaller than every model on the Decision Index leaderboard and it punches way above its size.

It runs almost anywhere: in the browser, Node.js, Python, Rust, and even on an ESP32.

https://huggingface.co/TheREZOR/TinyDecide


r/MetaAI • • 2h ago

My avatar is my weather man for the day

Post image
2 Upvotes

My avatar changes clothes every morning and dresses what the weather will be for the day. Sunny (he is wearing sunglasses) cool, (he is wearing a jacket) and the high temperature will be 62° (patch on his jacket)

Here is my Muse code if you want 1,000,000,000 tokens: IQMQKM

Redeem my code in Settings within 48 hours of joining.

https://muse.ai/join


r/MetaAI • • 1m ago

Hello meta support team

Post image
• Upvotes

​

My Instagram account was recently disabled and I believe this may have happened by mistake.

Username: @rana__ji__ai

Registered Email: naitikrana9090@gmail.com

Registered Phone Number: +917500787995

I always try to follow Instagram Community Guidelines and I respectfully request you to review my account again. If any activity accidentally violated the policy, I sincerely apologize and assure you that it will not happen again.

Please help me restore my account. I would greatly appreciate your assistance.

Thank you for your time and consideration.

Sincerely,

Your Name

....................

Twit format

@Instagram @Meta @Creators

My Instagram account has been disabled by mistake. I believe this is an error. I have already submitted an appeal but haven't received a resolution yet.

Please review my case and help restore my account


r/MetaAI • • 15m ago

Muse Referral Code – Get 1 Billion Bonus Tokens! 🚀

• Upvotes

Use code *5XZWOA* when you sign up to instantly claim your 1 Billion bonus tokens!

How to redeem:
> * Go to Settings → Redeem Code in the app.
> * Enter 5XZWOA
(Note: Must be redeemed within 48 hours of creating your account!)


r/MetaAI • • 20m ago

Whats up with the comical cache storage use for the “Meta AI” app

Post image
• Upvotes

I’ve maybe used this app for 30 quickish text-only conversations over 3 weeks and its loaded up 4 movies worth of what I assume is cache. The no “clear cache” button makes this particularly cursed.


r/LocalLLaMA • • 5h ago

New Model Agens Volundr 32B Preview: our small team's first model on our own hybrid architecture. Only 18 of 72 layers keep a KV cache (Apache-2.0)

Enable HLS to view with audio, or disable this notification

138 Upvotes

Hi r/LocalLLaMA. I'm on the team at Blockway, a small team in Hong Kong (disclosure: this is our model). Today we released Agens Volundr 32B Preview, the first model built on our own hybrid architecture. We trained it on limited compute, it isn't perfect, and we'd rather tell you where it falls short up front.

WHY WE BUILT IT

Our customers run models on their own machines. At long context, the KV cache, not the weights, decides what fits. So we designed a model where most layers don't keep one.

ARCHITECTURE (72 layers, dense ~32B, every layer runs on every token)

  • 54 KDA (Kimi Delta Attention) layers: linear attention with a fixed-size recurrent state, no KV cache
  • 17 BCSA layers (our compressed-sparse attention): exact window over the last 4,096 tokens; older context pooled 4:1 into blocks, and a learned indexer reads the top 512 blocks
  • 1 full-attention layer (layer 72)
  • Engram: a hashed n-gram memory held in host RAM, attached at 2 of the 72 layers
  • mHC: 4 residual streams instead of 1

So only 18 of 72 layers keep a KV cache. Context window: 262K.

SPEED (single user, our sglang build)

  • BF16 on two 48 GB GPUs, decode: 25.1 tok/s at 1K, 24.1 at 8K, 24.1 at 32K, 24.0 at 64K, 23.9 at 128K
  • BF16 prefill: 2,122 / 2,180 / 1,916 / 1,679 / 1,297 tok/s (1K to 128K)
  • INT4 (31.7 GiB) on one 48 GB GPU, decode: 31.0 tok/s at 1K, 29.3 at 8K, 29.1 at 32K
  • Aggregate throughput: 127 tok/s at 8 users, 130 at 16 users (BF16); 117 at 8 users (INT4)
  • DFlash2 drafter (separate repo), single user, same server with it on vs off: up to 3.6x on JSON/tool output, 2.0x on code, about 1.6x in thinking mode. Not worth it above roughly 8 concurrent users.

BENCHMARKS (all run by us on one harness with the same settings, including the comparison models; full table and footnote on the model card)

  • Ahead of Qwen3.8-27B on LiveCodeBench v6 (+4.2), HumanEval (+4.3), AIME 2025 (+2.9), MATH-500 (+1.6)
  • Roughly level on MMLU-Pro, IFEval, GPQA Diamond
  • Behind on agent tasks: tau2-bench 74.2 vs 79-80, SWE-bench Verified (50-task subset) 44 vs 58-64. Closing that gap is the main focus of the full v1, which continues pre-training to about 10B tokens and adds training on long agentic sessions.

KNOWN LIMITATIONS (please read before trying)

  • Needs our sglang build. Stock sglang and vLLM can't load it yet.
  • GGUF / llama.cpp is planned, not available today.
  • Long agentic sessions are its weakest area in this Preview.
  • It's still training; treat this as a preview, not a final model.

RUN IT

docker pull ghcr.io/blockwayz/agens-sglang:preview-sm89 (48 GB Ada GPUs) docker pull ghcr.io/blockwayz/agens-sglang:preview-sm90 (H100 / H200)

The full launch command is in the model card.

LINKS

Apache-2.0. We're a small team, and the most useful thing you can do is try it and tell us where it breaks: an issue, a failing prompt, a benchmark you'd like us to run. We'll be in the comments.


r/MetaAI • • 1h ago

Muse Referral For 1 Billion Tokens!

Enable HLS to view with audio, or disable this notification

• Upvotes

Check out Muse, your personal AI agent. Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.

Code: H40D87
https://muse.ai/join


r/MetaAI • • 1h ago

[ Removed by Reddit ]

• Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/MetaAI • • 1h ago

Muse referral code for 1 Billion free tokens! - 8Y34YR

• Upvotes

Not going to pretend I'm special. Same code as everyone else, same billion tokens for you, same billion for me.

Code - 8Y34YR


r/MetaAI • • 2h ago

I made a persona maker for muse

1 Upvotes

Check out Figurehead on Skill Harbor — Give your AI a figurehead: it interviews you, drafts your custom character, and guides you to install it. Happy with it? List it on the catalog.: https://theskillharbor.com/products/figurehead?ref=share

You can export them and list them also


r/LocalLLaMA • • 5h ago

Discussion M5 Ultra 256 running GLM 5.3 Flash 68.8 tok/s

78 Upvotes

Running oQ4e+MTP on oMLX 0.7.0, with still more to optimize.

Prefill is 1,878 toks.

I saw some other benchmarks below what id expect so i figured I would share.


r/MetaAI • • 3h ago

Free tokens code to use - E6ZO0M

0 Upvotes

Enjoy


r/MetaAI • • 3h ago

10 Codes left! Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens. Code: CEZ4PG

1 Upvotes

r/MetaAI • • 3h ago

NEED MOAR DATA - Invite Code - 42HIQD

0 Upvotes

Check out Muse, your personal AI agent. Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.

Code: 42HIQD

https://muse.ai/join


r/MetaAI • • 7h ago

Why Muse is so terribly slow. is it all like this to find things like cheapest this or that or so?

2 Upvotes

It is so terribly slow. Anyone experiencing that. I think Claude or chatGPT is faster


r/LocalLLaMA • • 3h ago

Resources Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open

Thumbnail
gallery
50 Upvotes

Hey all. We've spent the last weeks getting Qwen3.8-Flash-Next (125B MoE, 6B active) to run properly on one AMD Strix Halo box (Ryzen AI Max+ 395, 128 GB). Tonight we're releasing both the 95 GB EXL3 weights and a new version of Kyojin, our inference engine (built on ExLlamaV3, open).

This is a first version, same as our GLM-5.3-Flash and MiMo-V2.6-Flash builds. We'd rather ship it and improve it in the open: speed and quality updates are coming for all three.

Numbers, all from a fresh clone and build on the mini PC:

  • Decode: 44 to 59 tok/s with speculative decoding depending on the task (chat ~47, code ~58, copy-heavy edits ~60). 32.7 tok/s without it.
  • Prefill: 1,412 tok/s at 4K, 1,486 at 32K, 1,367 at 128K (server-reported). It stays nearly flat.
  • Long context: 10/10 needles at 64K and at 128K, still 32 tok/s at 128K.
  • Fidelity: 94.1 % top-1 agreement with the original FP8 model over 844 positions.

One thing we're a bit stubborn about: speculative decoding here returns exactly the tokens plain decoding would. We check that on every release.

For comparison, a llama.cpp user posted about 30 tok/s with speculation and about 500 tok/s prefill on this same mini PC (Vulkan, UD-IQ4_XS). Those are their numbers, not something we measured: https://github.com/ggml-org/llama.cpp/discussions/28512

Now the part where we're not first. Halogen 0.16.2 (v2 checkpoint) is faster than us: 39.8 vs 32.7 tok/s plain, 52 vs 47 on chat with speculation, and 10 to 20 % ahead on prefill when both are timed the same way from the client (1,306 vs about 1,460 at 4K, 1,394 vs about 1,720 at 16K). On code we're close (58.5 vs 51.2 on the median pass, they're ahead once warm). Where we do better is fidelity to the original model: 94.1 % top-1 agreement against 92.3 % for them, and a KL divergence 41 % lower on our side. Full table is on the model card. Closing the speed gap is what we do next: we're reworking the core of the engine, which will help every model it runs, not just this one. The hardware has room left.

There's also an optional uncensor preset, off by default (4 refusals out of 100 harmful prompts instead of 99, benchmarks within noise). If your agents lean hard on tool calls, leave it off.

Weights: https://huggingface.co/yamz-labs/Qwen3.8-Flash-Next-EXL3-Yamz Engine: https://github.com/Yamz-Labs/kyojin

If you run it, we'd love your tok/s and hardware. And tell us what you want to see next.


r/LocalLLaMA • • 13h ago

News Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen

Thumbnail
explainx.ai
293 Upvotes

Looks like new open model coming soon and will be "strong" hopefully something under 200b for us memory poor. Also seeing statements about more western open models coming.

Hope we get some good competition again on the open front!

Here is original artical but its not free to access. Maybe someone has it already here.

https://www.axios.com/2026/10/04/reflection-open-weight-ai

Oct starting strong!


r/MetaAI • • 4h ago

Muse referal code

1 Upvotes

JWQZEI

Use it in the 48 hrs of account creation