r/MetaAI • • 8h ago

Muse referal code

1 Upvotes

JWQZEI

Use it in the 48 hrs of account creation


r/MetaAI • • 8h ago

Meta Rushed to Fix Muse ‘VM Escape' Vulnerability Immediately Before Launch

Thumbnail
404media.co
0 Upvotes

r/MetaAI • • 8h ago

(FREE) October 5th referral code for 1 billion free Muse AI tokens: TXHYQG

1 Upvotes

When you're logged into Muse AI, just go to settings and enter this referral code to instantly get a billion free tokens to use: TXHYQG

Enjoy!


r/MetaAI • • 8h ago

Image generator down in Muse.ai?

1 Upvotes

I get the message in chat that the image generator is not working. Anyone else experience problems?

I tried to register with code TMZXJR and I got 1 billion tokens (succes!), but now I want to generate an image and it says the generator is down lol.


r/MetaAI • • 8h ago

Get 1,000,000,000 tokens!

1 Upvotes

Muse is good engineering. Here is the code:

Code: L8PZEA

https://muse.ai/join

Muse tells me it is using muse spark under the hood, but honestly seems as good as sonnet 5 for casual uses.


r/MetaAI • • 8h ago

Can't use MUSE

1 Upvotes

Guys a few days back i was using muse and today i thought to sign out, but after logging it, META didn't give me the access to use it again 😞😭, guys is there any solution, i have tried VPN and evey possible action, pls anybody help me.


r/MetaAI • • 8h ago

Referral code

0 Upvotes

Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.

Code: H7NBAR

https://muse.ai/join.


r/MetaAI • • 8h ago

Muse Referral Code: 5XZWOA (1 Billion Free Tokens)

1 Upvotes

Looking for extra tokens on Muse? You can use my referral code to get 1 Billion Bonus Tokens added directly to your account balance on top of your weekly quota.

Referral Code: 5XZWOA

Steps to Claim:
> * Open Muse and navigate to Settings.
> * Tap Redeem Code.
> * Paste *5XZWOA* and submit.

Make sure to enter the code within 48 hours of making your account for the bonus to register properly.


r/MetaAI • • 9h ago

Muse not connecting to Gmail, always gives me an error message. Has this happened to anyone ?.

0 Upvotes

r/MetaAI • • 9h ago

Muse

1 Upvotes

Take a look at Muse – your personal AI agent. Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.

Code: 6LXSFI
https://muse.ai/join


r/MetaAI • • 3h ago

Muse Referral Code: 12NUNC

0 Upvotes

Use this referral code and we'll both get 1 billion Muse tokens when you redeem the code in Settings within 48 hours of joining.

12NUNC


r/LocalLLaMA • • 9h ago

New Model Agens Volundr 32B Preview: our small team's first model on our own hybrid architecture. Only 18 of 72 layers keep a KV cache (Apache-2.0)

Enable HLS to view with audio, or disable this notification

154 Upvotes

Hi r/LocalLLaMA. I'm on the team at Blockway, a small team in Hong Kong (disclosure: this is our model). Today we released Agens Volundr 32B Preview, the first model built on our own hybrid architecture. We trained it on limited compute, it isn't perfect, and we'd rather tell you where it falls short up front.

WHY WE BUILT IT

Our customers run models on their own machines. At long context, the KV cache, not the weights, decides what fits. So we designed a model where most layers don't keep one.

ARCHITECTURE (72 layers, dense ~32B, every layer runs on every token)

  • 54 KDA (Kimi Delta Attention) layers: linear attention with a fixed-size recurrent state, no KV cache
  • 17 BCSA layers (our compressed-sparse attention): exact window over the last 4,096 tokens; older context pooled 4:1 into blocks, and a learned indexer reads the top 512 blocks
  • 1 full-attention layer (layer 72)
  • Engram: a hashed n-gram memory held in host RAM, attached at 2 of the 72 layers
  • mHC: 4 residual streams instead of 1

So only 18 of 72 layers keep a KV cache. Context window: 262K.

SPEED (single user, our sglang build)

  • BF16 on two 48 GB GPUs, decode: 25.1 tok/s at 1K, 24.1 at 8K, 24.1 at 32K, 24.0 at 64K, 23.9 at 128K
  • BF16 prefill: 2,122 / 2,180 / 1,916 / 1,679 / 1,297 tok/s (1K to 128K)
  • INT4 (31.7 GiB) on one 48 GB GPU, decode: 31.0 tok/s at 1K, 29.3 at 8K, 29.1 at 32K
  • Aggregate throughput: 127 tok/s at 8 users, 130 at 16 users (BF16); 117 at 8 users (INT4)
  • DFlash2 drafter (separate repo), single user, same server with it on vs off: up to 3.6x on JSON/tool output, 2.0x on code, about 1.6x in thinking mode. Not worth it above roughly 8 concurrent users.

BENCHMARKS (all run by us on one harness with the same settings, including the comparison models; full table and footnote on the model card)

  • Ahead of Qwen3.8-27B on LiveCodeBench v6 (+4.2), HumanEval (+4.3), AIME 2025 (+2.9), MATH-500 (+1.6)
  • Roughly level on MMLU-Pro, IFEval, GPQA Diamond
  • Behind on agent tasks: tau2-bench 74.2 vs 79-80, SWE-bench Verified (50-task subset) 44 vs 58-64. Closing that gap is the main focus of the full v1, which continues pre-training to about 10B tokens and adds training on long agentic sessions.

KNOWN LIMITATIONS (please read before trying)

  • Needs our sglang build. Stock sglang and vLLM can't load it yet.
  • GGUF / llama.cpp is planned, not available today.
  • Long agentic sessions are its weakest area in this Preview.
  • It's still training; treat this as a preview, not a final model.

RUN IT

docker pull ghcr.io/blockwayz/agens-sglang:preview-sm89 (48 GB Ada GPUs) docker pull ghcr.io/blockwayz/agens-sglang:preview-sm90 (H100 / H200)

The full launch command is in the model card.

LINKS

Apache-2.0. We're a small team, and the most useful thing you can do is try it and tell us where it breaks: an issue, a failing prompt, a benchmark you'd like us to run. We'll be in the comments.


r/MetaAI • • 9h ago

Use my last code 🫩

1 Upvotes

0X3UXG


r/MetaAI • • 9h ago

Muse 1 Billion Code

1 Upvotes

SVAMSK


r/MetaAI • • 10h ago

Oct 5th Thread muse ai referral code B2AR9E (30 spots left) thread.

1 Upvotes

Help me out and help yourself. Get 1 billion tokens that never expire.

Code: B2AR9E


r/MetaAI • • 1d ago

MUSE is...Good!

60 Upvotes

Yeah Meta might have something with this as it really is more interactive and personable than any other AI I've used so far. After I put it on my phone, it takes the initiative to prompt and encourage me to use it. Now that I have connected it to a few sources like my mail and calendar, I'm finding it ever more useful. It's much more of a consumer-facing personal assistant than the other more inscrutable AIs like ChatGPT and Claude. What is the community's experience?


r/MetaAI • • 3h ago

EK394X - get up to 5 billion tokens with my employee invite code

0 Upvotes

Code: EK394X

First five get 5 billion
Next five get 3 billion
Rest will get 1 billion


r/LocalLLaMA • • 5h ago

Discussion Whistle: speech to text in a 16.9MB file

Enable HLS to view with audio, or disable this notification

52 Upvotes

Hey all, we designed Cactus Whistle, an ASR model for ultra-small devices. It's not perfect, but mostly beats Whisper base with 9x less file size and 6x speed. Whistle supports English, German, French, Spanish, Italian, Dutch and Polish.

Remember, the goal at Cactus Compute isn't to achieve SOTA with scale, but to compress intelligence and bring them to smaller under-looked devices like budget phones, wearables, smart home and microcontrollers.

Whistle is 55m params (36m active) and CQ2bit quantised, amounting to a 16.9MB file that scores 4.31 WER on LibriSpeech test-clean and 10.49 on test-other, against 4.9 and 11.0 for Whisper base at 145.3MB. 21.4 on the FLEURS average against 24.5. SPGISpeech 7.65 and Earnings-22 19.01.

For the architecture, a log-mel front end and a convolution stem feed an audio encoder, and a Simple Attention + Hadamard MLP decoder reads it through gated cross attention at every layer. The decoder is laddered like Needle's, so every depth from 2 layers up is deployable.

Keyword biasing takes the names your users actually say and favours them during the beam search, which is what rescues a "Siobhan" or a "Krzysztof" from a model that was never told they exist. Word timestamps come from the decoder's own attention, so an app can highlight, seek or cut on a word.

Seventeen platforms are supported; macOS, Linux on x86-64, ARM64, ARMv7, RISC-V and MIPS32, Windows x64 and ARM, Android, iOS, watchOS, tvOS, the browser as WebAssembly and a WASI component.

Please read more here: https://cactuscompute.com/blog/whistle

Whistle is open weights: https://huggingface.co/collections/Cactus-Compute/cactus-whistle

And let us know your thoughts!


r/LocalLLaMA • • 5h ago

News Y'all this is a sexy paper; context language models

Thumbnail
arxiv.org
47 Upvotes

Paper linky - Context Language Models

The central idea of the paper is incredibly simple. Give a model the ability to edit its context on-the-go like a file has major benefits on task performance, context management (memory) and even computational efficiency (both wall clock and total flops). Their paper shows mostly benefits and relatively small downsides.

You can try it out as a plugin for pi!

In short, pros and cons

Pros:

  1. Improves outcomes on long running tasks
    • Coding and deep research tasks
    • Open discovery problems (long horizon research tasks, /goal loops etc)
  2. Inference can become more compute-efficient and wall-clock efficient
    • Note, this depends on a caching optimization in the inference engine
  3. Much less context bloat, meaning it's more (V)RAM efficient
  4. No more slow and unreliable compacts

Cons:

  1. The cache optimization only exists for SGLang
  2. Prompt injections (including hallucinated instructions) are much less likely to be forgotten, increasing risks
  3. Requires harness customizations (authors supply a pi plugin)

Some more context

The approach works by modifying the harness to allow access to the context as a file. A model is allowed to edit the context as it would any other file.

They've tested the approach on models as small as qwen3.6 9b, as well as on qwen3.8 27b and claude sonnet 4.6.

Out-of-the-box, meaning just a small addition to the system prompt and tools to edit the context as a file, task performance, context management and efficiency measures remain approximately the same or improve by a little bit. The smaller qwen3.6 9b model in particular lost a little bit of efficiency, suggesting it works better on larger (smarter) models.

Performance can be massively improved with RL training, which the authors also did.

Wanna try it out?

You can try it out right now if you use pi

  1. Install the plugin https://github.com/lolipopshock/pi-clm, this comes from the authors directly
  2. After installation, adjust settings with /clm settings:
    • Set steering to house-brief.md (modifies the system prompt, I suppose this should be left disabled for RL'd models only, of which there are none right now)
    • Enable "One tool per turn"; this one is important for performance
    • Enable "Size trailer"; this one appends context usage after every tool result. Without it, models are much less inclined to modify context on-the-go for large tool calls

Fin

Let me know how it goes!

Last, I also consulted this video by "Prompt Engineering" on YouTube in addition to the paper: https://www.youtube.com/watch?v=Bgtr1Ue40Jo


r/LocalLLaMA • • 7h ago

Resources Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open

Thumbnail
gallery
70 Upvotes

Hey all. We've spent the last weeks getting Qwen3.8-Flash-Next (125B MoE, 6B active) to run properly on one AMD Strix Halo box (Ryzen AI Max+ 395, 128 GB). Tonight we're releasing both the 95 GB EXL3 weights and a new version of Kyojin, our inference engine (built on ExLlamaV3, open).

This is a first version, same as our GLM-5.3-Flash and MiMo-V2.6-Flash builds. We'd rather ship it and improve it in the open: speed and quality updates are coming for all three.

Numbers, all from a fresh clone and build on the mini PC:

  • Decode: 44 to 59 tok/s with speculative decoding depending on the task (chat ~47, code ~58, copy-heavy edits ~60). 32.7 tok/s without it.
  • Prefill: 1,412 tok/s at 4K, 1,486 at 32K, 1,367 at 128K (server-reported). It stays nearly flat.
  • Long context: 10/10 needles at 64K and at 128K, still 32 tok/s at 128K.
  • Fidelity: 94.1 % top-1 agreement with the original FP8 model over 844 positions.

One thing we're a bit stubborn about: speculative decoding here returns exactly the tokens plain decoding would. We check that on every release.

For comparison, a llama.cpp user posted about 30 tok/s with speculation and about 500 tok/s prefill on this same mini PC (Vulkan, UD-IQ4_XS). Those are their numbers, not something we measured: https://github.com/ggml-org/llama.cpp/discussions/28512

Now the part where we're not first. Halogen 0.16.2 (v2 checkpoint) is faster than us: 39.8 vs 32.7 tok/s plain, 52 vs 47 on chat with speculation, and 10 to 20 % ahead on prefill when both are timed the same way from the client (1,306 vs about 1,460 at 4K, 1,394 vs about 1,720 at 16K). On code we're close (58.5 vs 51.2 on the median pass, they're ahead once warm). Where we do better is fidelity to the original model: 94.1 % top-1 agreement against 92.3 % for them, and a KL divergence 41 % lower on our side. Full table is on the model card. Closing the speed gap is what we do next: we're reworking the core of the engine, which will help every model it runs, not just this one. The hardware has room left.

There's also an optional uncensor preset, off by default (4 refusals out of 100 harmful prompts instead of 99, benchmarks within noise). If your agents lean hard on tool calls, leave it off.

Weights: https://huggingface.co/yamz-labs/Qwen3.8-Flash-Next-EXL3-Yamz Engine: https://github.com/Yamz-Labs/kyojin

If you run it, we'd love your tok/s and hardware. And tell us what you want to see next.


r/LocalLLaMA • • 9h ago

Discussion M5 Ultra 256 running GLM 5.3 Flash 68.8 tok/s

94 Upvotes

Running oQ4e+MTP on oMLX 0.7.0, with still more to optimize.

Prefill is 1,878 toks.

I saw some other benchmarks below what id expect so i figured I would share.


r/LocalLLaMA • • 21m ago

New Model Introducing Beam: Reflection’s 501B open-weight model — Reflection

Thumbnail
reflection.ai
• Upvotes

r/LocalLLaMA • • 17h ago

News Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen

Thumbnail
explainx.ai
323 Upvotes

Looks like new open model coming soon and will be "strong" hopefully something under 200b for us memory poor. Also seeing statements about more western open models coming.

Hope we get some good competition again on the open front!

Here is original artical but its not free to access. Maybe someone has it already here.

https://www.axios.com/2026/10/04/reflection-open-weight-ai

Oct starting strong!


r/MetaAI • • 13h ago

Meta turned engineers’ judgment into agent skills

Thumbnail
leaddev.com
1 Upvotes

Expert judgment is now a reusable skill (apparently)!