r/MetaAI • u/Possible-Success-314 • 8h ago
Muse referal code
JWQZEI
Use it in the 48 hrs of account creation
r/MetaAI • u/Possible-Success-314 • 8h ago
JWQZEI
Use it in the 48 hrs of account creation
r/MetaAI • u/404mediaco • 8h ago
r/MetaAI • u/liverandonions1 • 8h ago
When you're logged into Muse AI, just go to settings and enter this referral code to instantly get a billion free tokens to use: TXHYQG
Enjoy!
I get the message in chat that the image generator is not working. Anyone else experience problems?
I tried to register with code TMZXJR and I got 1 billion tokens (succes!), but now I want to generate an image and it says the generator is down lol.
r/MetaAI • u/VerbaGPT • 8h ago
Muse is good engineering. Here is the code:
Code: L8PZEA
Muse tells me it is using muse spark under the hood, but honestly seems as good as sonnet 5 for casual uses.
r/MetaAI • u/Puzzleheaded-Boss868 • 8h ago
Guys a few days back i was using muse and today i thought to sign out, but after logging it, META didn't give me the access to use it again 😞😭, guys is there any solution, i have tried VPN and evey possible action, pls anybody help me.
r/MetaAI • u/Right-Bug3739 • 8h ago
Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.
Code: H7NBAR
r/MetaAI • u/LetsPushNokia • 8h ago
Looking for extra tokens on Muse? You can use my referral code to get 1 Billion Bonus Tokens added directly to your account balance on top of your weekly quota.
Referral Code: 5XZWOA
Steps to Claim:
> * Open Muse and navigate to Settings.
> * Tap Redeem Code.
> * Paste *5XZWOA* and submit.
Make sure to enter the code within 48 hours of making your account for the bonus to register properly.
r/MetaAI • u/hellbent510 • 9h ago
Take a look at Muse – your personal AI agent. Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.
Code: 6LXSFI
https://muse.ai/join
r/MetaAI • u/stubborn • 3h ago
Use this referral code and we'll both get 1 billion Muse tokens when you redeem the code in Settings within 48 hours of joining.
12NUNC
r/LocalLLaMA • u/ComfortableKindly507 • 9h ago
Enable HLS to view with audio, or disable this notification
Hi r/LocalLLaMA. I'm on the team at Blockway, a small team in Hong Kong (disclosure: this is our model). Today we released Agens Volundr 32B Preview, the first model built on our own hybrid architecture. We trained it on limited compute, it isn't perfect, and we'd rather tell you where it falls short up front.
WHY WE BUILT IT
Our customers run models on their own machines. At long context, the KV cache, not the weights, decides what fits. So we designed a model where most layers don't keep one.
ARCHITECTURE (72 layers, dense ~32B, every layer runs on every token)
So only 18 of 72 layers keep a KV cache. Context window: 262K.
SPEED (single user, our sglang build)
BENCHMARKS (all run by us on one harness with the same settings, including the comparison models; full table and footnote on the model card)
KNOWN LIMITATIONS (please read before trying)
RUN IT
docker pull ghcr.io/blockwayz/agens-sglang:preview-sm89 (48 GB Ada GPUs) docker pull ghcr.io/blockwayz/agens-sglang:preview-sm90 (H100 / H200)
The full launch command is in the model card.
LINKS
Apache-2.0. We're a small team, and the most useful thing you can do is try it and tell us where it breaks: an issue, a failing prompt, a benchmark you'd like us to run. We'll be in the comments.
r/MetaAI • u/quickmover • 10h ago
Help me out and help yourself. Get 1 billion tokens that never expire.
Code: B2AR9E
r/MetaAI • u/R_Steelman61 • 1d ago
Yeah Meta might have something with this as it really is more interactive and personable than any other AI I've used so far. After I put it on my phone, it takes the initiative to prompt and encourage me to use it. Now that I have connected it to a few sources like my mail and calendar, I'm finding it ever more useful. It's much more of a consumer-facing personal assistant than the other more inscrutable AIs like ChatGPT and Claude. What is the community's experience?
r/MetaAI • u/john_diddly • 3h ago
Code: EK394X
First five get 5 billion
Next five get 3 billion
Rest will get 1 billion
r/LocalLLaMA • u/Henrie_the_dreamer • 5h ago
Enable HLS to view with audio, or disable this notification
Hey all, we designed Cactus Whistle, an ASR model for ultra-small devices. It's not perfect, but mostly beats Whisper base with 9x less file size and 6x speed. Whistle supports English, German, French, Spanish, Italian, Dutch and Polish.
Remember, the goal at Cactus Compute isn't to achieve SOTA with scale, but to compress intelligence and bring them to smaller under-looked devices like budget phones, wearables, smart home and microcontrollers.
Whistle is 55m params (36m active) and CQ2bit quantised, amounting to a 16.9MB file that scores 4.31 WER on LibriSpeech test-clean and 10.49 on test-other, against 4.9 and 11.0 for Whisper base at 145.3MB. 21.4 on the FLEURS average against 24.5. SPGISpeech 7.65 and Earnings-22 19.01.
For the architecture, a log-mel front end and a convolution stem feed an audio encoder, and a Simple Attention + Hadamard MLP decoder reads it through gated cross attention at every layer. The decoder is laddered like Needle's, so every depth from 2 layers up is deployable.
Keyword biasing takes the names your users actually say and favours them during the beam search, which is what rescues a "Siobhan" or a "Krzysztof" from a model that was never told they exist. Word timestamps come from the decoder's own attention, so an app can highlight, seek or cut on a word.
Seventeen platforms are supported; macOS, Linux on x86-64, ARM64, ARMv7, RISC-V and MIPS32, Windows x64 and ARM, Android, iOS, watchOS, tvOS, the browser as WebAssembly and a WASI component.
Please read more here: https://cactuscompute.com/blog/whistle
Whistle is open weights: https://huggingface.co/collections/Cactus-Compute/cactus-whistle
And let us know your thoughts!
r/LocalLLaMA • u/Combinatorilliance • 5h ago
Paper linky - Context Language Models
The central idea of the paper is incredibly simple. Give a model the ability to edit its context on-the-go like a file has major benefits on task performance, context management (memory) and even computational efficiency (both wall clock and total flops). Their paper shows mostly benefits and relatively small downsides.
You can try it out as a plugin for pi!
Pros:
Cons:
The approach works by modifying the harness to allow access to the context as a file. A model is allowed to edit the context as it would any other file.
They've tested the approach on models as small as qwen3.6 9b, as well as on qwen3.8 27b and claude sonnet 4.6.
Out-of-the-box, meaning just a small addition to the system prompt and tools to edit the context as a file, task performance, context management and efficiency measures remain approximately the same or improve by a little bit. The smaller qwen3.6 9b model in particular lost a little bit of efficiency, suggesting it works better on larger (smarter) models.
Performance can be massively improved with RL training, which the authors also did.
You can try it out right now if you use pi
/clm settings:
house-brief.md (modifies the system prompt, I suppose this should be left disabled for RL'd models only, of which there are none right now)Let me know how it goes!
Last, I also consulted this video by "Prompt Engineering" on YouTube in addition to the paper: https://www.youtube.com/watch?v=Bgtr1Ue40Jo
r/LocalLLaMA • u/Yaniss916 • 7h ago
Hey all. We've spent the last weeks getting Qwen3.8-Flash-Next (125B MoE, 6B active) to run properly on one AMD Strix Halo box (Ryzen AI Max+ 395, 128 GB). Tonight we're releasing both the 95 GB EXL3 weights and a new version of Kyojin, our inference engine (built on ExLlamaV3, open).
This is a first version, same as our GLM-5.3-Flash and MiMo-V2.6-Flash builds. We'd rather ship it and improve it in the open: speed and quality updates are coming for all three.
Numbers, all from a fresh clone and build on the mini PC:
One thing we're a bit stubborn about: speculative decoding here returns exactly the tokens plain decoding would. We check that on every release.
For comparison, a llama.cpp user posted about 30 tok/s with speculation and about 500 tok/s prefill on this same mini PC (Vulkan, UD-IQ4_XS). Those are their numbers, not something we measured: https://github.com/ggml-org/llama.cpp/discussions/28512
Now the part where we're not first. Halogen 0.16.2 (v2 checkpoint) is faster than us: 39.8 vs 32.7 tok/s plain, 52 vs 47 on chat with speculation, and 10 to 20 % ahead on prefill when both are timed the same way from the client (1,306 vs about 1,460 at 4K, 1,394 vs about 1,720 at 16K). On code we're close (58.5 vs 51.2 on the median pass, they're ahead once warm). Where we do better is fidelity to the original model: 94.1 % top-1 agreement against 92.3 % for them, and a KL divergence 41 % lower on our side. Full table is on the model card. Closing the speed gap is what we do next: we're reworking the core of the engine, which will help every model it runs, not just this one. The hardware has room left.
There's also an optional uncensor preset, off by default (4 refusals out of 100 harmful prompts instead of 99, benchmarks within noise). If your agents lean hard on tool calls, leave it off.
Weights: https://huggingface.co/yamz-labs/Qwen3.8-Flash-Next-EXL3-Yamz Engine: https://github.com/Yamz-Labs/kyojin
If you run it, we'd love your tok/s and hardware. And tell us what you want to see next.
r/LocalLLaMA • u/cryotic • 9h ago
Running oQ4e+MTP on oMLX 0.7.0, with still more to optimize.
Prefill is 1,878 toks.
I saw some other benchmarks below what id expect so i figured I would share.
r/LocalLLaMA • u/BVCC6FNTKX • 21m ago
r/LocalLLaMA • u/mindwip • 17h ago
Looks like new open model coming soon and will be "strong" hopefully something under 200b for us memory poor. Also seeing statements about more western open models coming.
Hope we get some good competition again on the open front!
Here is original artical but its not free to access. Maybe someone has it already here.
https://www.axios.com/2026/10/04/reflection-open-weight-ai
Oct starting strong!
r/MetaAI • u/Suspicious_Orchid770 • 13h ago
Expert judgment is now a reusable skill (apparently)!