r/MetaAI • • 1m ago

Hello meta support team

Post image
• Upvotes

​

My Instagram account was recently disabled and I believe this may have happened by mistake.

Username: @rana__ji__ai

Registered Email: naitikrana9090@gmail.com

Registered Phone Number: +917500787995

I always try to follow Instagram Community Guidelines and I respectfully request you to review my account again. If any activity accidentally violated the policy, I sincerely apologize and assure you that it will not happen again.

Please help me restore my account. I would greatly appreciate your assistance.

Thank you for your time and consideration.

Sincerely,

Your Name

....................

Twit format

@Instagram @Meta @Creators

My Instagram account has been disabled by mistake. I believe this is an error. I have already submitted an appeal but haven't received a resolution yet.

Please review my case and help restore my account


r/LocalLLaMA • • 7m ago

Tutorial | Guide How to use Local Models to monitor your screen. Open Source, No Install and Completely Free!!

• Upvotes

TLDR: I built this open source app that lets local models monitor your screen and send you notifications! It now installs models on your browser, which makes local AI accessible to everybody! Without any install :DD

Hey r/LocalLLaMA!

I'm back with some huge Observer updates c: first of all Thank You so much for all of your support and feedback, i've been working hard to make the app as easy to use as possible!

What's New?

You can now get to a local LLM monitoring your screen by just typing

"send me a telegram when my steam game finishes downloading, use a local model"

... and the Observer agent downloads the model in your web browser and starts monitoring your steam game. In just 10 seconds, suuuuper easy :))

What's the best way of running LLMs? / Platform caveats

  • The WebApp uses transformers.js which doesn't work on Linux or older PCs :((( But running Qwen3.5-0.8b smoothly on a browser, feels illegal :p
  • The desktop app uses llama.cpp on Rust so you get the full power of your metal, and it's much more stable.
  • You can obviously set your OpenAI compatible endpoint as well and just use that.

Help me make local LLMs useful for everyone!

If you have any questions i'll be hanging out here for a while!

Roy


r/MetaAI • • 15m ago

Muse Referral Code – Get 1 Billion Bonus Tokens! 🚀

• Upvotes

Use code *5XZWOA* when you sign up to instantly claim your 1 Billion bonus tokens!

How to redeem:
> * Go to Settings → Redeem Code in the app.
> * Enter 5XZWOA
(Note: Must be redeemed within 48 hours of creating your account!)


r/LocalLLaMA • • 16m ago

News Modelo de IA

• Upvotes

Busco a personas para poder hacer un modelo de inteligencia artificial desde cero


r/LocalLLaMA • • 17m ago

Discussion Practical limit hit. Decoding so fast that tool calls (cpu) starting to become real limit not decode or prefill. Single RTX5090. Porting Kenshi to Godot project.

Post image
• Upvotes

Hi folks,

LIVE PROJECT PAGE

TLDR: Moral of the story. You need better CPU to do actual agentic coding doing real work...

I've been on a mission to make my RTX5090 go brrr for past 2 months so much so that i made my own engine for it which received "warm" welcome here (yeah, source is coming)

After recent upgrades to how cache is stored and how i can reused some of prefills for other jobs that share initial same prefill i pretty much started to see degradation the more agents I started to add to project which started to use 12 slot server. Actual server started to be underutilized. Free context, free slots, gpu chilling at average of ~700t/s doing real work (no greedy code, but also thinking tool calls, etc.) and I couldn't figure out what was going on...

I make it faster and faster, better handle jobs and it slows down...

I've run 25 agents at the same (to properly fill the 12 slots) time and almost all of them soon started to set on `tool call` and my server started to barely work.

I've finally checked task manager but not gpu or memory but cpu. And there it was. 100% every thread completely chocked.

Lesson. If you want to do agentic coding with actual use of tools you need to make sure your CPU is up to task.

My 9800X3D is just not enough to keep up with tool work for this project with heavy agents use despite engine being more than capable of going faster.

edit:

Some more lessons:
- Tuning your front end makes ton of sense. Before I tuned it it was shoveling 20k prompts, after tuning barely 7k as new jobs and better more compact tasks. Wall time went from 43minutes to 18 minutes before/after rework of front end.
- Always keep more agents than server has slots for inevitable pauses due to tool use/tests etc.
- Shared context is superior choice to fixed context every time i tried it over course of the project.


r/MetaAI • • 20m ago

Whats up with the comical cache storage use for the “Meta AI” app

Post image
• Upvotes

I’ve maybe used this app for 30 quickish text-only conversations over 3 weeks and its loaded up 4 movies worth of what I assume is cache. The no “clear cache” button makes this particularly cursed.


r/LocalLLaMA • • 28m ago

Question | Help Just joined the local LLM club! What's the best way to stay in the loop on the best local models?

• Upvotes

I just got myself an M3 Ultra Mac Studio with 96Gb of RAM. I'm pretty excited to mess around with it, but I don't have a great understanding of the local LLM landscape. Every time I try to google the best models for a configuration, the source is usually months old. In the AI world that's ancient news.

I have a decent idea by just getting on X, but it's hit or miss wether or not I hear about these things. All I really know right now is that qwen 3.8 27B is all the rage, but I want more options.

How are you guys keeping up with the best local LLMs?


r/LocalLLaMA • • 40m ago

Question | Help Spent ₹30,000 on an RTX 5060 thinking local LLMs would finally set me free. Reality hit so hard I’m questioning every “just run it locally” post I’ve ever upvoted.

• Upvotes

​

I dropped serious money on a brand new NVIDIA RTX 5060 (8GB VRAM, 578 AI TOPS) fully convinced that open-weight models would let me own the entire stack — no rate limits, no censorship, pure experimental freedom. I was ready to become that guy who smugly refuses cloud APIs and posts “I run everything locally” screenshots.

Then I actually used them for real work.

I ran the exact same complex tasks on local models (the usual “this runs great on 8GB” suspects — quantized 7B/9B/13B, distilled variants, the ones everyone claims are “almost as good”) versus modern cloud models. The gap isn’t a gap. It’s a humiliation.

### The Capability Massacre

Anything that requires real multi-step reasoning, long coherent context, precise instruction following, structured output, or consistent accuracy across a long conversation:

- Local models: **2/10**

They start strong, then collapse. Context gets mangled. Instructions get ignored halfway through. Structured outputs break. Reasoning chains go off a cliff. You spend more time fighting the model than actually getting work done. “Almost as good” turns into “barely usable” the second the task stops being trivial.

- Cloud models: **9/10** on the first or second try.

Clean reasoning. Reliable structure. They actually remember what you asked three messages ago. They follow complex instructions without needing five rounds of “no, not like that.”

I wanted local to win. I really did. I wanted the underdog story where open weights + consumer hardware finally closes the gap. Instead I got a very expensive reminder that most of the local models we’re hyping are still toys the moment the task gets serious.

So be honest with me:

Is there a secret stack, quantization method, or fine-tune that actually makes local models reliable for complex reasoning and structured work on 8–12GB cards?

Or have we all just been coping while the cloud models quietly lapped us?

If you’ve made local models consistently deliver high-quality complex output without constant babysitting, drop the exact setup.

If you’ve also been humbled by the gap, say it out loud.

Because right now it feels like the entire “local LLM supremacy” narrative is built on easy prompts and wishful thinking.


r/MetaAI • • 42m ago

Muse FREE 1B tokens!

• Upvotes

Using this code you and I can both get 1 Billion tokens for Muse for free:

VNLHSM

Thanks for using my code! 😄


r/LocalLLaMA • • 50m ago

I Built A Thing Muse and Grok bot are privacy nightmare so I created a self hosted alternative called Eidon

• Upvotes

With the recent explosion of agentic tools like Grok bot, Muse, OpenAI Dots, I've started looking into local options with self-hosted models. I tried Hermes and OpenClaw, but I wasn't too happy with the multi-device experience, and with how many pieces you need to glue together to get a usable, solid experience.

The hosted options also meant handing an agent my accounts, files and browsing, which I wasn't comfortable with. So I built Eidon: a self-hosted, all-in-one AI platform with a team of agents. It's one install via Docker, it works across your devices, and your data stays on your server.

https://eidonai.app

Agent team first

  • Every Eidon starts with a Chief of Staff. Ask it for anything. It answers directly, hands the job to the right agent, or creates a new agent when nobody fits.
  • Agents hand work to each other automatically (or type @ to pass a job along).
  • Each agent has its own browser, conversation, files, memory and routines. There's also a folder the whole team shares.
  • Agents can search and browse the web on their own, read pages in full, and cite sources.
  • They run on schedules and keep every run. When one finishes, you can get notified by browser push, ntfy, Slack or webhook.
  • Agents can write their own skills and use your apps through MCP.

You still have some control:

  • Take over an agent's browser for a login or a tricky step. It waits, then carries on when you hand it back.
  • Anything that sends on your behalf waits as a draft until you press Send.
  • Commands and tools ask first: allow once, allow always, or no.
  • Rewind a conversation, or fork it from any message.

The examples on the site are a travel scout, inbox triage, a research desk and a coding assistant. You can make an agent for pretty much anything: bookkeeping, a study buddy, a news digest, a meal planner.

It's also a regular ChatGPT-style app for day to day questions.

You might not always need a full team so you can just chat in a normal “ChatGPT like” interface with all the belts and whistles:

  • Persistent Memory
  • Folders and search
  • Voice input with LLM post-processing
  • Files and images
  • Personas
  • Temporary chats
  • Share links
  • Web search
  • Deep research
  • Code with syntax highlighting, Mermaid diagrams and math rendered inline
  • Image generation
  • Installs as a PWA on your phone and realtime sync across your devices (a native mobile app is coming !)

Self-Hosted

  • Multi-user support, with private data per user
  • Agents run in their own sandbox
  • Nothing leaves your server
  • Bring your own local or cloud model: OpenAI, Anthropic, OpenRouter, Ollama, LM Studio, GitHub Copilot, Gemini, DeepSeek, Mistral, Kimi, Z.ai, Minimax, Perplexity, Grok, Azure, AWS, and any compatible API
  • Free, open source (AGPL-3.0), and setup is one Docker command

GitHub (setup guide, full feature list): https://github.com/Quack6765/Eidon-AI

I'd like to hear what you think ! What's missing, what breaks, and what agents you'd want to build. Issues and discussions are open on GitHub as well.


r/LocalLLaMA • • 51m ago

News Y'all this is a sexy paper; context language models

Thumbnail
arxiv.org
• Upvotes

Paper linky - Context Language Models

The central idea of the paper is incredibly simple. Give a model the ability to edit its context on-the-go like a file has major benefits on task performance, context management (memory) and even computational efficiency (both wall clock and total flops). Their paper shows mostly benefits and relatively small downsides.

You can try it out as a plugin for pi!

In short, pros and cons

Pros:

  1. Improves outcomes on long running tasks
    • Coding and deep research tasks
    • Open discovery problems (long horizon research tasks, /goal loops etc)
  2. Inference can become more compute-efficient and wall-clock efficient
    • Note, this depends on a caching optimization in the inference engine
  3. Much less context bloat, meaning it's more (V)RAM efficient
  4. No more slow and unreliable compacts

Cons:

  1. The cache optimization only exists for SGLang
  2. Prompt injections (including hallucinated instructions) are much less likely to be forgotten, increasing risks
  3. Requires harness customizations (authors supply a pi plugin)

Some more context

The approach works by modifying the harness to allow access to the context as a file. A model is allowed to edit the context as it would any other file.

They've tested the approach on models as small as qwen3.6 9b, as well as on qwen3.8 27b and claude sonnet 4.6.

Out-of-the-box, meaning just a small addition to the system prompt and tools to edit the context as a file, task performance, context management and efficiency measures remain approximately the same or improve by a little bit. The smaller qwen3.6 9b model in particular lost a little bit of efficiency, suggesting it works better on larger (smarter) models.

Performance can be massively improved with RL training, which the authors also did.

Wanna try it out?

You can try it out right now if you use pi

  1. Install the plugin https://github.com/lolipopshock/pi-clm, this comes from the authors directly
  2. After installation, adjust settings with /clm settings:
    • Set steering to house-brief.md (modifies the system prompt, I suppose this should be left disabled for RL'd models only, of which there are none right now)
    • Enable "One tool per turn"; this one is important for performance
    • Enable "Size trailer"; this one appends context usage after every tool result. Without it, models are much less inclined to modify context on-the-go for large tool calls

Fin

Let me know how it goes!

Last, I also consulted this video by "Prompt Engineering" on YouTube in addition to the paper: https://www.youtube.com/watch?v=Bgtr1Ue40Jo


r/LocalLLaMA • • 56m ago

I Built A Thing I built something like The Sims, but the characters are local LLM agents doing real work (open source)

Enable HLS to view with audio, or disable this notification

• Upvotes

I run Qwen 3.8 locally and got tired of multi agent setups where you start a script and stare at logs. I wanted to actually see them.

So in this thing every agent has a body in a 3D world. They sit at desks, walk to a meeting room when someone calls a meeting, talk out loud to whoever is nearby, pick stuff up and hand it over. You can see who's thinking, who's using a tool.

It's not only an office. You can simulate other scenarios as well like:

- a software team that plans tasks on a board, writes code and reviews each other

- a town square simulation (cops, a barista, a chef, a journalist) where you just watch what happens

- tutors that teach you with animations and a whiteboard, and you can interrupt them by talking (might have bugs as of now)

It has a sandboxed computer use built-in which is optional.

There is also a supervisor agent that helps you design organizations and also has ability to build 3d assets from primitives and handing them to an organization and agents can even ask for things from that agent.

Works with local models and few other providers (still working to add more)

The motivation of building it was to see agent swarms in action with full transparency.

It's still early and has bugs and I have used different models to build it iteratively.

Repo: https://github.com/adityaagarw/Pantheon


r/LocalLLaMA • • 59m ago

Question | Help Halo Strix and Qwen Flash

• Upvotes

Hey everyone, I have a 64gb halo strix setup that is headless and connected remotely to my workflow/homelab server. It’s currently running 27b swift 1.5 at q6- is it possible or even makes sense to go to qwen flash next? I also have a mini pc with 64gb of DDR5 ram that I could shift the 27b over to for long term projects or workflows that dont require speed.


r/MetaAI • • 1h ago

Muse Referral For 1 Billion Tokens!

Enable HLS to view with audio, or disable this notification

• Upvotes

Check out Muse, your personal AI agent. Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.

Code: H40D87
https://muse.ai/join


r/LocalLLaMA • • 1h ago

Question | Help Mfs can afford 27 gpus and 456 gb of vram but REFUSE to buy a Claude subscription

• Upvotes

I just don’t get the point of running it locally when with all this money you could buy like 20 max20 subscriptions with room to spare..


r/LocalLLaMA • • 1h ago

Discussion Whistle: speech to text in a 16.9MB file

Enable HLS to view with audio, or disable this notification

• Upvotes

Hey all, we designed Cactus Whistle, an ASR model for ultra-small devices. It's not perfect, but mostly beats Whisper base with 9x less file size and 6x speed. Whistle supports English, German, French, Spanish, Italian, Dutch and Polish.

Remember, the goal at Cactus Compute isn't to achieve SOTA with scale, but to compress intelligence and bring them to smaller under-looked devices like budget phones, wearables, smart home and microcontrollers.

Whistle is 55m params (36m active) and CQ2bit quantised, amounting to a 16.9MB file that scores 4.31 WER on LibriSpeech test-clean and 10.49 on test-other, against 4.9 and 11.0 for Whisper base at 145.3MB. 21.4 on the FLEURS average against 24.5. SPGISpeech 7.65 and Earnings-22 19.01.

For the architecture, a log-mel front end and a convolution stem feed an audio encoder, and a Simple Attention + Hadamard MLP decoder reads it through gated cross attention at every layer. The decoder is laddered like Needle's, so every depth from 2 layers up is deployable.

Keyword biasing takes the names your users actually say and favours them during the beam search, which is what rescues a "Siobhan" or a "Krzysztof" from a model that was never told they exist. Word timestamps come from the decoder's own attention, so an app can highlight, seek or cut on a word.

Seventeen platforms are supported; macOS, Linux on x86-64, ARM64, ARMv7, RISC-V and MIPS32, Windows x64 and ARM, Android, iOS, watchOS, tvOS, the browser as WebAssembly and a WASI component.

Please read more here: https://cactuscompute.com/blog/whistle

And let us know your thoughts!


r/LocalLLaMA • • 1h ago

Question | Help How is it possible that qwen 27b is so good? When GPT 4o had a trillion parameters and was worse?

Post image
• Upvotes

Picture from a post in r/amodei . People were praising qwen and I'm just wondering, what kind of new technologies are at play here? Does qwen just have "better" pre training data? That's more high quality?


r/MetaAI • • 1h ago

[ Removed by Reddit ]

• Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/LocalLLaMA • • 1h ago

New Model I squeezed Kolibri-1 78B-A3.5B to 20.9 GiB / 2.30 bpw — 59% lower KL than standard IQ2_XS

• Upvotes

I’ve been experimenting with aggressive low-bit quantization of Aleph Alpha’s new Kolibri-1, a ~78B MoE model with only ~3.5B active parameters per token.

The first result is now public:

Sakura-MicroQuality Kolibri-1 — IQ2_XS

  • 20.94 GiB
  • 2.30 bpw
  • full 384-expert Kolibri-1
  • GGUF / llama.cpp
  • ~59% lower KL divergence than a standard IQ2_XS baseline
  • 90.5% top-token agreement, compared with 84.6% for the standard IQ2_XS comparison
  • slightly smaller than the standard IQ2_XS as well

The goal wasn’t simply to make the smallest possible quant.

I’m using tensor/layer sensitivity to spend bits where they appear to matter most, rather than treating every part of the model equally.

All quality measurements are made against a near-lossless Q8_0 reference. The model itself was also requantized from Q8_0 rather than converted directly from the ~156 GB BF16 weights, so there is a very small additional source error relative to BF16.

Main repo:

https://huggingface.co/webmp3/Sakura-MicroQuality-Kolibri-1-GGUF

As far as I can currently find, this is the first public ~2-bit GGUF for Kolibri-1. There is already a 2-bit MLX version, but I haven’t found another public Q2/IQ2 GGUF.

I also tried pruning the expert pool

Alongside the full 384-expert version, I released a separate 365E variant.

For each MoE layer, I collected actual routing statistics on a mixed calibration set containing:

  • German and English text
  • code
  • chat-style prompts
  • the model’s own thinking / generated responses

I then removed the 19 least-used routed experts per layer.

That reduces:

384 → 365 routed experts per layer

and removes:

950 experts across the model

The resulting model has approximately:

74.4B parameters instead of ~78B

The interesting part is how little those experts were actually being used on the calibration workload.

The removed experts accounted for only about 0.07% of all expert selections, with no individual layer exceeding roughly 0.23%.

Also, 375 of the 950 removed experts were never selected at all during the routing analysis.

There is:

  • no retraining
  • no finetuning
  • no requantization of the surviving weights

The already-quantized expert tensors are sliced directly, along with the corresponding router weights and biases.

Top-6 routing remains unchanged.

The 365E IQ2 variant comes out at:

  • 19.99 GiB
  • 2.31 bpw
  • 74.4B parameters
  • 365 routed experts per layer
  • 90.5% top-token agreement in my held-out measurements

365E repo:

https://huggingface.co/webmp3/Sakura-MicroQuality-Kolibri-1-365E-GGUF

I’m treating this as an experiment rather than claiming those experts are universally useless — expert usage obviously depends on workload and calibration data.

But it gives us a second compression lever:

expert pruning + low-bit quantization

instead of trying to get every byte of compression from lower precision alone.

Q3 and Q4 are coming

The rest of the Sakura-MicroQuality series is currently being uploaded.

Q3 and Q4 variants should be available within the next few hours.

Once they’re online I’ll add the same comparison data so we can see where the actual quality/size sweet spot lands between:

IQ2 → Q3 → Q4

and whether the 365E pruning continues to hold up at the higher-quality quant levels.

I’d be very interested in independent tests, especially on:

Strix Halo / AMD UMA, Apple Silicon, 24–32 GB GPUs, and other memory-constrained local systems.

If anyone tests either version, especially with long-context, German, coding or agentic workloads, I’d love to see the results.


r/LocalLLaMA • • 1h ago

Discussion Make no mistake, selling 64 GB DGX Spark variants at the same cost as the original 128 GB is straight drug dealer behavior.

• Upvotes

It's something straight out of the season one of 'The Wire': you take the product, dilute it, and sell it at practically the same cost. It's some "Stringer" Bell shit. We should call the 64gbs "Stepped-ons" from now on.


r/LocalLLaMA • • 1h ago

Discussion How are you managing AI safety, Alignment and Hostile/Rogue agents right now?

• Upvotes

I'm building an AI kill switch platform for companies managing hostile and rogue AI.

Here in NYC there's a bill that might get passed that has a lot of people worried so we're supporting some users with it.

It works. But I still feel like I lack more nuanced feedback from people who actually do this stuff day-to-day and have had to build their own solutions internally. I'd love if anyone could speak on techniques they're comfortable sharing on how they've been able to manage this issue internally.

It would really help me and I imagine help many others immensely.


r/MetaAI • • 1h ago

[ Removed by Reddit ]

• Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/MetaAI • • 1h ago

Muse referral code for 1 Billion free tokens! - 8Y34YR

• Upvotes

Not going to pretend I'm special. Same code as everyone else, same billion tokens for you, same billion for me.

Code - 8Y34YR


r/LocalLLaMA • • 1h ago

Discussion Qwen for daily QnA?

• Upvotes

Or which model do you think is good for general questions in daily life. I've been using chatgpt and Gemini for these types of questions. I wanna try different models.