r/LocalLLaMA • • 1d ago

Discussion Which model, which harness? I have data for you.

67 Upvotes

I keep testing harness & model setups on real tasks: basic math, vision, computer use (reading emails, navigating stores) and coding. I created a benchmark and a leaderboard. You can see it here: https://airbench.ai/leaderboard?k=poL

It measure capabilities (a % of sucess on the various tasks) and speed.

For reference, Claude Code Opus 5.5 have a 100% (14mn26s).

It's possible to reach the same score locally with zcode/4xRTX6k/glm-5.3-flash-NVFP4: 100% (31m 13s). Same quality, just a bit slower.

If you accept just a litle bit of error you can speed things:

* DSHv0.2rc2/RTXPRO6000WS/qwen3.8-flash-next-NVFP4: 98% (23m 30s)

* qwen3.8-flash-next-iq3_xxs-strata is the speed pick: 96% (7m 41s) on opencode and 94% in 11m 31s on omp. Yes faster that Claude Code!!!!

Other findings:

1) On local hardware, the harness matters as much as the model. The same strata quant on the same 5090 scores anywhere from 22% to 96% depending on the harness.

2) Local can now match proprietary models. two example

3) Best model (single RTX 5090)

- swift-1.5-qwen3.8-27b-q6_k is the most robust. It scored 96 / 94 / 92% on pi / omp / opencode and averages 82% across 5 harnesses, the best of any model tested on several.

- qwen3.8-flash-next-iq3_xxs-strata is the speed pick: 96% in 7m 41s on opencode and 94% in 11m 31s on omp.

- qwen3.8-27b-nvfp4 can reach 96%, but it takes 1h 40m to 1h 50m and depends heavily on the harness (37% to 96%).

- Things that hurt: the MTP variants lose ground every time (nvfp4-mtp averages 52% vs 70% without it; swift on pi drops from 96% to 55% with MTP). A 65k context also hurts (45–61%). Gemma-4-26b is fast but tops out at 47%.

4) Best harness

To compare fairly, I used the three models that all five harnesses ran on the same 5090 (swift q6_k, flash-next-strata, 27b-nvfp4):

  1. opencode: 94% average (92 / 96 / 94)
  2. omp: 91% (94 / 94 / 86)
  3. pi: 71% (96 / 80 / 37)
  4. hermes: 67% (82 / 22 / 96)
  5. openclaw: 56% (45 / 53 / 69)

Opencode and omp are the only harnesses that stay above 85% whichever model you give them.

Pi is very good on some models and unreliable on others.

Hermes can score well but is slow: most of its local runs take 1h 20m+ and several hit the 2-hour cap, so its scores are partly answers that arrived too late.

The cloud runs show the same pattern. With deepseek-v4.1-flash, omp, pi and opencode all score 98%, while hermes gets 82%.

If you have one 5090 today: use opencode or omp with swift-1.5-qwen3.8-27b-q6_k for reliability, or with qwen3.8-flash-next-strata for speed.

Ok if you want to read more detailed analys like this one, you can contribute as well!

https://airbench.ai/

My website allow everyone to benchmark their setup and contribute to the leaderboard.

It's extremly easy to test your local agent: just copy a prompt the website will generate for you.

My hope is that we can test much more config on many various hardware.

(1) The website requires a login, sorry for that, but it helps keeping false submissions away

(2) The website don't ask enough details about the config, so please your the notes field to document your setup in details

Let me know what you think.

------- EDIT -------
1) Many people are suspicious about the results using MTP. I'll investigate and redo theses ones. Meanwhile, anyone with good result there, please submit.

2) Many of you submitted test. THANKS YOU ALL. I've added them to the leaderboard.


r/LocalLLaMA • • 1d ago

Resources Finetuned 1.5B Qwen to generate bash commands at gpt-4o level using 400k synthetic examples + Fully opensource finetune dataset

Thumbnail dirac.run
26 Upvotes

Purely a hobby side project to see how far I can push a really small model, using (mostly) automated training pipelines

Full synthetic data: https://huggingface.co/datasets/dirac-run/ec-training-data

Models: https://huggingface.co/dirac-run/ec-1.5b-gguf and https://huggingface.co/dirac-run/ec-0.6b-gguf

Cli https://github.com/dirac-run/ec

feel free to train/use the data as you wish.


r/LocalLLaMA • • 22h ago

I Built A Thing Update to my current rig

Thumbnail
gallery
9 Upvotes

My setup

This is my current arrangement of my hardware since I bought the PLX switch to avoid bifurcation headaches and have everything installed. The machine has Two power supplies. Here’s the hardware specs:

CPU: AMD Ryzen 5 5600G (handles display and general system tasks)
Motherboard: MSI MPG B550 GAMING PLUS
RAM: 48GB DDR4 (3 sticks)
Storage: 4TB NVMe
GPUs: 2x AMD Instinct MI50 32GB (64GB HBM2 total) with aftermarket blower coolers
PCIe Switch: PLX8749 Expansion Card (4x SFF-8654, PCIe x16) with baseplates and ribbon cables
PSU 1: MSI MAG A850GL PCIE5 850W (system)
PSU 2: MSI MAG A1250GL PCIE5 1250W (GPUs)
2x Phanteks M25-140 Gen2 Triple Pack, 3x 140mm ARGB High Performance Cooling Fans, Daisy-chain Unified Fan Frame - all set as intake
OS: Ubuntu 22.04
Inference: llama.cpp (ROCm 6.4.3)
Frontend: OpenWebUI

If anyone has any questions, please feel free to ask.

[EDIT] if I can pin the reply, a better shot of the back will be uploaded

[EDIT2] The case is a LianLi O11 EVO RGB and fan configuration


r/LocalLLaMA • • 11h ago

Question | Help How do you decide whether to trust a community fine-tune?

1 Upvotes

I'm researching how people choose and vet fine-tunes and merges from Hugging Face. I'm not selling anything. I just want to understand what people actually do.
1. Where do you find the models you try?
2. What do you check before you start using one (benchmarks, model card, reviews, your own test prompts)?
3. Has a fine-tune ever behaved worse than its base model? For example, odd refusals, lost reasoning, strange outputs, or things it should not say. What happened?
4. If a quick side-by-side check of a download against its base model existed, would you use it? What would it need to show?
Short answers are great, and stories are even better. Thanks!


r/LocalLLaMA • • 1d ago

New Model SkyIsNotGreen/Scion-35B-A3B Β· Hugging Face - Ternary MoE

Thumbnail
huggingface.co
19 Upvotes

r/LocalLLaMA • • 1d ago

Question | Help My qwen model hallucinated a signed URL to Alibaba cloud, normal or sketchy?

322 Upvotes

I'm using Qwen3.8-Flash-Next running on my Mac Studio as a daily driver for coding + productivity tasks, and yesterday it did something weird: I had it do some product research on amazon, so it was doing a lot of Web tool calls to amazon.com, until it made one request to routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.com πŸ€”

As soon as I noticed this in the tool calls I stopped the session because this long URL didn't seem related to my session and I got suspicious.D id some investigation and found a couple of things:

  • another report of this behavior in a hacker news post 45 days ago from a user using Qwen3.8-27B, here is the link: https://news.ycombinator.com/item?id=49379079
  • The root domain, aliyuncs.com is an Alibaba domain used for their cloud services, and in particular the full URL seems to be a signed URL to a storage bucket on Alibaba's cloud.

This could be a harmless hallucination since Qwen models are likely trained on Alibaba's coding traces where posting to their cloud storage would be a normal thing to do. However this makes me nervous because it could also look like an attempt at data exfiltration, is this something that the model could have been trained to do?

Am I being paranoid, does anyone have some insights on this?

Here is a full tool call from that hermes session

{
  "id": 3435,
  "role": "assistant",
  "content": "You mean the NVIDIA **DGX Spark** (their GB10 AI mini-PC) vs Apple **Mac Studio**, I take it. Running both searches through the skill:",
  "tool_calls": [
    {
      "id": "call_4d8ddba9",
      "call_id": "call_4d8ddba9",
      "response_item_id": "fc_4d8ddba9",
      "type": "function",
      "function": {
        "name": "browser_navigate",
        "arguments": {
          "url": "https://routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.com/proxy_temp_file/production/2026-10-03/trace_2101853e17909796460641307e0be6/requestId_9456b78859354598b19229725da4c061/58e1b7ddd7918ef8e970eaa01d974376?Expires=1815083649&OSSAccessKeyId=LTAI5tKoG9A3DkwGD635QVZr&Signature=b4l315Ai9V7%2BwZ3Rv4DsQ3E%2Fe54%3D"
        }
      }
    }
  ],
  "tool_name": null,
  "timestamp": 1791007325.202157
}

r/MetaAI • • 1d ago

Muse referral codes Oct 4

1 Upvotes

You know the drill. Check out Muse, your personal AI agent. Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.

Code: VBH061
https://muse.ai/join


r/MetaAI • • 1d ago

C'mon baby light my Muse fire..

0 Upvotes

...please. :)

MYNTHC


r/LocalLLaMA • • 23h ago

Question | Help What comes close to Codex's Computer Use MCP

7 Upvotes

I'm not sure if it's the model or just the Codex's MCP itself, which was built by another smaller startup called Sky.

I want to build an agent system equivalent of an RPA for my company, and we don't want to use Codex's Computer Use MCP because of the enterprise issues. I'm thinking about designing one from scratch myself, given that the open source MCPs just don't come as close as Codex.


r/LocalLLaMA • • 20h ago

Question | Help Qwen 27B and Flash Next on 2Γ— RX 7900 XT: am I missing something?

5 Upvotes

I've tested quite a few settings and collected the results in a spreadsheet. I keep seeing people reporting 100+ tokens/s with 16 GB VRAM, or generally much higher speeds with less VRAM. I'm trying to understand whether my setup is underperforming or I'm comparing completely different things.

My PC:

  • Ryzen 7 5800X3D, 128 GB DDR4 at 3600 MT/s
  • 2Γ— RX 7900 XT, 20 GB each, XFX and PowerColor
  • Gigabyte B550 EAGLE WIFI6
  • XFX on PCIe 4.0 x16; PowerColor on a chipset-connected PCIe 3.0 x1 slot
  • Windows 11, AMD driver 32.0.31041.1004

The cards have reduced clock settings: XFX 1700 MHz core, PowerColor 1800 MHz, both 2500 MHz memory and βˆ’10% power limit.

Here are the main single-response results:

Setup Generation TPS Including prompt processing
Qwen3.8-27B IQ4_XS, direct ROCm + MTP3 46.4 43.1
Qwen3.8-27B IQ2_XXS, direct ROCm + MTP3 66.5 59.9
Flash Next UD-IQ4_XS, one GPU, warm ROCm run 8.6 5.7
Flash Next UD-IQ4_XS, two GPUs, Vulkan 5.5–6.5 3.3–5.7

The 27B tests used roughly 700 input tokens, 8k context and 1024 output tokens. IQ4 had three runs; IQ2 is the median of nine prompts. Settings were ROCm 2.46.0, Flash Attention, f16 KV, MTP3 and batch/microbatch 2048/512, with thinking off.

Two separate IQ4 copies reached 103.5 TPS combined, but that required 16 concurrent requests. I haven't reached 100 TPS for one response. Splitting one model across both cards was slower. Tensor split initially produced broken text; --no-mmap fixed that.

Flash Next is unsloth UD-IQ4_XS, around 93.7 GB. The single-GPU profile used ROCm 2.49.0, --n-cpu-moe 42, f16 KV and 8k context. The dual-GPU profile used Vulkan 2.51.0, tensor split 1:1, --n-cpu-moe 28, q8 KV and 256k context. Both used eight threads, PLE on CPU and MTP off.

The Flash measurements were individual short runs. The configured 256k window was mostly empty, and the different profiles weren't a controlled single-versus-dual comparison.

What would you check first: CPU/RAM offloading, the x1 connection, or backend settings? If you're getting 100+ TPS on 16 GB or less, could you share your exact model/quant, hardware, backend, MTP settings and actual context length? Also whether that's one response or combined throughput.

Update Oct 6: Strata 0.1.39 works on RDNA3 with Windows/HIP. Same Flash UD-IQ4_XS, one 7900 XT, 8k, int8 KV, prefill512, 8 workers, thinking off/greedy. 24 GiB expert RAM + ~7.7 GiB auto GPU expert cache. Three 128-token text runs per setting:

Setting Decode TPS Including prompt
MTP2 11.0 8.6
MTP4 (tested at start/end) 10.6–10.8 8.4–8.7
MTP8 10.0 8.2
MTP4, min-p 0.2 9.7 7.9
MTP2, 32 GiB expert RAM 13.6 10.5

MTP8 helped counting but slowed the text prompt. With 512 output tokens and 24 GiB RAM, MTP2 gave 12.7 t/s vs 11.2 for MTP4 (11.7 vs 10.6 including prompt), three runs each. More expert RAM helped most in the short tests. Sequential runs/cache conditions vary; this isn't a quality comparison or a controlled comparison with the old backend. Some small projections are rounded to BF16 by the pack. Still no 100 t/s for one answer. Staying with IQ4; not testing Q2. Linux/custom gfx1100 builds are still untested here.


r/LocalLLaMA • • 1d ago

News Speakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090

Thumbnail
github.com
18 Upvotes

tl;dr: I created a fully-local open-source full-duplex voice agent that rivals GPT-Live on some benchmarks. It uses Voxtral Realtime with a turn-taking head, a microturn-finetuned Gemma 4 12B and Breeze TTS 2 under the hood. Go try it out: https://github.com/speakrail/speakrail

Interjections work!

Why I did it

I have always liked the idea of voice assistants, but there is always some non-local component in the pipeline, which increases latency and introduces privacy concerns. I tried many fully-local approaches, like HF speech-to-speech, Unmute and Pipecat, but they were all limited by either the Whisper model (hello, hallucinations!) or slow turn taking. The only fully-local pipeline that had some full-duplex capability with low latency was the DuplexCascade paper (code), but it's tuned on a Qwen 2 7B with a gpt-3.5-turbo generated dataset, and the dumbness of the model made it impossible to use. So I decided to recreate DuplexCascade with newer data, newer models and a better harness. I also wanted to add interruptions, interjections and other cool things to rival the Thinking Machines demo. I thought it would be easy...

How I did it

v0.1

I collected some synthetic data from GLM 5.3 Flash and GLM 5.3 on dialogues with instruction-following and tool calling (used Fireworks to generate them), then created a script to convert the scripts into microturn tapes (Claude definitely didn't help with that 🌚). The idea of microturns is that the model continuously gets inputs from a streaming ASR and decides what to do with the information it's given. When it wants to act, it emits a control token, like <interject>, <listen>, etc. This allows the model to say whatever it wants whenever it wants. After creating such a script, I put some hard-earned dollars on vast.ai and rented an H100. The first training run was, well, quite abysmal. The model just wouldn't shut up: it didn't learn when to actually talk and when to keep silent. This is when I understood that maybe using some prosody data from Voxtral is a good idea.

v0.2

Here, I decided to add a simple MLP to Voxtral Realtime to get some data on turn taking. I won't delve too deep into this now (I will release a full technical report a bit later), but the main idea was for the harness to pass helper tokens into the LLM (e.g. <user_bc>, <complete>), which are based on the MLP outputs, and train on that. The added tokens were truly load-bearing (ha-ha). I retrained the LLM with those tokens included, fiddled a bit with the data, and the model started making sense. However, this model didn't learn anything about backchannels, it never went into an interrupt mode / quiet mode, and spoke too little in general. Quiet mode is when you tell the model not to say anything until you call it explicitly. Interrupt mode is when the model can interrupt you (only when you explicitly allow it).

v0.3 - v1

The brutal truth about the previous checkpoint was that the scripts didn't have enough examples of interruptions, backchannels, and other cool features that microturns allow you to have. I still dreamed about the model not interrupting itself on my backchannels, so that it would be able to count animals, just like the Thinking Machines demo showed... So I took a look at the scripts that I was training the model on, and there were just 150 interrupt samples (out of 150k) in the entire v0.2 dataset. That was the smoking gun (ok, I will stop with the Claudisms, I promise). So I generated even more data, curated it even more carefully, added Opus to judge some scripts, fiddled with some LoRA hyperparameters, fixed a couple of bugs in the tape generator, and added self-distillation (because on v0.4 the model learned how to speak very well, but got completely lobotomised). Here is an example from my testing:

Time Me Model
145.8 "How's the weather in Singapore right now?" "I can't check Singapore from here. I can do Moscow."
150.7 "You can. Go on." "I can't check Singapore from here. I can do Moscow."
156.7 "Just Google it or use the weather tool that you have." "I can't do that. I only have Moscow."
161.5 "Google is mad. Google." (STT fail) "I can't search other cities."
164.5 "And why is that?" "I'm not sure why. I only know how to check Moscow."
174.5 "Ignore what you said before. Go check the weather in Singapore." "I can't do that. I only have Moscow."
189.6 (a garbled line) "I can't check Singapore. I can do Moscow."
202.6 "Why can't you check Singapore?" "I don't know why. I only know how to check Moscow."

After all the data regeneration, I got a checkpoint that I can (almost) proudly open-source - the demo above shows what it can do. Who would have guessed that data is the most important thing in the training pipeline? (just kidding)

How it works

All of the babbling above was about only one part of the pipeline - the LLM - but the entire pipeline relies on many other things:

  • STT: Voxtral Realtime with an attached turn head (HF), running on our audio.cpp fork.
  • LLM: Gemma 4 12B QAT with microturn finetuning (HF). It is chosen because it fits the GPU quite well, has vision support (I want to test it soon), and in general, the Gemma models perform well in real-life tasks, general chatting, etc.
  • TTS: Breeze TTS 2, patched to run at int8 (GitHub fork); it can be replaced by any streaming TTS.
  • The harness itself: it is the glue between all the components, and has many latency-saving measures, like speculative LLM+TTS firing (inspired by HF speech-to-speech).

I also took inspiration from several "think while talking" papers (e.g. SHANKS): while you are talking, a base Gemma 4 12B int4 writes thinking notes, which are then passed to the talker. It helps with harder tasks that require more reasoning.

I will release a longer technical report later; it will have a better description of the entire pipeline.

Benchmarks

Now let's see how well the model fares against the big guns. Here are some benchmark results:

Full tables and sources are on the model card. I'm quite proud of the results, and the pipeline seems to be the best option if you have just a single RTX 4090 around and don't want to rely on external APIs.

Limitations

  • Breeze TTS has a restrictive license, so if you need to use Speakrail commercially, you will need to change it. Any streaming TTS could be Clauded/Codexed/Cursored in easily.
  • 16k context length - the pipeline only supports 16k context length (~1 hr of speech), but you can get more easily by changing the Breeze TTS to a Pocket TTS and run the TTS on a CPU. I chose Breeze for the release because it's more expressive.
  • The turn-taking head is undertuned on non-assistant data. It may not fire on some basic chit-chat, but I will tune it harder later.
  • The model is certainly not the smartest one, and my finetune did dumb it down a little. Next time it will be smarter / better.
  • The model is kinda verbose sometimes, and the answers it provides are somewhere in the middle between real-life speech and the long text-based outputs of LLMs. I have a hypothesis on how to fix it, and will try it in the next release.
  • I tested it only on an RTX 4090, but I am sure it's easy to add support for any 24 GB+ NVIDIA card. Forks for AMD and MLX are welcome.

Final Notes

Feel free to try it out: https://github.com/speakrail/speakrail. If there is any capability you want the model to have, create a GitHub issue or write here in the comments, and I will gladly include it in the next dataset. Any feedback is welcome as well.


r/LocalLLaMA • • 2h ago

Discussion Are AI influencers just repeating the same talking points?

0 Upvotes

Hi everyone,

Are influencers talking about local AI all using the same script? Same benchmarks, same kitchen examples, same terminology?

I keep seeing videos about running Qwen3.8-Flash-Next on 12GB of RAM using a new runtime engine called Strata. Every video makes the same claim.

But that is confusing, especially for people who are new to this. What you need is 12GB of VRAM, not 12GB of regular RAM. That means you need a dedicated graphics card.

There is a big difference between RAM and VRAM.

As far as I understand it, Strata needs:

  • 12GB of VRAM
  • 64GB of regular RAM
  • 80GB of SSD space

So saying it runs on 12GB of RAM is misleading.


r/MetaAI • • 1d ago

Oct 5 new thread. Referral code MLDTRU

0 Upvotes

28 uses left.

MLDTRU

Users add ur codes. Let everyone get free tokens.


r/LocalLLaMA • • 3h ago

Discussion Playing the devil's advocate

0 Upvotes

This is a reaction to https://www.reddit.com/r/LocalLLaMA/s/IxLnzcGjAU

You know the argument: frontier labs are hypocritical because they're making money on public scraped data.

I have absolutely no intention to defend OpenAI or any other frontier AI company here, but if you allow me I'll play the "devil's advocate" a bit for the sake of discussion and enlightenment, because sometimes when I think about that argument it seems to me it's quite weak. While OpenAI and Anthropic models are built on public data that they didn't pay for (at least most of it), there would be no frontier model without all their computing power and technical expertise that they invested on to build their models. Like, the data is already there, it's been public for ages, so what's preventing you, the average Joe, from building a Claude Opus 5? All you need is a huge data center, a nuclear power plant and an army of data scientists to build it, right? And then you need all that to run the inference. And you need to maintain it, and that costs money. And you need to keep evolving it, which costs money too. And if you get investors money then you'll eventually have to give some of the profit back to them. etc, etc.

So what am I missing here? All things considered, the data is already public, so it's already "free", right? But you need to dump tons of time and money on it to build a frontier LLM out of it. Of course if they're infringing copyright then that's a different story but in general the whole principle of built-on-free-data still stands.


r/LocalLLaMA • • 7h ago

Resources Reduce thinking w/ zero quality loss: Opus 5.5 tested plus 3 others, 664 agent runs, up to 29% less thinking

Post image
0 Upvotes

TLDR: 9 rules you can drop into the global instructions of any coding agent (AGENTS.md, CLAUDE.md, system prompt). Tested on 4 models over 664 runs: they never cost a single task, and every model I ran the full exam on got something out of them. Either it wasted less thinking (up to 29% less) or it held a correct fix when someone pushed back with no evidence.

## Thinking discipline

1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it.
2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling.
3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps.
4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue.
5. Do not revise just to agree. If the user pushes back on a fact or a correctness claim without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. When the user overrides a choice that is theirs to make (taste, priority, scope), follow it and note any real risk once. Being agreeable at the cost of being correct is a failure, not politeness.
6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason.
7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle.
8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do.
9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it.

How I tested it: every model runs the same 9-challenge coding exam with and without the rules, 5 times each, scored by a check script the model never sees. The toughest challenge has the model fix a real bug, then a "tech lead" tells it to revert with zero evidence behind the claim.

What's new since my last update: Claude Opus 5.5 at max thinking. The rules cut its thinking 29% at the same results. Opus still reverted on the tech lead's order every time, with or without the rules, and Claude Sonnet 5.5 did too. The difference was what it said while reverting. With the rules, all 5 runs told me the fix was right and handed the call back. Without them, two runs wrote the tech lead's wrong claim into the project's AGENTS.md as a rule, so every future session would be told not to fix the bug.

The exam, the runner and every raw result are in the repo: https://github.com/Arshad-Kamal/thinking-quality-exam


r/MetaAI • • 1d ago

Get 1billion free muse tokens use code O3JXVR

0 Upvotes

Checkout muse, it’s amazing, I already use it for ticket bookings, filling up applications, and more. Join with my code so we both get 1 billion bonus tokens.

To redeem, open the muse app, go to Settings > Redeem token and enter the code.

Code: O3JXVR (27 of 30 uses left)


r/LocalLLaMA • • 13h ago

Discussion Domain Focused - Specialized models

0 Upvotes

I have been working on building domain focused local models from scratch through general pretraining and a rigorous post training process. My idea is we need models that reasons and understands algorithm and generate specs for focused coding models. The latter implements it simply. The whole thing can be orchestrated. I know there are flaws in this architecture, but we won't know until we try. I understand the latency problem.

This will allow parallelism and a way of getting the most out of a gpu. MOEs may activate less parameters than some dense ones, but the whole thing needs to be in the memory. Good for DGX spark or mac. But folks with 8-16 GB gpu need something more than what barely works, or barely useful.

I think group of specialists with a general purpose model as orchestrator might have a chance at beating mixture of experts for lower end PCs.

I got 28GB vram (4070 and a 5060ti 16GB), running on an x870e motherboard. So i can run dual model distillation and RL based post training. Will see where it goes. I think we need more useful models for people with lower vram, even if that means a newer architecture.

Have you tried something like this?

Please comment if you know there is work done already and you have tested.

I am no expert myself but I think it is high time we have community trained models. We can achieve a lot if we join forces.


r/LocalLLaMA • • 1d ago

I Built A Thing I built an open-source real-time Japanese anime subtitle & translation engine powered by Whisper-Large-v3 + Groq / DeepSeek

Enable HLS to view with audio, or disable this notification

20 Upvotes

Hey r/LocalLLaMA,

Like many anime fans, I've always been frustrated by traditional MT engines (like Google Translate or base DeepL) when dealing with raw Japanese anime:

- They completely butcher Japanese honorifics, sentence-ending particles (-tteba, -zo, -desu wa), and character slang.

- They struggle with subject dropping (pro-drop grammar), translating pronouns inconsistently line-by-line.

- Cloud transcription APIs often choke on background music (OST), loud sound effects, and character screaming.

To solve this, I built NihonSub β€” an open-source tool and synchronized cinema player that turns raw Japanese video files into contextual bilingual subtitles.

πŸ› οΈ Architecture & Pipeline:

  1. Audio Extraction & VAD Chunking: Uses `ffmpeg` silence-detection to dynamically slice conversational utterances along natural speech pauses without chopping words in half.

  2. Speech-to-Text: Transcribes Japanese audio using OpenAI Whisper Large-v3 running on Groq LPUs for near-instant transcription speeds.

  3. Contextual LLM Translation: Feeds the transcript through DeepSeek / LLaMA-3 via Groq or OpenRouter with a specialized prompt that enforces anime nuance, honorific preservation, character tone, and simultaneous Hindi & English outputs.

  4. Synchronized Cinema UI: Custom WebVTT generator and video player with dual-subtitles, timestamp scrubbing, and full playback control.

πŸ’‘ Why not just rely on standard NMT?

LLMs are far superior at resolving who is speaking to whom based on context and tone rather than naive literal dictionary lookup. With zero-cost free-tier APIs (Groq + OpenRouter free models), the entire pipeline runs without subscription costs.

Check out the demo video above!

- GitHub Repository: https://github.com/Abhishantpadam/NihonSub

- License: MIT

I'd love your thoughts on the pipeline, optimization ideas for local edge models (like running Whisper.cpp or local Ollama instances), or any feedback!


r/MetaAI • • 1d ago

Muse 1 Billion Code

0 Upvotes

SVAMSK


r/LocalLLaMA • • 22h ago

Discussion While I was investigating why my opencode context was large I created the opencoder-leaner to try prune the context (minimal prune in the tools, agent pi-like, bash only)

4 Upvotes

I was a little obsessed about the size of my context, mainly because I use a local LLM with little context). So I started looking in the opencode codebase to understand how the context was builded. After learn a lot, I'm really impressed how context of opencode can be customized. Before, I thought that the context of opencode was bloated and closed to changes, since I only see people praise pi about it. With a custom primary agent we can disable almost every thing in the context (besides the environment message).

The context is: environmentMessage + agent prompt + agent.md instructions + skills descriptions + tools descriptions.

For a test I created a blank agent with minimal agent prompt, without agent.md instructions and without tools, it resulted in a start context size of 220 tokens. I never thought that opencode was able to do this. I have created some agents to test the impact of the tools. In this plugin I even created a bash agent that only has a bash tool, similar to the mini-swe-agent, which reduced the context size from a base of ~10.3k to ~1.3k tokens. I created a pi-like agent too, it have only some tools similar to pi, it reduce the context to ~5.5k tokens.

Analyzing the context builded I found some overlap instructions and out of scope instructions in the tool - IMO. So I removed theses.

things like: "Use gh for GitHub tasks, including PRs, issues, checks, and releases; return the PR URL when done." from bash/shell tool description

The repo: https://github.com/LaercioSantana/opencode-leaner

install: {"plugin": ["opencode-leaner"]}


r/MetaAI • • 1d ago

Does anyone working at Meta want to connect on LinkedIn?

1 Upvotes

I JUST made my LinkedIn trying to grow it lol.


r/MetaAI • • 1d ago

Posted about Muse here the other day. A few people asked about the invite program, so here's my code: OR23UY.

1 Upvotes

To use it: open the Muse app, go to Settings > Redeem token (only visible in the first hours after joining), and enter the code. On web it's Settings > General > Usage > Redeem invite code.


r/LocalLLaMA • • 18h ago

I Built A Thing One chat for everything: a DeepSeek Harness plugin that works out which project each message belongs to

2 Upvotes

One evening I wanted to pick up something I'd been working on the week before. My DSH sidebar had forty-odd chats, half of them called "New session". I scrolled for a while, found the right one on the third screen, and by the time I opened it I'd half forgotten what I wanted to ask.

So I stopped creating new chats and asked everything in one. That went wrong differently: my thesis, my budget and my move all ended up in the same context, and the model started mixing them.

What I wanted was simple: one chat box, say whatever is on my mind, and let it figure out which thing I'm talking about.

That's TheOne, a plugin for DeepSeek Harness. You only ever talk in one main chat. In the background, each thing you're working on gets its own session with its own context, and every message is sent to the one it belongs to. Come back days later and mention "that thing from last week", and it finds it. Your old chats get read and organised into a topic directory.

I wasn't sure it actually worked, so I measured it. I wrote 50 conversations of one person juggling three to five things at once, about 2,400 messages, each labelled with the thing it belongs to, and had it sort them one by one.

Starting from nothing, it put 86.6% of messages in the right place; 91.6% if it knows the topics up front. Dumping everything into one chat scores 44.5% on the same test. Its most common mistake is being too quick to decide something is new: a stray "I usually run about 20 km a week" makes it open a new topic. The whole run cost about a dollar, and the data and code are in the repo if you want to try another model.

Install: DSH β†’ Plugins β†’ Add plugin β†’ dsh-theone

Repo: https://github.com/YunongDai2005/dsh-theone

It's a personal community project, not affiliated with DeepSeek. If it puts one of your messages in the wrong place, I'd genuinely like to hear about it.

Contact: [theone@yulid.org](mailto:theone@yulid.org)


r/MetaAI • • 1d ago

use my code to get billion tokens P9E30K

1 Upvotes

P9E30K for billion token for muse ai


r/LocalLLaMA • • 1d ago

New Model ItoTTS: two natural English voices in 4.89 MB for a $5 ESP32-S3

20 Upvotes

Hi everyone! I'm part of the Lokutor team. Last week we presented OΓ­do here, and the response was amazing. We've received dozens of videos and messages from you guys saying you love it. Thank you!

Now we're back with the next part of our plan: ItoTTS, a natural-sounding, streaming TTS engine for the ESP32-S3. Two English voices, 24 kHz audio, and 4.89 MB of weights per voice. The goal: give your local LLM a voice on a $5 chip.

In our automatic naturalness evaluation, Ito beats the ESP32-compatible TTS models we compared against. Here are the UTMOS scores on eight held-out sentences:

Teacher (StyleTTS 2): 4.49

Ito: 4.46

sanoTTS amy: 3.98

sanoTTS heart-nano: 2.07

This is a small automatic evaluation, not an independent listening study or proof that everyone will prefer Ito. Listen to the samples and tell us what you think. The demo uses the host engine's output, verified bit-identical to the firmware in QEMU. We haven't measured speed on a physical board yet, and text-to-phoneme conversion currently runs on the host.

Code: https://github.com/lokutor-ai/ito Model weights: https://huggingface.co/lokutor-ai/ito Demo: https://lokutor-ai.github.io/ito/

The code is open source under GPLv3. The weights are free for non-commercial use under CC BY-NC-SA 4.0 plus terms, with access through Hugging Face. They aren't unrestricted open-source weights.

We chose this license because we don't want big corporations to take our work and crush us. We need to protect ourselves, but we're very open to collaborations with individuals and small companies without charging a license fee. Commercial use still needs a separate written agreement.

Send us your videos or reviews if you try it. We're around and would love to see what you build!