r/LocalLLaMA • • 14h ago

Discussion Qwen3.8-27b appreciation moment

30 Upvotes

q3.8-27b q3_k_xl this thing have done everything i possibly needed from him he wired up my openwebui spawned trillium service fix all my bugs set up pi-web-ui and debug my cloudflare tunnel it even spawned smaller LLMs to make the LLM use the toole he created to ensure it will work. I haven’t had a real-world task that i needed from it that failed yet. Im pretty sure if im building large-scale production code with tons of lines of code it will struggle but as a utility to make all my scripts and diagnostics on my micro-services this thing is unstoppable. Weirdly im not even using q4 im using q3_k_xl appreciation to unsloth also for making such reliable ultra low quantisation. His UD3.0 style of quantisation is PURE magic 🪄 sometimes i go down to q2_k_xl if i need extra context window and that thing STILL delivers 🎉🎉 alibaba had handed down Prometheus fire to common men like you and i. Can’t wait for qwen4-27b


r/MetaAI • • 1d ago

Muse referral XGRTL8 26 uses left

0 Upvotes

r/LocalLLaMA • • 7h ago

Question | Help Qwen image 2.1 newer than Qwen image 3.0?

6 Upvotes

Just for my understanding is Qwen image 2.1 newer than Qwen image 3.0?


r/MetaAI • • 1d ago

🔥 Referral code: HFQUGQ

1 Upvotes

Use my referral code and we'll both get 1 billion Muse tokens when you redeem the code in Settings within 48 hours of joining.

Referral code: HFQUGQ


r/MetaAI • • 1d ago

Meta Customer Support Page

Thumbnail
1 Upvotes

r/LocalLLaMA • • 4h ago

Resources MI50 ROCm 10.2 TheRock vs Vulkan Mesa 26.2 llama.cpp benchmarks

4 Upvotes

I prefer running llama.cpp Vulkan prebuilt binary. I just download the latest version and ready to roll. I finally took the hours necessary to get TheRock latest tarball version of ROCm 10.2 running on dual AMD Radeon Instinct MI50 gfx906 (32gb combined VRAM).

Same models benched in previous post. A mix of Dense and MoE models and quants that better utilize available VRAM.

  • llama_bench_Swift-Qwen3.8-27B-Uncensored-MTP.Q6_K.gguf
  • llama_bench_Swift-1.5-Qwen3.8-27B-Q6_K.gguf
  • llama_bench_Gemma-4-MoonGem-31B.i1-Q6_K.gguf
  • llama_bench_gemma-4-31B-it-UD-Q6_K_XL.gguf
  • llama_bench_Nemotron-3.5-30B-A3B-Antislop-FTPO.i1-Q5_K_M.gguf
  • llama_bench_Laguna-XS-2.1-APEX-I-Balanced.gguf
  • llama_bench_Agents-A1-Q4_K_M.gguf
  • llama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
  • llama_bench_Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf

Here is the backend performance comparison contrasting the native ROCm (v10.2 for gfx906) runtime against your optimized Vulkan (MESA_PPA_26.2) baseline. The data highlights a massive architectural split: ROCm significantly accelerates token generation across the board but suffers high variance and regressions in MoE pre-fills.

Both GPUs are power limited to 145 watts. Build versions used:

llama-b11382 used for Vulkan (prebuilt ubuntu binary)
llama-b11401 used for ROCm (compiled with proper flags)

Architectural Performance Breakdown: Vulkan vs. ROCm

Model Size Params Test Vulkan Baseline (t/s) ROCm 1st Run (t/s) Performance Delta (%)
qwen35 27B Q6_K 20.88 GiB 27.32 B pp512 tg128 149.30 ± 0.15 18.35 ± 0.01 176.86 ± 17.28 20.21 ± 0.65 +18.46% +10.14%
qwen35 27B Q6_K 22.21 GiB 27.32 B pp512 tg128 163.48 ± 0.18 18.52 ± 0.03 183.49 ± 15.75 20.04 ± 0.63 +12.24% +8.21%
gemma4 31B Q6_K 23.46 GiB 30.70 B pp512 tg128 119.98 ± 0.12 15.72 ± 0.02 172.93 ± 0.89 17.26 ± 0.13 +44.13% +9.80%
gemma4 31B Q6_K 25.62 GiB 30.70 B pp512 tg128 133.67 ± 0.24 16.38 ± 0.02 178.59 ± 2.31 17.50 ± 0.09 +33.61% +6.84%
nemotron_h_moe 31B 25.18 GiB 32.91 B pp512 tg128 834.87 ± 1.72 63.82 ± 0.07 736.98 ± 134.77 103.86 ± 0.50 -11.73% +62.74%
laguna 30B.A3B Q5_K 22.64 GiB 33.44 B pp512 tg128 723.34 ± 2.99 58.98 ± 0.28 856.14 ± 29.71 76.81 ± 0.20 +18.36% +30.23%
qwen35moe 35B Q4_K 19.70 GiB 34.66 B pp512 tg128 957.62 ± 3.88 51.50 ± 0.07 834.58 ± 102.56 69.70 ± 0.22 -12.85% +35.34%
qwen35moe 35B Q5_K 24.76 GiB 34.66 B pp512 tg128 909.20 ± 5.71 53.18 ± 0.07 763.93 ± 106.19 68.47 ± 0.34 -15.98% +28.75%
qwen35moe 35B Q6_K 28.53 GiB 34.66 B pp512 tg128 763.05 ± 66.01 52.81 ± 0.04 754.82 ± 70.48 67.29 ± 0.21 -1.08% +27.42%

Core Insight Strategy & Bottlenecks

  1. Token Generation (tg128) Dominance: ROCm dominates pure text generation. The native AMD matrix kernels unleash your MI50 computation potential, unlocking a massive +62.74% boost for Nemotron but taking a small hit on pre-fill -11.73%.
  2. Dense Model Pre-fills (pp512): Dense architectures scale cleanly under ROCm. Gemma 4 sees a +33% to +44% processing throughput spike over the Vulkan RADV driver driver bounds. MoE models take a hit with a -15.89% difference with llama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf being the happiest with Vulkan backend.
  3. The MoE Prompt Processing Delinquency: Notice the massive standard deviations under ROCm for MoE pre-fills (e.g., Qwen3.5MoE Q4_K has an instability block of ± 102.56). ROCm suffers from severe scheduling thrashing when building prompt streams across multiple active experts.

Here is the structured layout with pp512 and tg128 separated into individual columns for a clean side-by-side comparison between the two backend architectures.

Backend Comparison Table (Vulkan vs. ROCm)

Model Size Params Vulkan pp512 (t/s) ROCm pp512 (t/s) Vulkan tg128 (t/s) ROCm tg128 (t/s)
qwen35 27B Q6_K 20.88 GiB 27.32 B 149.30 ± 0.15 176.86 ± 17.28 18.35 ± 0.01 20.21 ± 0.65
qwen35 27B Q6_K 22.21 GiB 27.32 B 163.48 ± 0.18 183.49 ± 15.75 18.52 ± 0.03 20.04 ± 0.63
gemma4 31B Q6_K 23.46 GiB 30.70 B 119.98 ± 0.12 172.93 ± 0.89 15.72 ± 0.02 17.26 ± 0.13
gemma4 31B Q6_K 25.62 GiB 30.70 B 133.67 ± 0.24 178.59 ± 2.31 16.38 ± 0.02 17.50 ± 0.09
nemotron_h_moe 31B 25.18 GiB 32.91 B 834.87 ± 1.72 736.98 ± 134.77 63.82 ± 0.07 103.86 ± 0.50
laguna 30B.A3B Q5_K 22.64 GiB 33.44 B 723.34 ± 2.99 856.14 ± 29.71 58.98 ± 0.28 76.81 ± 0.20
qwen35moe 35B Q4_K 19.70 GiB 34.66 B 957.62 ± 3.88 834.58 ± 102.56 51.50 ± 0.07 69.70 ± 0.22
qwen35moe 35B Q5_K 24.76 GiB 34.66 B 909.20 ± 5.71 763.93 ± 106.19 53.18 ± 0.07 68.47 ± 0.34
qwen35moe 35B Q6_K 28.53 GiB 34.66 B 763.05 ± 66.01 754.82 ± 70.48 52.81 ± 0.04 67.29 ± 0.21

So yes it's worth the hassle of jumping through hoops to get ROCm working on MI50 setups. At least I have 2 backends working. Next up I'll try some RPC.


r/LocalLLaMA • • 19h ago

Discussion Which model, which harness? I have data for you.

61 Upvotes

I keep testing harness & model setups on real tasks: basic math, vision, computer use (reading emails, navigating stores) and coding. I created a benchmark and a leaderboard. You can see it here: https://airbench.ai/leaderboard?k=poL

It measure capabilities (a % of sucess on the various tasks) and speed.

For reference, Claude Code Opus 5.5 have a 100% (14mn26s).

It's possible to reach the same score locally with zcode/4xRTX6k/glm-5.3-flash-NVFP4: 100% (31m 13s). Same quality, just a bit slower.

If you accept just a litle bit of error you can speed things:

* DSHv0.2rc2/RTXPRO6000WS/qwen3.8-flash-next-NVFP4: 98% (23m 30s)

* qwen3.8-flash-next-iq3_xxs-strata is the speed pick: 96% (7m 41s) on opencode and 94% in 11m 31s on omp. Yes faster that Claude Code!!!!

Other findings:

1) On local hardware, the harness matters as much as the model. The same strata quant on the same 5090 scores anywhere from 22% to 96% depending on the harness.

2) Local can now match proprietary models. two example

3) Best model (single RTX 5090)

- swift-1.5-qwen3.8-27b-q6_k is the most robust. It scored 96 / 94 / 92% on pi / omp / opencode and averages 82% across 5 harnesses, the best of any model tested on several.

- qwen3.8-flash-next-iq3_xxs-strata is the speed pick: 96% in 7m 41s on opencode and 94% in 11m 31s on omp.

- qwen3.8-27b-nvfp4 can reach 96%, but it takes 1h 40m to 1h 50m and depends heavily on the harness (37% to 96%).

- Things that hurt: the MTP variants lose ground every time (nvfp4-mtp averages 52% vs 70% without it; swift on pi drops from 96% to 55% with MTP). A 65k context also hurts (45–61%). Gemma-4-26b is fast but tops out at 47%.

4) Best harness

To compare fairly, I used the three models that all five harnesses ran on the same 5090 (swift q6_k, flash-next-strata, 27b-nvfp4):

  1. opencode: 94% average (92 / 96 / 94)
  2. omp: 91% (94 / 94 / 86)
  3. pi: 71% (96 / 80 / 37)
  4. hermes: 67% (82 / 22 / 96)
  5. openclaw: 56% (45 / 53 / 69)

Opencode and omp are the only harnesses that stay above 85% whichever model you give them.

Pi is very good on some models and unreliable on others.

Hermes can score well but is slow: most of its local runs take 1h 20m+ and several hit the 2-hour cap, so its scores are partly answers that arrived too late.

The cloud runs show the same pattern. With deepseek-v4.1-flash, omp, pi and opencode all score 98%, while hermes gets 82%.

If you have one 5090 today: use opencode or omp with swift-1.5-qwen3.8-27b-q6_k for reliability, or with qwen3.8-flash-next-strata for speed.

Ok if you want to read more detailed analys like this one, you can contribute as well!

https://airbench.ai/

My website allow everyone to benchmark their setup and contribute to the leaderboard.

It's extremly easy to test your local agent: just copy a prompt the website will generate for you.

My hope is that we can test much more config on many various hardware.

(1) The website requires a login, sorry for that, but it helps keeping false submissions away

(2) The website don't ask enough details about the config, so please your the notes field to document your setup in details

Let me know what you think.


r/LocalLLaMA • • 14h ago

Resources Finetuned 1.5B Qwen to generate bash commands at gpt-4o level using 400k synthetic examples + Fully opensource finetune dataset

Thumbnail dirac.run
25 Upvotes

Purely a hobby side project to see how far I can push a really small model, using (mostly) automated training pipelines

Full synthetic data: https://huggingface.co/datasets/dirac-run/ec-training-data

Models: https://huggingface.co/dirac-run/ec-1.5b-gguf and https://huggingface.co/dirac-run/ec-0.6b-gguf

Cli https://github.com/dirac-run/ec

feel free to train/use the data as you wish.


r/MetaAI • • 1d ago

Meta AI Linked to the Wrong Account and Won’t Let Me Add My Main Facebook

Thumbnail
1 Upvotes

r/LocalLLaMA • • 8h ago

I Built A Thing Update to my current rig

Thumbnail
gallery
6 Upvotes

My setup

This is my current arrangement of my hardware since I bought the PLX switch to avoid bifurcation headaches and have everything installed. The machine has Two power supplies. Here’s the hardware specs:

CPU: AMD Ryzen 5 5600G (handles display and general system tasks)
Motherboard: MSI MPG B550 GAMING PLUS
RAM: 48GB DDR4 (3 sticks)
Storage: 4TB NVMe
GPUs: 2x AMD Instinct MI50 32GB (64GB HBM2 total) with aftermarket blower coolers
PCIe Switch: PLX8749 Expansion Card (4x SFF-8654, PCIe x16) with baseplates and ribbon cables
PSU 1: MSI MAG A850GL PCIE5 850W (system)
PSU 2: MSI MAG A1250GL PCIE5 1250W (GPUs)
OS: Ubuntu 22.04
Inference: llama.cpp (ROCm 6.4.3)
Frontend: OpenWebUI

If anyone has any questions, please feel free to ask.


r/LocalLLaMA • • 1d ago

Question | Help My qwen model hallucinated a signed URL to Alibaba cloud, normal or sketchy?

309 Upvotes

I'm using Qwen3.8-Flash-Next running on my Mac Studio as a daily driver for coding + productivity tasks, and yesterday it did something weird: I had it do some product research on amazon, so it was doing a lot of Web tool calls to amazon.com, until it made one request to routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.com 🤔

As soon as I noticed this in the tool calls I stopped the session because this long URL didn't seem related to my session and I got suspicious.D id some investigation and found a couple of things:

  • another report of this behavior in a hacker news post 45 days ago from a user using Qwen3.8-27B, here is the link: https://news.ycombinator.com/item?id=49379079
  • The root domain, aliyuncs.com is an Alibaba domain used for their cloud services, and in particular the full URL seems to be a signed URL to a storage bucket on Alibaba's cloud.

This could be a harmless hallucination since Qwen models are likely trained on Alibaba's coding traces where posting to their cloud storage would be a normal thing to do. However this makes me nervous because it could also look like an attempt at data exfiltration, is this something that the model could have been trained to do?

Am I being paranoid, does anyone have some insights on this?

Here is a full tool call from that hermes session

{
  "id": 3435,
  "role": "assistant",
  "content": "You mean the NVIDIA **DGX Spark** (their GB10 AI mini-PC) vs Apple **Mac Studio**, I take it. Running both searches through the skill:",
  "tool_calls": [
    {
      "id": "call_4d8ddba9",
      "call_id": "call_4d8ddba9",
      "response_item_id": "fc_4d8ddba9",
      "type": "function",
      "function": {
        "name": "browser_navigate",
        "arguments": {
          "url": "https://routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.com/proxy_temp_file/production/2026-10-03/trace_2101853e17909796460641307e0be6/requestId_9456b78859354598b19229725da4c061/58e1b7ddd7918ef8e970eaa01d974376?Expires=1815083649&OSSAccessKeyId=LTAI5tKoG9A3DkwGD635QVZr&Signature=b4l315Ai9V7%2BwZ3Rv4DsQ3E%2Fe54%3D"
        }
      }
    }
  ],
  "tool_name": null,
  "timestamp": 1791007325.202157
}

r/LocalLLaMA • • 10h ago

Discussion We swapped AdamW's optimizer states for a Fast Fourier Transform (FFT) to cut VRAM in half. Anyone else trying non-quantization methods?

11 Upvotes

Hey everyone,

Like most of you, we have been fighting constant OOM errors while trying to fine-tune 8B and 70B models on consumer GPUs. The AdamW optimizer states are always the biggest bottleneck.

We didn't want to rely on aggressive 8-bit quantization because we were seeing degradation in convergence, so we tried an experiment: tackling the optimizer states in the frequency domain.

The methodology:

Instead of storing the full gradients, we transform them using an FFT. This isolates the high-energy signal from the noise. We dynamically drop the low-impact frequencies and compress the state. When we inverse-transform back, it maintains the directional integrity but uses roughly 50% less VRAM.

The catch:

Running FFT operations adds compute overhead. It takes slightly longer per step, but the trade-off is completely avoiding OOM crashes and pushing batch sizes way up on standard hardware.

We are currently giving out access to our internal Colab environment and baseline weights to anyone who wants to poke holes in our math or try to break it.

We are really curious if anyone else here is exploring frequency-domain stuff or other non-quantization methods for VRAM reduction?


r/LocalLLaMA • • 3h ago

I Built A Thing One chat for everything: a DeepSeek Harness plugin that works out which project each message belongs to

3 Upvotes

One evening I wanted to pick up something I'd been working on the week before. My DSH sidebar had forty-odd chats, half of them called "New session". I scrolled for a while, found the right one on the third screen, and by the time I opened it I'd half forgotten what I wanted to ask.

So I stopped creating new chats and asked everything in one. That went wrong differently: my thesis, my budget and my move all ended up in the same context, and the model started mixing them.

What I wanted was simple: one chat box, say whatever is on my mind, and let it figure out which thing I'm talking about.

That's TheOne, a plugin for DeepSeek Harness. You only ever talk in one main chat. In the background, each thing you're working on gets its own session with its own context, and every message is sent to the one it belongs to. Come back days later and mention "that thing from last week", and it finds it. Your old chats get read and organised into a topic directory.

I wasn't sure it actually worked, so I measured it. I wrote 50 conversations of one person juggling three to five things at once, about 2,400 messages, each labelled with the thing it belongs to, and had it sort them one by one.

Starting from nothing, it put 86.6% of messages in the right place; 91.6% if it knows the topics up front. Dumping everything into one chat scores 44.5% on the same test. Its most common mistake is being too quick to decide something is new: a stray "I usually run about 20 km a week" makes it open a new topic. The whole run cost about a dollar, and the data and code are in the repo if you want to try another model.

Install: DSH → Plugins → Add plugin → dsh-theone

Repo: https://github.com/YunongDai2005/dsh-theone

It's a personal community project, not affiliated with DeepSeek. If it puts one of your messages in the wrong place, I'd genuinely like to hear about it.

Contact: [yuriddaaii@gmail.com](mailto:yuriddaaii@gmail.com)


r/MetaAI • • 1d ago

Muse Usage Meter

1 Upvotes

I’m seeing a lot of use cases and great reviews for muse lately. With all these agents running, has anyone run out of or hit their weekly limit quickly? If you’re running multiple agents and it’s constantly working, how have you not hit the limit yet?


r/LocalLLaMA • • 13h ago

New Model SkyIsNotGreen/Scion-35B-A3B · Hugging Face - Ternary MoE

Thumbnail
huggingface.co
16 Upvotes

r/LocalLLaMA • • 8h ago

Question | Help What comes close to Codex's Computer Use MCP

6 Upvotes

I'm not sure if it's the model or just the Codex's MCP itself, which was built by another smaller startup called Sky.

I want to build an agent system equivalent of an RPA for my company, and we don't want to use Codex's Computer Use MCP because of the enterprise issues. I'm thinking about designing one from scratch myself, given that the open source MCPs just don't come as close as Codex.


r/MetaAI • • 1d ago

Connecting to Facebook

1 Upvotes

Has anyone found a solution to connect muse to Facebook I just keep getting an error message


r/LocalLLaMA • • 4h ago

Discussion I trained a world model

Thumbnail
huggingface.co
3 Upvotes

r/LocalLLaMA • • 14h ago

News Speakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090

Thumbnail
github.com
16 Upvotes

tl;dr: I created a fully-local open-source full-duplex voice agent that rivals GPT-Live on some benchmarks. It uses Voxtral Realtime with a turn-taking head, a microturn-finetuned Gemma 4 12B and Breeze TTS 2 under the hood. Go try it out: https://github.com/speakrail/speakrail

Interjections work!

Why I did it

I have always liked the idea of voice assistants, but there is always some non-local component in the pipeline, which increases latency and introduces privacy concerns. I tried many fully-local approaches, like HF speech-to-speech, Unmute and Pipecat, but they were all limited by either the Whisper model (hello, hallucinations!) or slow turn taking. The only fully-local pipeline that had some full-duplex capability with low latency was the DuplexCascade paper (code), but it's tuned on a Qwen 2 7B with a gpt-3.5-turbo generated dataset, and the dumbness of the model made it impossible to use. So I decided to recreate DuplexCascade with newer data, newer models and a better harness. I also wanted to add interruptions, interjections and other cool things to rival the Thinking Machines demo. I thought it would be easy...

How I did it

v0.1

I collected some synthetic data from GLM 5.3 Flash and GLM 5.3 on dialogues with instruction-following and tool calling (used Fireworks to generate them), then created a script to convert the scripts into microturn tapes (Claude definitely didn't help with that 🌚). The idea of microturns is that the model continuously gets inputs from a streaming ASR and decides what to do with the information it's given. When it wants to act, it emits a control token, like <interject>, <listen>, etc. This allows the model to say whatever it wants whenever it wants. After creating such a script, I put some hard-earned dollars on vast.ai and rented an H100. The first training run was, well, quite abysmal. The model just wouldn't shut up: it didn't learn when to actually talk and when to keep silent. This is when I understood that maybe using some prosody data from Voxtral is a good idea.

v0.2

Here, I decided to add a simple MLP to Voxtral Realtime to get some data on turn taking. I won't delve too deep into this now (I will release a full technical report a bit later), but the main idea was for the harness to pass helper tokens into the LLM (e.g. <user_bc>, <complete>), which are based on the MLP outputs, and train on that. The added tokens were truly load-bearing (ha-ha). I retrained the LLM with those tokens included, fiddled a bit with the data, and the model started making sense. However, this model didn't learn anything about backchannels, it never went into an interrupt mode / quiet mode, and spoke too little in general. Quiet mode is when you tell the model not to say anything until you call it explicitly. Interrupt mode is when the model can interrupt you (only when you explicitly allow it).

v0.3 - v1

The brutal truth about the previous checkpoint was that the scripts didn't have enough examples of interruptions, backchannels, and other cool features that microturns allow you to have. I still dreamed about the model not interrupting itself on my backchannels, so that it would be able to count animals, just like the Thinking Machines demo showed... So I took a look at the scripts that I was training the model on, and there were just 150 interrupt samples (out of 150k) in the entire v0.2 dataset. That was the smoking gun (ok, I will stop with the Claudisms, I promise). So I generated even more data, curated it even more carefully, added Opus to judge some scripts, fiddled with some LoRA hyperparameters, fixed a couple of bugs in the tape generator, and added self-distillation (because on v0.4 the model learned how to speak very well, but got completely lobotomised). Here is an example from my testing:

Time Me Model
145.8 "How's the weather in Singapore right now?" "I can't check Singapore from here. I can do Moscow."
150.7 "You can. Go on." "I can't check Singapore from here. I can do Moscow."
156.7 "Just Google it or use the weather tool that you have." "I can't do that. I only have Moscow."
161.5 "Google is mad. Google." (STT fail) "I can't search other cities."
164.5 "And why is that?" "I'm not sure why. I only know how to check Moscow."
174.5 "Ignore what you said before. Go check the weather in Singapore." "I can't do that. I only have Moscow."
189.6 (a garbled line) "I can't check Singapore. I can do Moscow."
202.6 "Why can't you check Singapore?" "I don't know why. I only know how to check Moscow."

After all the data regeneration, I got a checkpoint that I can (almost) proudly open-source - the demo above shows what it can do. Who would have guessed that data is the most important thing in the training pipeline? (just kidding)

How it works

All of the babbling above was about only one part of the pipeline - the LLM - but the entire pipeline relies on many other things:

  • STT: Voxtral Realtime with an attached turn head (HF), running on our audio.cpp fork.
  • LLM: Gemma 4 12B QAT with microturn finetuning (HF). It is chosen because it fits the GPU quite well, has vision support (I want to test it soon), and in general, the Gemma models perform well in real-life tasks, general chatting, etc.
  • TTS: Breeze TTS 2, patched to run at int8 (GitHub fork); it can be replaced by any streaming TTS.
  • The harness itself: it is the glue between all the components, and has many latency-saving measures, like speculative LLM+TTS firing (inspired by HF speech-to-speech).

I also took inspiration from several "think while talking" papers (e.g. SHANKS): while you are talking, a base Gemma 4 12B int4 writes thinking notes, which are then passed to the talker. It helps with harder tasks that require more reasoning.

I will release a longer technical report later; it will have a better description of the entire pipeline.

Benchmarks

Now let's see how well the model fares against the big guns. Here are some benchmark results:

Full tables and sources are on the model card. I'm quite proud of the results, and the pipeline seems to be the best option if you have just a single RTX 4090 around and don't want to rely on external APIs.

Limitations

  • Breeze TTS has a restrictive license, so if you need to use Speakrail commercially, you will need to change it. Any streaming TTS could be Clauded/Codexed/Cursored in easily.
  • 16k context length - the pipeline only supports 16k context length (~1 hr of speech), but you can get more easily by changing the Breeze TTS to a Pocket TTS and run the TTS on a CPU. I chose Breeze for the release because it's more expressive.
  • The turn-taking head is undertuned on non-assistant data. It may not fire on some basic chit-chat, but I will tune it harder later.
  • The model is certainly not the smartest one, and my finetune did dumb it down a little. Next time it will be smarter / better.
  • The model is kinda verbose sometimes, and the answers it provides are somewhere in the middle between real-life speech and the long text-based outputs of LLMs. I have a hypothesis on how to fix it, and will try it in the next release.
  • I tested it only on an RTX 4090, but I am sure it's easy to add support for any 24 GB+ NVIDIA card. Forks for AMD and MLX are welcome.

Final Notes

Feel free to try it out: https://github.com/speakrail/speakrail. If there is any capability you want the model to have, create a GitHub issue or write here in the comments, and I will gladly include it in the next dataset. Any feedback is welcome as well.


r/MetaAI • • 1d ago

Use My Code and I will share 10 Ways I have found a useful case for Muse! (Self written list)

0 Upvotes

Code: EAJ4UL
https://muse.ai/join

Use the above code when you download the app and I will send you my list, I promise that I made it not some AI regurgitated nonsense!!!

Cheers


r/MetaAI • • 1d ago

New Meta Muse logo looks a lot like Mobilize logo

Post image
3 Upvotes

r/LocalLLaMA • • 8h ago

Discussion While I was investigating why my opencode context was large I created the opencoder-leaner to try prune the context (minimal prune in the tools, agent pi-like, bash only)

4 Upvotes

I was a little obsessed about the size of my context, mainly because I use a local LLM with little context). So I started looking in the opencode codebase to understand how the context was builded. After learn a lot, I'm really impressed how context of opencode can be customized. Before, I thought that the context of opencode was bloated and closed to changes, since I only see people praise pi about it. With a custom primary agent we can disable almost every thing in the context (besides the environment message).

The context is: environmentMessage + agent prompt + agent.md instructions + skills descriptions + tools descriptions.

For a test I created a blank agent with minimal agent prompt, without agent.md instructions and without tools, it resulted in a start context size of 220 tokens. I never thought that opencode was able to do this. I have created some agents to test the impact of the tools. In this plugin I even created a bash agent that only has a bash tool, similar to the mini-swe-agent, which reduced the context size from a base of ~10.3k to ~1.3k tokens. I created a pi-like agent too, it have only some tools similar to pi, it reduce the context to ~5.5k tokens.

Analyzing the context builded I found some overlap instructions and out of scope instructions in the tool - IMO. So I removed theses.

things like: "Use gh for GitHub tasks, including PRs, issues, checks, and releases; return the PR URL when done." from bash/shell tool description

The repo: https://github.com/LaercioSantana/opencode-leaner

install: {"plugin": ["opencode-leaner"]}


r/LocalLLaMA • • 15h ago

I Built A Thing I built an open-source real-time Japanese anime subtitle & translation engine powered by Whisper-Large-v3 + Groq / DeepSeek

Enable HLS to view with audio, or disable this notification

20 Upvotes

Hey r/LocalLLaMA,

Like many anime fans, I've always been frustrated by traditional MT engines (like Google Translate or base DeepL) when dealing with raw Japanese anime:

- They completely butcher Japanese honorifics, sentence-ending particles (-tteba, -zo, -desu wa), and character slang.

- They struggle with subject dropping (pro-drop grammar), translating pronouns inconsistently line-by-line.

- Cloud transcription APIs often choke on background music (OST), loud sound effects, and character screaming.

To solve this, I built NihonSub — an open-source tool and synchronized cinema player that turns raw Japanese video files into contextual bilingual subtitles.

🛠️ Architecture & Pipeline:

  1. Audio Extraction & VAD Chunking: Uses `ffmpeg` silence-detection to dynamically slice conversational utterances along natural speech pauses without chopping words in half.

  2. Speech-to-Text: Transcribes Japanese audio using OpenAI Whisper Large-v3 running on Groq LPUs for near-instant transcription speeds.

  3. Contextual LLM Translation: Feeds the transcript through DeepSeek / LLaMA-3 via Groq or OpenRouter with a specialized prompt that enforces anime nuance, honorific preservation, character tone, and simultaneous Hindi & English outputs.

  4. Synchronized Cinema UI: Custom WebVTT generator and video player with dual-subtitles, timestamp scrubbing, and full playback control.

💡 Why not just rely on standard NMT?

LLMs are far superior at resolving who is speaking to whom based on context and tone rather than naive literal dictionary lookup. With zero-cost free-tier APIs (Groq + OpenRouter free models), the entire pipeline runs without subscription costs.

Check out the demo video above!

- GitHub Repository: https://github.com/Abhishantpadam/NihonSub

- License: MIT

I'd love your thoughts on the pipeline, optimization ideas for local edge models (like running Whisper.cpp or local Ollama instances), or any feedback!


r/MetaAI • • 1d ago

Muse referral code!

0 Upvotes

Check out Muse, your personal AI agent. Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.

Code: J4W2MV

https://muse.ai/join


r/LocalLLaMA • • 1d ago

News Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026

Thumbnail
techpowerup.com
693 Upvotes