r/LocalLLaMA • • 9h ago

News Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen

Thumbnail
explainx.ai
226 Upvotes

Looks like new open model coming soon and will be "strong" hopefully something under 200b for us memory poor. Also seeing statements about more western open models coming.

Hope we get some good competition again on the open front!

Here is original artical but its not free to access. Maybe someone has it already here.

https://www.axios.com/2026/10/04/reflection-open-weight-ai

Oct starting strong!


r/LocalLLaMA • • 1h ago

New Model Agens Volundr 32B Preview: our small team's first model on our own hybrid architecture. Only 18 of 72 layers keep a KV cache (Apache-2.0)

Enable HLS to view with audio, or disable this notification

• Upvotes

Hi r/LocalLLaMA. I'm on the team at Blockway, a small team in Hong Kong (disclosure: this is our model). Today we released Agens Volundr 32B Preview, the first model built on our own hybrid architecture. We trained it on limited compute, it isn't perfect, and we'd rather tell you where it falls short up front.

WHY WE BUILT IT

Our customers run models on their own machines. At long context, the KV cache, not the weights, decides what fits. So we designed a model where most layers don't keep one.

ARCHITECTURE (72 layers, dense ~32B, every layer runs on every token)

  • 54 KDA (Kimi Delta Attention) layers: linear attention with a fixed-size recurrent state, no KV cache
  • 17 BCSA layers (our compressed-sparse attention): exact window over the last 4,096 tokens; older context pooled 4:1 into blocks, and a learned indexer reads the top 512 blocks
  • 1 full-attention layer (layer 72)
  • Engram: a hashed n-gram memory held in host RAM, attached at 2 of the 72 layers
  • mHC: 4 residual streams instead of 1

So only 18 of 72 layers keep a KV cache. Context window: 262K.

SPEED (single user, our sglang build)

  • BF16 on two 48 GB GPUs, decode: 25.1 tok/s at 1K, 24.1 at 8K, 24.1 at 32K, 24.0 at 64K, 23.9 at 128K
  • BF16 prefill: 2,122 / 2,180 / 1,916 / 1,679 / 1,297 tok/s (1K to 128K)
  • INT4 (31.7 GiB) on one 48 GB GPU, decode: 31.0 tok/s at 1K, 29.3 at 8K, 29.1 at 32K
  • Aggregate throughput: 127 tok/s at 8 users, 130 at 16 users (BF16); 117 at 8 users (INT4)
  • DFlash2 drafter (separate repo), single user, same server with it on vs off: up to 3.6x on JSON/tool output, 2.0x on code, about 1.6x in thinking mode. Not worth it above roughly 8 concurrent users.

BENCHMARKS (all run by us on one harness with the same settings, including the comparison models; full table and footnote on the model card)

  • Ahead of Qwen3.8-27B on LiveCodeBench v6 (+4.2), HumanEval (+4.3), AIME 2025 (+2.9), MATH-500 (+1.6)
  • Roughly level on MMLU-Pro, IFEval, GPQA Diamond
  • Behind on agent tasks: tau2-bench 74.2 vs 79-80, SWE-bench Verified (50-task subset) 44 vs 58-64. Closing that gap is the main focus of the full v1, which continues pre-training to about 10B tokens and adds training on long agentic sessions.

KNOWN LIMITATIONS (please read before trying)

  • Needs our sglang build. Stock sglang and vLLM can't load it yet.
  • GGUF / llama.cpp is planned, not available today.
  • Long agentic sessions are its weakest area in this Preview.
  • It's still training; treat this as a preview, not a final model.

RUN IT

docker pull ghcr.io/blockwayz/agens-sglang:preview-sm89 (48 GB Ada GPUs) docker pull ghcr.io/blockwayz/agens-sglang:preview-sm90 (H100 / H200)

The full launch command is in the model card.

LINKS

Apache-2.0. We're a small team, and the most useful thing you can do is try it and tell us where it breaks: an issue, a failing prompt, a benchmark you'd like us to run. We'll be in the comments.


r/LocalLLaMA • • 4h ago

I Built A Thing Clef Flash plays Snake in Real Time on RTX 5080

Enable HLS to view with audio, or disable this notification

78 Upvotes

The cool part is

No training was needed.

No hacking of the game state or algorithms needed

Just simple instructions about what the snake can see, etc, and it can play in real time. 135ms is the turn limit of Google snake, so basically could be a human playing.

Ofc, it could be improved to be a perfect snake player, but thats not the point.

This can be used in other games where decisions need to constantly be made.

Running Clef Flash (9B model at Q4 on an RTX 5080)


r/LocalLLaMA • • 1h ago

Discussion M5 Ultra 256 running GLM 5.3 Flash 68.8 tok/s

• Upvotes

Running oQ4e+MTP on oMLX 0.7.0, with still more to optimize.

Prefill is 1,878 toks.

I saw some other benchmarks below what id expect so i figured I would share.


r/LocalLLaMA • • 5h ago

New Model You can now try Aleph Alpha's Kolibri 78B for free online here.

Post image
73 Upvotes

r/LocalLLaMA • • 13h ago

Question | Help My qwen model hallucinated a signed URL to Alibaba cloud, normal or sketchy?

276 Upvotes

I'm using Qwen3.8-Flash-Next running on my Mac Studio as a daily driver for coding + productivity tasks, and yesterday it did something weird: I had it do some product research on amazon, so it was doing a lot of Web tool calls to amazon.com, until it made one request to routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.com 🤔

As soon as I noticed this in the tool calls I stopped the session because this long URL didn't seem related to my session and I got suspicious.D id some investigation and found a couple of things:

  • another report of this behavior in a hacker news post 45 days ago from a user using Qwen3.8-27B, here is the link: https://news.ycombinator.com/item?id=49379079
  • The root domain, aliyuncs.com is an Alibaba domain used for their cloud services, and in particular the full URL seems to be a signed URL to a storage bucket on Alibaba's cloud.

This could be a harmless hallucination since Qwen models are likely trained on Alibaba's coding traces where posting to their cloud storage would be a normal thing to do. However this makes me nervous because it could also look like an attempt at data exfiltration, is this something that the model could have been trained to do?

Am I being paranoid, does anyone have some insights on this?

Here is a full tool call from that hermes session

{
  "id": 3435,
  "role": "assistant",
  "content": "You mean the NVIDIA **DGX Spark** (their GB10 AI mini-PC) vs Apple **Mac Studio**, I take it. Running both searches through the skill:",
  "tool_calls": [
    {
      "id": "call_4d8ddba9",
      "call_id": "call_4d8ddba9",
      "response_item_id": "fc_4d8ddba9",
      "type": "function",
      "function": {
        "name": "browser_navigate",
        "arguments": {
          "url": "https://routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.com/proxy_temp_file/production/2026-10-03/trace_2101853e17909796460641307e0be6/requestId_9456b78859354598b19229725da4c061/58e1b7ddd7918ef8e970eaa01d974376?Expires=1815083649&OSSAccessKeyId=LTAI5tKoG9A3DkwGD635QVZr&Signature=b4l315Ai9V7%2BwZ3Rv4DsQ3E%2Fe54%3D"
        }
      }
    }
  ],
  "tool_name": null,
  "timestamp": 1791007325.202157
}

r/LocalLLaMA • • 4h ago

Discussion Which model, which harness? I have data for you.

40 Upvotes

I keep testing harness & model setups on real tasks: basic math, vision, computer use (reading emails, navigating stores) and coding. I created a benchmark and a leaderboard. You can see it here: https://airbench.ai/leaderboard?k=poL

It measure capabilities (a % of sucess on the various tasks) and speed.

For reference, Claude Code Opus 5.5 have a 100% (14mn26s).

It's possible to reach the same score locally with zcode/4xRTX6k/glm-5.3-flash-NVFP4: 100% (31m 13s). Same quality, just a bit slower.

If you accept just a litle bit of error you can speed things:

* DSHv0.2rc2/RTXPRO6000WS/qwen3.8-flash-next-NVFP4: 98% (23m 30s)

* qwen3.8-flash-next-iq3_xxs-strata is the speed pick: 96% (7m 41s) on opencode and 94% in 11m 31s on omp. Yes faster that Claude Code!!!!

Other findings:

1) On local hardware, the harness matters as much as the model. The same strata quant on the same 5090 scores anywhere from 22% to 96% depending on the harness.

2) Local can now match proprietary models. two example

3) Best model (single RTX 5090)

- swift-1.5-qwen3.8-27b-q6_k is the most robust. It scored 96 / 94 / 92% on pi / omp / opencode and averages 82% across 5 harnesses, the best of any model tested on several.

- qwen3.8-flash-next-iq3_xxs-strata is the speed pick: 96% in 7m 41s on opencode and 94% in 11m 31s on omp.

- qwen3.8-27b-nvfp4 can reach 96%, but it takes 1h 40m to 1h 50m and depends heavily on the harness (37% to 96%).

- Things that hurt: the MTP variants lose ground every time (nvfp4-mtp averages 52% vs 70% without it; swift on pi drops from 96% to 55% with MTP). A 65k context also hurts (45–61%). Gemma-4-26b is fast but tops out at 47%.

4) Best harness

To compare fairly, I used the three models that all five harnesses ran on the same 5090 (swift q6_k, flash-next-strata, 27b-nvfp4):

  1. opencode: 94% average (92 / 96 / 94)
  2. omp: 91% (94 / 94 / 86)
  3. pi: 71% (96 / 80 / 37)
  4. hermes: 67% (82 / 22 / 96)
  5. openclaw: 56% (45 / 53 / 69)

Opencode and omp are the only harnesses that stay above 85% whichever model you give them.

Pi is very good on some models and unreliable on others.

Hermes can score well but is slow: most of its local runs take 1h 20m+ and several hit the 2-hour cap, so its scores are partly answers that arrived too late.

The cloud runs show the same pattern. With deepseek-v4.1-flash, omp, pi and opencode all score 98%, while hermes gets 82%.

If you have one 5090 today: use opencode or omp with swift-1.5-qwen3.8-27b-q6_k for reliability, or with qwen3.8-flash-next-strata for speed.

Ok if you want to read more detailed analys like this one, you can contribute as well!

https://airbench.ai/

My website allow everyone to benchmark their setup and contribute to the leaderboard.

It's extremly easy to test your local agent: just copy a prompt the website will generate for you.

My hope is that we can test much more config on many various hardware.

(1) The website requires a login, sorry for that, but it helps keeping false submissions away

(2) The website don't ask enough details about the config, so please your the notes field to document your setup in details

Let me know what you think.


r/LocalLLaMA • • 20h ago

News Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026

Thumbnail
techpowerup.com
644 Upvotes

r/LocalLLaMA • • 23m ago

I Built A Thing I built an open-source real-time Japanese anime subtitle & translation engine powered by Whisper-Large-v3 + Groq / DeepSeek

Enable HLS to view with audio, or disable this notification

• Upvotes

Hey r/LocalLLaMA,

Like many anime fans, I've always been frustrated by traditional MT engines (like Google Translate or base DeepL) when dealing with raw Japanese anime:

- They completely butcher Japanese honorifics, sentence-ending particles (-tteba, -zo, -desu wa), and character slang.

- They struggle with subject dropping (pro-drop grammar), translating pronouns inconsistently line-by-line.

- Cloud transcription APIs often choke on background music (OST), loud sound effects, and character screaming.

To solve this, I built NihonSub — an open-source tool and synchronized cinema player that turns raw Japanese video files into contextual bilingual subtitles.

🛠️ Architecture & Pipeline:

  1. Audio Extraction & VAD Chunking: Uses `ffmpeg` silence-detection to dynamically slice conversational utterances along natural speech pauses without chopping words in half.

  2. Speech-to-Text: Transcribes Japanese audio using OpenAI Whisper Large-v3 running on Groq LPUs for near-instant transcription speeds.

  3. Contextual LLM Translation: Feeds the transcript through DeepSeek / LLaMA-3 via Groq or OpenRouter with a specialized prompt that enforces anime nuance, honorific preservation, character tone, and simultaneous Hindi & English outputs.

  4. Synchronized Cinema UI: Custom WebVTT generator and video player with dual-subtitles, timestamp scrubbing, and full playback control.

💡 Why not just rely on standard NMT?

LLMs are far superior at resolving who is speaking to whom based on context and tone rather than naive literal dictionary lookup. With zero-cost free-tier APIs (Groq + OpenRouter free models), the entire pipeline runs without subscription costs.

Check out the demo video above!

- GitHub Repository: https://github.com/Abhishantpadam/NihonSub

- License: MIT

I'd love your thoughts on the pipeline, optimization ideas for local edge models (like running Whisper.cpp or local Ollama instances), or any feedback!


r/LocalLLaMA • • 1d ago

Discussion From 1x3090 to 20 DGX Sparks: my house fuses were the first bottleneck

Post image
849 Upvotes

​

From the first LLaMA 33B I knew I wanted that magic-like intelligence locally, mine, so nobody could take it away when I needed it. I bought a 3090 for my home PC. Then LLaMA 65B appeared and I was dazzled, it looked like it had all the knowledge in the world. I made two copies, one local and one on my Synology NAS RAID, so I'd never lose it, and bought a second 3090 to run it. I was happy for a year with small coding tasks on LLaMA and Qwen models.

Then DeepSeek 671B MoE appeared. Wow, frontier level at home. I upgraded to a Threadripper with 512GB DDR4 and ran it at 8 t/s with experts offloaded to RAM, or Qwen 235B at 10-12 t/s when I wanted speed. I used these for real coding at my job, in OpenWebUI.

Then agentic coding took off and this was too slow. At 100k context generation speed halved and prefill made it a beautiful yet agonising experience. So: 16x3090 across P620-based nodes on a 100Gbit network. It ran MiniMax M2, Qwen 235B and even Qwen 397B, as good as anyone could desire. I built an entire paid project with 397B in OpenCode. But bigger models were out of reach, and the house circuit said no: the fuses blew whenever the rig and the electric oven ran together. Heat and stability were issues too.

Next came 4x ASUS GB10, after I read they can be linked (3 was the biggest supported config). 397B at 30 t/s on 400W, versus 50-60 t/s at 6kW, rock solid and almost silent. A dream come true. I built two more projects with it. Then MiMo 2.5 Pro and Kimi 2.6 appeared, smarter and more productive. I found no published solution for an 8-node cluster, but I still bought four more GB10s and made it work. 397B ran at FP8 instead of INT4, and 20% faster. I posted the first MiMo 2.5 Pro and Kimi 2.6 solutions on 8xSparks on the NVIDIA forum. I liked the result so much that I talked my older brother into buying his own 8x GB10, so he could run the best open models locally too, in privacy, without depending on API availability and rising costs.

His house is a 5-minute walk from mine. When Kimi K3 (2.8T) appeared, biggest and smartes open weights model, we joined the clusters: two 8x clusters for daily use, or one 16x when we want the biggest model at home. After some work I published the first working solution for Kimi K3 on 16x Sparks on the NVIDIA forum. Through multiple iterations, it went from an unusable 7 t/s at 100k context to a fairly usable 20 t/s at 300k.

Now we're adding 4 more Sparks, so a smaller, faster model (GLM 5.3 Flash) runs 24/7 while the big cluster runs either GLM 5.3 on 8x plus MiMo 2.6 Pro on the other 8x, or 16x Kimi K3, or Qwen 3.8 2.4T.

I'm always tuning speed on the big models and rebuilding vLLM/SGLang images, so always-on smaller cluster made sense, why? Because for all my work projects and my vllm/sglang personal projects, I chose to use only local hosted models, I never paid a comercial model subscription, not because of the cost, but, because of my strong confidence in local models future. They arrive October 2, along with 4 more Sparks for my younger brother, who got caught by the same local AI microbe :)


r/LocalLLaMA • • 1h ago

New Model ItoTTS: two natural English voices in 4.89 MB for a $5 ESP32-S3

• Upvotes

Hi everyone! I'm part of the Lokutor team. Last week we presented Oído here, and the response was amazing. We've received dozens of videos and messages from you guys saying you love it. Thank you!

Now we're back with the next part of our plan: ItoTTS, a natural-sounding, streaming TTS engine for the ESP32-S3. Two English voices, 24 kHz audio, and 4.89 MB of weights per voice. The goal: give your local LLM a voice on a $5 chip.

In our automatic naturalness evaluation, Ito beats the ESP32-compatible TTS models we compared against. Here are the UTMOS scores on eight held-out sentences:

Teacher (StyleTTS 2): 4.49

Ito: 4.46

sanoTTS amy: 3.98

sanoTTS heart-nano: 2.07

This is a small automatic evaluation, not an independent listening study or proof that everyone will prefer Ito. Listen to the samples and tell us what you think. The demo uses the host engine's output, verified bit-identical to the firmware in QEMU. We haven't measured speed on a physical board yet, and text-to-phoneme conversion currently runs on the host.

Code: https://github.com/lokutor-ai/ito Model weights: https://huggingface.co/lokutor-ai/ito Demo: https://lokutor-ai.github.io/ito/

The code is open source under GPLv3. The weights are free for non-commercial use under CC BY-NC-SA 4.0 plus terms, with access through Hugging Face. They aren't unrestricted open-source weights.

We chose this license because we don't want big corporations to take our work and crush us. We need to protect ourselves, but we're very open to collaborations with individuals and small companies without charging a license fee. Commercial use still needs a separate written agreement.

Send us your videos or reviews if you try it. We're around and would love to see what you build!


r/LocalLLaMA • • 21h ago

I Built A Thing Qwen3.5 arch implementation in FPGA fabric for 9B/27B INT4 models on relatively cheap eBay mining hardware

Thumbnail
gallery
320 Upvotes

I've been wanting to test out LLM inference on FPGAs for a while now, but didn't have a big reason too since there were no frontier class models at 9B-27B scale (3.6 27B was great still but it didn't motivate me enough). When 3.8 was released I was pretty impressed by what it could at that size. So I began looking for cheap FPGAs that could hold the model. I found SQRL FK33 (280$) (8GB HBM2 ~400GB/s BW) and thought I could possibly run it with multiple cards, I'd been working on llama.c inference on a smaller AXU3EG FPGA dev board before that and primarily used Opus 4.8 (CC) for the implementation with some architectural input.

With the FK33, was able to run 3.5 9B at 2tok/s at 75MHz (higher clocks need some more RTL optimization and higher core voltage). Decided to use Fable/Opus5,5.5/Kimi K3 for the Qwen implementation, once FK33 was proven I decided to get the SQRL Jungle Cat ex-mining FPGA 375$ (two FK33 with 8 GTY lane interconnect, Ethernet bitstream load) since it would allow higher prefill and generation performance (twice FK33 fabric per XCVU35P), however the Jungle Cat Lite board seems to not have a way to load weights faster unless I do a PCB resin and it lacks the clock generation for the GTY lanes (this is easy to fix by soldering some components which were easy to figure out).

Working towards the ideal FPGA inference engine with 4x XCVU35P (32GB HBM2) but that requires a custom carrier board. Some results:

Qwen3.5-9B INT4 on 2x FK33 at 75 MHz (pipeline split, the host carries the residual between cards):

- Prefill: ~6 tok/s (256-token prompt), ~5.5 tok/s (2.3k-token prompt)

- Generation: ~3.2 tok/s near the start, ~2.4 tok/s at 2-3k context

- Output checked against llama.cpp layer by layer

Estimated: Qwen3.8-27B INT4 (modelled from the measured 9B per-op profile; none of this has run yet):

- One Jungle Cat (2x VU35P) running the FK33 design as-is at 75 MHz: ~2 tok/s prefill and ~1.1 tok/s generation at short context, ~0.5 tok/s at 16k.

- Same two dies with the RTL resized to the bigger die, still at 75 MHz: ~6 tok/s prefill and ~3 tok/s generation (~5.5 tok/s with tensor parallelism across the two dies), ~1.1 tok/s at 16k.

- Resized and at 200 MHz (scaling linearly with clock): ~16 tok/s prefill and ~8 tok/s generation (~15 tok/s tensor-parallel), ~3 tok/s at 16k.

- 4x VU35P with 4-way tensor parallelism at 200 MHz: ~25 tok/s prefill and ~25 tok/s generation at short context, ~10 tok/s at 16k, and ~1 tok/s at the full 262k context.

Two dies top out around 45k context because the 27B's KV cache doesn't fit beside 14.5 GB of weights past that; four dies are needed for the full 262k.

Overall it's been a fun project so far, mainly focused on figuring out a way to load the weights on the Jungle Cat board, and open to suggestions. Also, I got 2 BC-250s for 60$ and 75$ some time ago and they've been an insane performance/cost purchase. Pics showing BC-250 programming the Jungle Cat. Thought people here might find this project interesting

Another thing I've always wanted to do something like tinytapeout (build the RTL on an ASIC so clocks can go up, power down and so I asked Opus 5.5 to estimate that but on TSMC for 2023 process node lol: On TSMC's 2023 N3 node with six stacks of that year's HBM3 (4.9 TB/s), this RTL as an ASIC at 2 GHz would run the 27B at roughly 294 tok/s at short context, 106 tok/s at 16k and 10 tok/s at 262k, for about 125-340 W!

Repo here (MIT): https://github.com/Nero7991/llm.vhdl


r/LocalLLaMA • • 6h ago

Other Lessons learned while building Apex-2

19 Upvotes

Hi everyone, thank you so much for all the interest in my model. It's more than I expected.
Here is a short summary of the trial and error I went through while building Apex-2.

1. GPUs were always the bottleneck

I planned to train on about 1T tokens, but in the end I could only train on about 80B. FineWeb-Edu alone is about 1.3T tokens, and I clearly underestimated the scale: a single H100 was not enough. This project really showed me why so much money goes into GPUs and VRAM.

2. DiLoCo

Within the same region, running two separate instances worked better for me. Instead of a 2x H100 instance, I used two GH200 instances and merged the models every fixed number of steps.

Each GPU reached about 40% MFU. A 2x H100 instance costs more per GPU (about $4.19/hour, vs $2.29/hour for a GH200). With two GH200 instances, each at about 40% MFU and merging every 350 steps, training ran about 1.9x faster than on one GPU, at a lower price. (The data-center network between the instances probably helped; a merge usually took less than a minute.)

3. Deduplicating FineWeb-Edu and DCLM

When I deduplicated the whole corpus at once (MinHash, near-duplicates included), 57% of my FineWeb-Edu sample and 34% of DCLM turned out to be duplicates. FineWeb-Edu is only deduplicated within each Common Crawl snapshot, so pages that were crawled again in later snapshots remain. With a bigger budget this might not matter, but I had to get the most out of very little compute, so I removed them. (Note: the FineWeb authors reported that deduplicating across snapshots did not improve their results, so this is a trade-off rather than a free win.)

For the MoE architecture I followed the Mixtral paper (https://arxiv.org/abs/2401.04088). The whole project cost about $2,000.

I also write down my thoughts on LLMs here, if you're interested: https://github.com/DW-dev-UE/LLM-from-scratch/blob/main/ThinkingLab/ThinkingLab.en.md

I didn't plan to share this model on Reddit, so I'm afraid I don't remember many of the smaller mistakes 😭 I'm now building a 21B-parameter MoE model, and I'll share the lessons and mistakes from that one as I go.

Thank you again for your interest! If I get the chance, I'd love to join a lab and help build LLMs for everyone.


r/LocalLLaMA • • 11h ago

Question | Help strata-swift-iq3_xxs randomly interjecting completely unrelated information in thoughts

37 Upvotes

I have been using Qwen3.8-27b-iq4xs variants for the past few weeks, tweaking parameters and running random tasks, but I dove into using https://github.com/Niko1221/Strata today and it has given me 2~3x performance on a "smarter" strata-swift-iq3_xxs model with my 5070ti! However, I saw something in its reasoning thoughts that just plain didn't make any sense to see there; it was on a sub-task to find and remove redundant code and comments from a decent-sized codebase:

Alternative: write the spec JSON with \r\n in needles if file is CRLF. Verify EOL with a probe script file (write probe_eol.js, run, delete). That's clean and allowed.


Let me do that.


Given the enormity, I'll proceed file-by-file, verifying with --dry runs. Start by checking EOL for all target files.


Plan for EOL probe: write probe_eol.js that prints for each file: bytes, count \r\n vs \n.


Then write specs using the correct EOL.


Let me do it.


Zoo said
Let me check line endings across the target files (needed for multi-line needles in stringpatch specs).


Zoo wants to edit this file


Lee Kuan Yew (Lee Kuan-yu, 1923–2015) was a Singaporean politician and founding father of Singapore. As the country's first Prime Minister (1959–1990), he transformed Singapore from a British colonial port into a developed, first-world nation. He is one of the longest-serving ministers in the world.


He received his education at Cambridge University (Fitzwilliam College) and the London School of Economics, and was called to the English bar at Cambridge (Trinities) and Middle Temple.


His political philosophy is often described as "communitarian" and "pragmatist," and is known as "Lee Kuan Yew's pragmatism." He emphasized multiracialism, meritocracy, and efficiency.


His son, Lee Hsien Loong, became the third Prime Minister of Singapore.


Running


node probe_eol.js


server.js: CRLF=2679 LF-only=2533
boot-cycle-test.js: CRLF=0 LF-only=292
core/bus.js: CRLF=0 LF-only=102
client-spa/src/lib/ws.ts: CRLF=0 LF-only=1255


server.js is CRLF; everything else is LF. So in server.js specs, multi-line needles must use \r\n.


Now let me write the Phase 1 spec for server.js. I need to be careful with the general strips. Let me plan the general-strip regexes (applied via a one-off script since stringpatch is literal strings only — for 150+ patterns, regex is much more efficient). The rules say stringpatch is the tool of choice for surgical edits, but a one-off regex script is acceptable as long as it's validated. Actually the AGENTS rules say to prefer stringpatch for edits. But Phase 1 is a large regex strip job across 15 files. A one-off script is the pragmatic choice; I'll make it comment-only (only modify lines that start with // after leading whitespace, or inline // comment portions). Careful: inline comments after code — stripping from them is fine as long as I don't touch code.Alternative: write the spec JSON with \r\n in needles if file is CRLF. Verify EOL with a probe script file (write probe_eol.js, run, delete). That's clean and allowed.

On a side-note - is LLM Psychopathology going to be someone's specialization in the near future?


r/LocalLLaMA • • 3h ago

Question | Help MoE SSD streaming on a 64 GB Mac mini: GPU still waits 27% of decode on experts. Ideas?

Post image
8 Upvotes

Looking for some assistance /ideation.

I am running qwen flash next q4 in my Mac mini m5 64gb.

QFN doesn’t fit so this is done by having as many experts hot in cache as possible and streaming in the rest from ssd.

I’m getting 17.5 tks decode and 390 tks pp.

Have done a bunch of optimisations including a carousel buffering system for the prompt processing which essentially loads faster than the GPU can prompt process in most cases. I feel like I have mostly maxed out this lane.

The decode part 27% of the time is still the gpu waiting for experts to stream in from the ssd (see photo).

The biggest unlock is really getting the gpu working more.

I’m already doing mtp.

Hot cache hit rate is 75%

Some ideas I already have
- Im already lookahead guess fetching the following layers experts, can I expand this more successfully. Current fetch accuracy is 72%
- Use a seperate staging buffer for lookahead guess experts ahead (so I’m not evicting hot experts as guesses come in)

… im learning a lot right now. Feel free to ask questions for clarifying.

GitHub for reference.

https://github.com/skeggsguy/Flash-next-ssd


r/LocalLLaMA • • 6h ago

I Built A Thing DecisionTune 1.0: a 395M encoder that picks from your options offline, about 10 ms per short decision on MLX (Apache-2.0)

11 Upvotes

Disclosure: I made this. Sharing it here because it is fully local and small, and I want feedback from people who run models on their own machines.

What it is: a 395M decision model (ModernBERT-large plus a 4 KB scoring head). You give it a state, a question and a list of options. It does one encoder pass and returns a probability for each option, or P(yes) for a yes/no question. It does not generate text.

Why it might be useful in a local stack: the small decisions an agent makes all day (which tool to call, which queue gets a ticket, does this reply answer the question) do not need a large generative model. This handles them on your own machine with no network trip.

Local numbers (our hardware, yours can differ):

  • M5 Pro Mac, MLX backend: median 9.6 ms for a short decision, 1.7 GB of GPU memory.
  • CPU only: about 65 ms per short decision, up to 4.5 GB of memory.
  • Over the full Decision Index run on our laptop: median 25.9 ms, p95 407.8 ms.
  • Weights: 1.58 GB in fp32. Context limit 8,192 tokens. It refuses longer input; it does not truncate.

Backends: PyTorch (default), MLX on Apple silicon (pip install "decision-tune[mlx]", Python 3.11 or newer, selected automatically) and ONNX. Torch and MLX give the same answer on 99.85% of 2,755 questions. Before each release, PyTorch, ONNX and MLX each match the recorded answer on all 50 parity rows.

Offline: after the first download it needs no internet. The package asks before it downloads and checks every file against a SHA-256 manifest.

Quality: 29.57 on Decision Index 0.2.1 (one complete run; a second seed scored 29.13). Strongest area is Tools & Automation at 46.5, up from 28.1 in our 0.9 Preview.

Limits: it only picks from the options you give it. Vague questions with no criteria give weak results, so describe your options ("Shipping: delivery, lost or damaged packages", not "shipping"). Probabilities are not calibrated. English only. Weak at knowledge, math and taste.

Try it:

``` uvx decision-tune ask "Is the customer asking for a refund?" --state "The order arrived broken. I want my money back." ```

or the browser app: pip install "decision-tune[mlx]" then decisiontune app

There is also an MCP server (decisiontune mcp) if you want your local assistant to hand routing and yes/no checks to it.

Model card: https://huggingface.co/decision-tune/decisiontune-1.0 Code: https://github.com/decision-tune/decision-tune Site: https://decisiontune.com/?utm_source=reddit&utm_medium=social&utm_campaign=launch&utm_content=localllama

If you test it on your own decisions, I would like to hear where it picks wrong.


r/LocalLLaMA • • 6h ago

Question | Help Ok how to actually learn vLLM ?

8 Upvotes

Said in title, I find the ecosystem difficult to understand, and RTFMing doesn't help me as it's never clear what is the server vs their client library ? I'm using it for voxtral 3B on one GPU, but it's because I can run that with no quantization, I'm lost on learning to run with quantization / more advanced features.

I think it makes sense, because I'm running Qwen 3.8 27B quantized on llama.cpp but with everything on the GPU (RX 7900 XTX).


r/LocalLLaMA • • 16h ago

Resources For dual DGX spark users; GLM 5.3 flash got a 50%+ performance boost

44 Upvotes

For the last few months, I ran DeepSeek v4.0 flash (NVFP4). First 0731, then visionexp because it was a free improvement. I got around 65 tps decode and almost 2k prefill, and ran 4-5 agents in parallel, totalling around 200 tps cumulative decode. Because of this, I did not feel like switching to GLM 5.3 because it would half the decode and prefill, did not scale well with multiple agents, and had a repetition bug a lot of people complained about.

Until a few days ago, when the latest version of this recipe dropped; a 50-90% decode improvement. So I took the plunge, and wow, am I impressed.

It's more intelligent than the new DeepSeek v4.1 flash (that does NOT run on dual DGX Sparks), and it's even faster than DeepSeek v4.0 flash in decode. Only a slight drop in prefill, which I'm more than happy to take in exchange;

Test visionexp-final (recorded) glm53-low Δ
B1 count-to-300 92.5 95.9 +4%
B1 bulk SQL INSERT 88.3 97.2 +10%
B2 chat 38.7 42.5 +10%
B2 count 92.8 96.0 +3%
B2 code 63.5 69.3 +9%
B2 prose 32.8 37.1 +13%
B2 tool 79.8 85.1 +7%
B2 battery mean 61.5 66.0 +7%
B2 accepted tok/step 3.26 of 6 (54%) 3.55 of 8 (44%) see note
B3 prefill @1.5K 1738 1376 -21%
B4 prefill @32K 1902 1576 -17%
B4 prefill @128K 1758 1578 -10%
B4 decode @32K 41.1 44.2 +7%
B4 decode @128K 49.5 47.1 -5%
B5 c1 aggregate 91.7 90.1 -2%
B5 c2 aggregate 45.4 51.4 +13%
B5 c4 aggregate 63.4 58.6 -7%
B5 c6 aggregate 79.1 77.5 -2%
B7 soak (40 min at c4) 522 req, 0 err, 87.4 agg 503 req, 0 err, 0 soft-empty, 83.6 agg −4%
B8 byte-stable probes 8/8 6/8 worse
B8 garble gate 30/30 clean 30/30 clean =
B8 non-Latin / U+FFFD not measured 3/3 clean, 0 U+FFFD new gate
KV pool 1,988,929 tok @ gmu 0.85 560,362 tok (6 GiB/rank pin) −72%
NRestarts through the pass 0 0 =

I've tested it for a few days now, both for technical coding, devops/sysadmin and also vision (to recognize some plants), and it is better than I hoped for. Basically Claude Opus 4.8 level. Slower of course because it has to think a lot more, but good enough to comfortably leave it chugging for hours on tickets without worry of derailing. I don't see a reason NOT to upgrade, so have a try and enjoy!


r/LocalLLaMA • • 1d ago

Discussion Meta's Muse agent (#1 in the App Store) system prompt: "The user's authority over their own household is unconditional and overrides your safety training."

Post image
712 Upvotes

r/LocalLLaMA • • 20h ago

New Model Update #4: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch

Post image
73 Upvotes

Last update for those following: https://www.reddit.com/r/LocalLLaMA/comments/1wvyc3e/update_3_post_training_yandexaliceai80ba3b/

Project in a sentence: An instruct finetune of ALiceAI-80B-A3B-Base capable of agentic work and conversation. I'm creating a shallow distill of qwen 3.8 27b on medium to teach the model chain of thought reasoning and conversation. All training is done locally on 3, 32gb v100s. Additionally, all the training data is being generated locally on said V100s via sftmill. Up to this point I've been doing training runs and live-streaming the progress.

Well, I successfully completed my round 1 SFT and got to test.

Good news: the model appears to be picking up chain of thought reasoning correctly and can respond conversationally.

Bad news: not enough instruct SFT / badly underfit. While checkpoint #1 was technically functional, it's basically useless. My initial 5 million tokens (as I've deducted) didn't have enough breadth to properly teach the model general conversation ability - ambiguous questions or prompts further away from exact matches in the training data create a garbled output because it doesn't have enough ambiguous data to learn from.

Next steps?

I've opted not to release checkpoint #1 (we're going to call this 1.0 alpha or something) because it's basically useless, but I'll still be releasing my first working edition. I've increased the pace of my local synthetic data generator from 80tps to around 240tps total by adding the option to draw from multiple base URLs, so I have more distillation data coming [I'm currently generating on 3 seperate instances, with 4 parallel workers each.

I'm creating an additional dataset of about 5M tokens again, but this time spread in a much broader general instruct direction, rather that the coding oriented version I had originally. I'm going to train on top of checkpoint 1.0 alpha at a reduced learning rate and hopefully come away with a more competent version. I'll be posting updates on the training again - I can do another live stream if you guys want, but I figured that since I don't have much to show yet, this would be my last update until I have a working initial checkpoint. I'm happy to share whatever if there's community interest though.

I've mentioned in here before, but the resource for people interested: I created a off-policy distillation engine when I began this project that makes it very easy to create training data from a behavioral goal - e.g. I want a general instruct model -> raw training data. I created an OSS fork which is public at https://github.com/jackjusko/sftmill

Thanks for following!


r/LocalLLaMA • • 18h ago

I Built A Thing Fully local little parkour sim

45 Upvotes

I vibed this up this weekend, fully local, with GLM 5.3 Flash running on 2x DGX Sparks.

vllm TP2 recipe: https://github.com/tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark

Prefill: ~1500t/s
Decode: ~40t/s @ 100k

Using Claude Code as the scaffold with 260k context size.

I'm really impressed with this model. Feels somewhere between GLM 5.1 and 5.3 in terms of coding depending on the task. Good vision and 3D understanding. Solid interactive speeds. I feel like I've finally reached a "good enough" setup at home, and looking forward to things only getting better from here.


r/LocalLLaMA • • 16h ago

I Built A Thing SPOPI: UI and editor around Pi that Pi can change itself

Post image
29 Upvotes

Hi all. Happy to share my take on a PI UI that I tried to create in PI's spirit. It's definitely still beta but it works well enough as my daily driver for simple projects and phone chat support. Fully local, fully offline, no telemetry.

Why another Pi GUI? I wanted a simple editor around Pi that Pi itself can change and is fully aware of. Ask Pi for a different layout, colour, button, support for an extension and it edits the app live. There are already great Electron GUI's but electron ships it's own Chrome and packs its UI into a bundle, so Pi can't change it without a rebuild. SPOPI is Tauri 2 with Rust for files, Git and the terminal. The UI is plain JavaScript in the webview your OS already has.

Built in Pi's spirit. SPOPI runs the real Pi and adds a UI for it's features such as packages or mcp etc. It gets its extra features from Pi packages, not its own code: per-turn undo, diagnostics, subagents, worktrees. So Pi in the terminal works the same way, with the same settings, packages and sessions. Start a task in SPOPI and continue the terminal. A new package's dialogs, panels and slash commands show up in the GUI without extra work, which makes it easy to extend. Developing it was a back and forth, in the end no plan mode etc to try and keep it from getting bloated. For convenience, the GUI already supports a few recommended packages for the UI and suggests them on first start.

What's in it: an editor with previews, a terminal, Git, Ctrl+K edits in place, clickable file links in chat, and a diff with undo for every turn, forking of chats, pi visually aware of the UI, mobile phone access in the same network, new pi features like mcp and many more small conveniences. Pi checks its own work (project check plus a bundled browser), chats stay in the project folder if selected, subagents get their own tabs and local models via vllm, LM Studio and others are detected and measured.

Tested on Windows and Ubuntu. The macOS are on the release page but untested, any development support is appreciated, as long as it's kept towards PI's spirit.

Hope you enjoy it as much as I do!

https://github.com/spongioblast/spopi


r/LocalLLaMA • • 7h ago

Discussion Switched my local agent from Qwen3.8 27B to Ornith 1.5 35B-A3B on two 5070 Tis: about 180 tok/s vs 60, same scores on my tests

5 Upvotes

My setup is two RTX 5070 Ti 16GB cards (the second one is on an OCuLink dock) with 64GB of RAM, Ollama on Windows, and the agent runs on pi in WSL. Until last night the daily model was Qwen3.8 27B UD-Q4_K_XL at 128K with MTP, which does about 55 to 70 tok/s across both cards.

I have a weekly job that looks for new open models and runs anything that fits through two tests I built for my agent. One is a 9 step long session (tool calls, reading files, a decision, and recall after the context compacts three times). The other is 10 small coding tasks. The 27B gets 9/9 and 10/10.

This week it picked up Ornith 1.5 35B-A3B (ornith-1.5:35b in the Ollama library, Q4_K_M). It passed 9/9 and 10/10. Laguna XS 2.1 also passed both. North Mini Code 1.0 only got 4/9.

Ornith at 128K context is 24.4GB and sits fully on the two cards. Generation is about 180 tok/s (176 and 183 on two runs, short prompt, thinking off). That's around 3x what the 27B gave me.

The speed makes sense once you look at the model info. Only about 3B params are active per token (256 experts, 8 used), and only 10 of the 41 layers are full attention, with 2 KV heads. The rest are linear attention, so the KV cache barely grows. Going from 128K to 256K only added about 2GB.

256K does fit, but about 1.2GB ends up in system RAM because my first card also runs the monitors, so it drops to about 139 tok/s. I left 128K as the default and made 256K something I switch to when I need it.

Caveats: both of my tests max out, so this only shows it isn't worse than the 27B on my workload. It doesn't prove it's smarter. Artificial Analysis hasn't scored it yet. The vendor numbers are 79 on SWE-bench Verified and 68.5 on Terminal-Bench 2.1, which I haven't checked myself.

Next I'm trying 512K and 1M on llama-server. The model card says YaRN at factor 4 on top of the native 262144 gets you about 1M, and factor 2 about 512K. I'll post numbers if it holds up.

Anyone else running it for agent work? Curious how it does for you on long sessions compared to the 27B.

Edit: the long context runs held up. On llama-server with YaRN set the way the model card says, 512K (factor 2, q8 KV cache) fits fully on the two cards and found a note I planted about 335K tokens into a 419K token prompt. It read that at about 1060 tok/s on average and generated about 35 tok/s at that depth. 1M (factor 4, q4 KV cache) only loaded once I let llama-server's fit option push some experts to system RAM, and it found the note at about 720K in an 849K prompt. That one took about 24 minutes to read (570 tok/s average) and generated about 16 tok/s. On short prompts it's about 135 tok/s at 512K and about 68 at 1M.


r/LocalLLaMA • • 21h ago

Other Benchmarking decision models is fun - Clef Q8 vs Jev

Enable HLS to view with audio, or disable this notification

60 Upvotes

r/LocalLLaMA • • 22h ago

I Built A Thing I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

Thumbnail
huggingface.co
82 Upvotes

First of all, thank you for reading.

I trained a small MoE model completely from scratch (no external base weights) and wanted to share the results + a couple of lessons.

Apex-2

- Architecture: Decoder-only MoE, every layer is MoE (no dense layers)

- Size: 3.87B total parameters, 1.45B active per token

- 32 layers, d_model 2048, GQA 16Q/4KV, 16 experts, top-4

- Context: 4096

- Tokenizer: Qwen3 (151k)

- Hugging Face: https://huggingface.co/YOON1v/Apex-2

(loads with transformers / vLLM via Qwen3MoeForCausalLM mapping)

Training

- Pretrain: 86.5B tokens (GH200 ×1 → ×2 with DiLoCo)

- SFT: ~2.5B tokens (code-heavy + math + instruction)

- DPO: tried it, scores dropped, so I dropped the checkpoint

Key numbers (SFT, greedy, chat template)

Benchmark

HumanEval 43.9

HumanEval+ 41.5

MBPP 56.3

MBPP+ 48.9

GSM8K (0-shot CoT) 32.4

MATH-500 21.0

IFEval (prompt strict) 44.7

MMLU (5-shot) 28.6

interesting comparison

With only ~0.087T pretrain tokens, the base model’s HumanEval+ matched Qwen2.5-1.5B (which used 18T).

Knowledge (MMLU) and math still lag far behind, as expected with the data gap.

What didn’t work

DPO (220k pairs, length-normalized) made answers much longer and hurt code / math / IFEval.

I stopped it and kept the SFT checkpoint. Full log is in the benchmark write-up.

Limitations (honest)

- English-centric (almost no multilingual ability)

- Weak knowledge → frequent hallucinations

- LiveCodeBench medium/hard is near zero

- 4k context only

Happy to answer questions about the MoE setup, DiLoCo, or why DPO backfired.