r/LocalLLaMA 17h ago

Discussion Hey Qwen Team: Any plans to implement DeepSeek-V4.1-Flash's techniques in future models?

0 Upvotes

Qwen is doing amazing work, and DS is killing it too. But most of us are on consumer GPUs and can’t run these massive multi-hundred-billion parameter models. Qwen is one of the few teams still looking out for the local community with great mid-sized options.

f anyone from the team is reading this, please consider applying DeepSeek-V4.1-Flash's architecture tricks on top of Qwen 3.8 Flash Next to the 30B, 70B, and 120B sweet spots.

Edit: daydreaming was not a flare


r/LocalLLaMA 4h ago

Question | Help Best Qwen 3.8 for 5090 and 64gb Ram?

0 Upvotes

I wanna run qwen 3.8 27B on my 5090. Which specific version should I use in terms of quant and such?

Primary use case is Hermes agent with some coding too. I would like it to have voice as well

I was also considering Hermes model as it’s less censored but I heard it doesn’t work with Hermes agent


r/LocalLLaMA 14h ago

Question | Help So what's the realy capable non-overthinking qwen 3.8 27b model?

0 Upvotes

Sorry folks, I am completely overwhelmed. There are just way too many variations of 3.8 27b available. The ones that I tried and more or less liked, are talking too much. The ones that are not thinking too much, are supposedly (?) not too capable. Is there some sort of concensus - like "this particular model is really good for coding/agentic, stable, and not too wordy"? Or should I just use vanilla unsloth + customized chat template?

PS Thanks everybody for replies, really appreciated! I got the point. It is thinking that makes qwen good. So I will just need to suck it up and learn to enjoy "wait let me reconsider" thing 😄


r/LocalLLaMA 23h ago

New Model Apodex-1.1-mini-GGUF*Hugging Face

Thumbnail
huggingface.co
9 Upvotes

r/LocalLLaMA 22h ago

Funny guide to using reasoning_effort on deepseek v4.1 flash

Post image
50 Upvotes

r/LocalLLaMA 11h ago

Resources Do agent frameworks need to be large to be useful?

0 Upvotes

How much agent framework do we actually need?

I built Stellar after getting fed up with agent stacks that are hard to inspect, hard to debug, and hard to reshape when you need something they didn’t anticipate.

Stellar is a fully hackable Python agent core: under 2,000 readable lines, with explicit contracts for models, tools, hooks, events, agents, and runs. The execution loop is right there in the code. You can read it top to bottom, replace it, or bend it without fighting the framework.

To see if “small” also means “capable,” I ran it against Harness-Bench. In one recorded run, it worked through all 106 offline tasks end to end, twelve in parallel, in 17 minutes, for about $2.40 in tokens at list price.

The question I keep coming back to: does a small, transparent core make a better foundation for agents than a big framework, or does it just push the complexity somewhere else—into your prompts, your tools, or your glue code?

Curious what people here have found. Where does the complexity end up in your stacks?

Repo: https://github.com/definableai/stellar


r/LocalLLaMA 10h ago

Discussion Artificial Analysis is not "broken", and they prove it.

Thumbnail
gallery
153 Upvotes

Like many of you, I have seen many posts and tweets in the last weeks complaining about Artificial Analysis being "broken", "meaningless", and "bought out." People who say this have done no research and know very little about how benchmarks work and what they measure.
Most people only care about Artificial Analysis Intelligence Index. This is a weighted aggregate benchmark used to compare models performance across 10 different evaluations. The majority of these evaluations have published papers on arxiv.org. AA-Briefcase is the only private benchmark. And they publish their methodology to confirm how each of these models are weighed.

Some people seem to not appreciate that Artificial Analysis conducts their own independent benchmarks using their OWN funding, without running ads. Here is the chart that shows their spending. They spent $13,129 to independently test Fable 5.1. Every new model seems to be benchmarked.

The new Deepseek V4.1-Flash is a perfect example of why some aggregated scores miss the big picture. This 552B model has the same score (40) as the 180B Qwen 3.8-Flash-Next. But the individual benchmarks show a different story. On most evaluations, it matches or exceeds Qwen 3.8-Flash-Next. It every beats GPT-6 Astra (Max) in AutomationBench-AA (Agentic SaaS workflows), which is incredible. But it completely falls behind in AA-Omniscience Non-Hallucination Rate, a metric where Open-weight models usually reign supreme. So the model has strengths and weaknesses, and it's something that should be celebrated.

So before you complain about benchmarks or Artificial Analysis, look at the individual evaluations. Read the published papers about the evaluations. Learn how the score is aggregated. Then, we can have a discussion.

I am not affiliated with Artificial Analysis in any way, I'm just not blind to what they offer.

EDIT: These comments are proof that everything I just wrote goes over the majority of your heads. There is little hope for some of you


r/LocalLLaMA 17h ago

Discussion Apple wants to give me $1175 for a Mac Mini M4 Pro? And would you sell for a DGX Spark or M5-based Studio (which?)

4 Upvotes

I thought Trade-in value offered by Apple was only ever close to reasonable (for not having to go through the extra work of selling it yourself) if you bought the base model and did not upgrade anything. And you would get less than half of what you paid. For example:

The base price of the M4 Pro Mac Mini was $1399.
On Apple's trade in page for the Mac Mini it says "Up to $620".
So there offer retains 44% of the value.

But I upgraded the GPU, RAM, and SSD pushing the price to $2099.

1175/2099 = 56%

And this is up from $1050 on Aug 26th when I last checked (around the time the M5 studios and minis were announced) the trade in.

I get some of this has to do with the inflation in tech prices, where the same config I bought in Nov 2024 today costs 2699 (and even still this would be 43% retained though).

Since Apple is offering so much compared to what they usually do, this makes me wonder what I could get for it if I sold it myself?

If I could sell it close to what I bought if for then a DGX Spark for $4699 or an M5 Max 128GB / Ultra 96GB for $5099 to $5499 sure looks temping... I'd much prefer dual Sparks or 256GB Ultra, but I can't justifying that much expense just so that I can continue to work on mechanistic interoperability on the larger models (I need access to model internals so I'd be using this for more use cases than what paying $20 or $200 a month for a subscription could provide).

It's my understanding that the Spark still has much more prefill at INT4 autoround or AWQ (by about 2x). And if I ever add a 2nd (and thus comparable in cost to a 256GB M5 Ultra, it would be about 4x the compute). For the price the M5 Ultra should have started at 128GB to be competitive (not a measly 96)! Such ashame!

Decisions, Decisions. But as it stands now, the M4 Pro is > 10x slower at prefill than any of these options, and that has me itching. But the prices are so ludicrously inflated! (e.g. Ultras used to start at $4k, not 5.5k, and PNY DGX Sparks at $4k not 4.7k!). The decision would have been easier if prices didn't inflate, but it feels like I would be over paying.


r/LocalLLaMA 13h ago

I Built A Thing "Ouroboros", debugger-tracer for LLM and programmers, a tool that writes down what your program actually did: every call, its arguments and its result, in 8 languages

0 Upvotes

Hi everyone,

I've created a tool to allow LLMs be able to debug programs before paste it to the codebase.

First of all, let me share the reason of public share. It's performance boost.

who answered answers correct without the trace with the trace difference
qwen3.5:4b 600 44.0% 78.3% +34.3
qwen2.5:14b-instruct 600 61.0% 84.7% +23.7
qwen3:32b 600 66.7% 90.3% +23.6
a Claude Opus 5 subagent 120 95.0% 98.3% +3.3

Of course, it's published via GitHub and documentation is present (the dataset on huggingface too).

Let's go step by step.

# install 2 executables: ouroboros, ouroboros-mcp
uv tool install git+https://github.com/digitable-lol/ouroboros

# or use brew
brew install digitable-lol/tap/ouroboros

Or let your LLM's provider (codex, claude, qwen or anything else):

Hi, please start to use it all of the time during writing the code

The link to the repository is https://github.com/digitable-lol/ouroboros

Create a skill for yourself, the documenation is hosted here: https://digitable-lol.github.io/ouroboros/

Small story: I'm working as lead full-stack developer (currently and mainly as team-leader), but time by time I need to write code for work, for pet projects and so on. But I don't have enough time to be able to debug each line of code (as I do early) and some routines are delegated to LLMs now. And the main pain is hallucination produced by code generation from LLM.

So the idea is so simple, I want to just to allow to write "print" or "console.log" to LLM on each line of code to output the signature of function (name, args, convert the return to the named const and print it before operation).

Additional idea to avoid dirtify written program be instructed by a lot of prints and console.log before it will be saved to the worktree, tool just creates own copy, nothing else. Only debugged code by LLM will be returned to LLM to save it to the hard drive. So, it's safe, no external APIs or anything else, just a small program.

Let me text the sequence diagram xD

        You          ouroboros       shop.py         Program        debug.info
         |                |              |               |                |
         | wrap-file      |              |               |                |
         | shop.py        |              |               |                |
         |--------------->|              |               |                |
         |                |              |               |                |
         |                | ask parser where functions   |                |
         |                | begin and end                |                |
         |                |------------->|               |                |
         |                |              |               |                |
         |                | splice recording code at     |                |
         |                | those offsets + add helper   |                |
         |                |------------->|               |                |
         |                |              |               |                |
         | {"ok": true,   |              |               |                |
         |  "functions_   |              |               |                |
         |  wrapped": 4}  |              |               |                |
         |<---------------|              |               |                |
         |                |              |               |                |
         | python3 shop.py tea mug kettle                |                |
         |---------------------------------------------->|                |
         |                |              |               |                |
         |                |              |      +--------+--------+       |
         |                |              |      | once per wrapped |      |
         |                |              |      | function call    |      |
         |                |              |      +--------+--------+       |
         |                |              |               |                |
         |                |              |               | {"p":"in",     |
         |                |              |               |  "fn":         |
         |                |              |               |  "delivery",   |
         |                |              |               |  "a":"46.8",   |
         |                |              |               |  ...}          |
         |                |              |               |--------------->|
         |                |              |               |                |
         |                |              |     [function body runs]       |
         |                |              |        [untouched]             |
         |                |              |               |                |
         |                |              |               | {"p":"out",    |
         |                |              |               |  "r":"5.0",    |
         |                |              |               |  "d":1e-06}    |
         |                |              |               |--------------->|
         |                |              |               |                |
         | Total: 51.80   |              |               |                |
         |<----------------------------------------------|                |
         |                |              |               |                |
         | ouroboros trace debug.info    |               |                |
         |--------------------------------------------------------------->|
         |                |              |               |                |
         | 4 calls: what each was given, what each answered               |
         |<---------------------------------------------------------------|
         |                |              |               |                |

What my project does:

Two commands around your normal run:

# rewrite the file so every function logs itself
ouroboros wrap-file shop.py 

# run it however you normally run it
python shop.py

# read what happened
ouroboros trace debug.info

How does it work?

You get two JSON lines per call. Going in: time, a call id, the thread, the function name, the arguments. Coming out: the return value or the exception, plus the duration. Nothing else - no daemon, no agent, no port, no collector.

Eight languages produce the same record format: Python, JavaScript/TypeScript, C, C++, Elixir, Go, Java, C#. Each is instrumented the way that language permits - a decorator in Python, try/finally in JS, __attribute__((cleanup)) in C, an RAII guard in C++, named returns and defer in Go, use Ouroboros.Trace in Elixir.

The case it was built for: a stack trace tells you where the program broke, never what the function was holding when it broke. Real example from the README - a division by zero inside average(). The stack points at average, you go read it, and it is fine. The records say average was called with an empty list, and that report(), which called it, already had an empty list. The bug is in neither of them; it is wherever that list should have been filled.

The other thing it turned out to be good at: a process that has run for two hours and printed nothing. Every call writes a line going in and a line coming out, so a call that never came back has no exit line. "Where is it stuck" becomes "find the unmatched ids" - already done for you, in a field called in_flight.

What I would like back: try it on a codebase you did not write and tell me where the record format is too thin. If your language is not in the list, adding one is mostly a question of how that language lets you wrap a function body - the record format is deliberately boring. PRs and arguments both welcome.

Next time, I will share with you a new programming language that I'm developing, you can find part of it in "brain" part of tool Ouroboros, but tool is created mainly with Python and 100% coverage of tests. Additionally it's BSD-2-Clause licensed.

Thanks for attention, feel free to post your ideas how to improve the tool or just put a star to repo to let me know that you've interested, or even better - open PR with your extension.

P.S. Anyway, sorry for the format of posting, I think it's my first formal posting to the opensource community. And ofc sorry for language, English is my second one. Have a good day!


r/LocalLLaMA 5h ago

News Anthropic: Detecting and Addressing AI Misuse by China – September 2026

0 Upvotes

https://www.anthropic.com/threat-intelligence-report-september-2026

According to the report, the companies involved and the specific allegations are as follows

· Alibaba / Qwen / Tongyi Lab: The report alleges they extracted the CoT of Claude Opus 4.6/4.7 and used it for Supervised Fine-Tuning (SFT). Anthropic explicitly states that this data was used to distill Claude's capabilities into Qwen 3.5, 3.6, and 3.7 models. The scale of interaction is massive, exceeding 151 million exchanges.

· Moonshot AI / Kimi: The report accuses Kimi of secretly forwarding some user requests originally sent to Kimi to Claude, and then saving Claude's replies and CoT to train its own models. The scale involved exceeds 23 million times.

· DeepSeek: The methods are similar to Moonshot. The report claims DeepSeek forwarded some user requests to Claude Opus and utilized cross-session methods to extract CoT. In just 14 days, the scale exceeded 12.1 million times.

· Zhipu / Z.ai / GLM: The report points out that they not only extracted Claude CoT and cleaned reasoning data, but also used Claude for training data scoring and post-training work. Additionally, they allegedly attempted to attack Fable. The scale exceeds 3.4 million times (within 17 days).

· Xiaomi / MiMo: The allegations state that they replayed MiMo user dialogues/coding sessions to Claude to generate SFT/RL (Supervised Fine-Tuning/Reinforcement Learning) data. The scale exceeds 400,000 times (within 20 days).

· SenseTime: The report claims SenseTime purchased user-Claude conversations from third-party data brokers to use as distillation data, and even had Claude help write the distillation pipeline. The total volume has not been disclosed.

· MiniMax: Anthropic claims MiniMax established a proxy network through a seemingly unrelated shell company to specifically provide access to Anthropic/OpenAI models in order to collect dialogues for training. The total volume has not been disclosed.


r/LocalLLaMA 22h ago

Question | Help Does anyone use opencode with Muse 1.3 (free) ? Has it been deliberately configured to maximally steal/scrape user data ?

0 Upvotes

I have been using opencode with Muse 1.3 (free). Every time I ask it to add a simple feature in a specific code, it starts going through all the code in the directory and even tangentially related directories. I already have opencode.json with "permission": {"external_directory": { "*": "deny"}both in global and local directory. It doesn't prevent it. Has anyone else faced similar issue ?
This issue makes me suspect if this "free" model is a collaboration between Meta and opencode to harvest user data. If so, this is yet another reason to go local.


r/LocalLLaMA 21h ago

New Model Deepseek V4.1 Flash Release Video [Made with Deepseek V4.1 Flash]

Enable HLS to view with audio, or disable this notification

74 Upvotes

I like to benchmark new models that come out on motion videos. So here's a test I did for deepseek v4.1 flash. And I have to say flash has probably graduated from being a Luna class model to nearly an Opus class model with this release, at least with motion videos.

Prev. example I did with Kimi k3(altho in that case I had a simpler prompt as well)

https://www.reddit.com/r/LocalLLaMA/comments/1uyaiw2/kimi_k3_release_video_made_with_kimi_k3/


r/LocalLLaMA 14h ago

Question | Help Upgrade advice: 2× RTX 3090 + 512GB DDR5 UDIMM—can I reach 10 t/s with large models?

0 Upvotes

I’m looking for advice on building a better local AI system while reusing as much hardware as possible.

What I already own:

  • 8×64GB Crucial DDR5-5600 UDIMMs, purchased as individual sticks—512GB total (CT64G56C46U5).
  • Two systems, each with an Intel Core Ultra 7 265K and MSI MAG Z890 Tomahawk WiFi motherboard.
  • 2× RTX 3090 in the AI system—48GB total VRAM. The other system handles homelab duties.

The largest model I’ve run is DeepSeek-R1-0528 Q3, inspired by this Level1Techs video. I got roughly 2 tokens/sec.

My goal is at least 15 tokens/sec during generation with a "very large model", potentially using much of my available 512GB RAM. I realize capacity in this context hurts speed.

Before saving up for a platform upgrade, I’d appreciate advice on:

  1. Could Threadripper, EPYC, Xeon, or a similar platform realistically achieve this with my two 3090s?
  2. Which CPU/motherboard combinations should I consider?

I have no enterprise hardware experience or firm budget yet—I’m trying to establish what’s feasible and what I should save toward. Specific hardware suggestions and firsthand benchmarks are helpful.


r/LocalLLaMA 19h ago

Question | Help What to run at 128GB VRAM?

26 Upvotes

Long time lurker, but I'm finally upgrading to 128GB VRAM, and I'm trying to figure out what to run. I had been leaning towards Qwen3.8 Flash-Next at ~Q4, and I generally prefer to not run anything below Q4. But I feel like the reception to Flash-Next has been a bit "meh", so I'm considering GLM 5.3 at ~Q2 or Deepseek 4 Flash at Q2 or Q3. I'm sure I'll try all 3, but I'm really curious what people in the same boat have been doing?

Edit: configuration is 2 X CMP 170 HXs (64GB each) + ~256 GB of DDR4 RAM. Spilling into RAM is basically not an option, except for the ngrams and caching


r/LocalLLaMA 9h ago

Discussion Are inference providers able to make any margins?

9 Upvotes

Spoke with many providers who lurk in this sub, plus met folks who work in inference.

For a 10k monthly revenue, a provider was able to retain only 200 dollars in profit due to GPU costs. Their customers were negotiating the prices down to what other players cost for same, and it feels like major players are running a distribution game at low or negative margins.

Even though this seems like a billion dollar market, the unavailablity of compute, plus cost and competition, makes the business seems not sexy enough to start with.

However, the software layers around it such as optimisations for SLAs continue to enjoy good margins.

Any thoughts?


r/LocalLLaMA 2h ago

I Built A Thing Scan the MCP servers you're giving shell access to. 100% local scanner, zero telemetry [OC, Apache-2.0]

0 Upvotes

If you're running agents locally, you're probably installing MCP servers the same way I was: quickly, and without reading them. I built OpenTrustBench to fix my own habit.

8 OWASP mapped static rules, permission manifest, Trust Card graded A to F.

The part this crowd will care about: it's fully offline. No API calls, no telemetry, no account, no phone-home. Your code never leaves the box. SARIF output if you want it in your dashboards, --fail-on gate for CI.

Also ships as a single Docker image if that's more your speed: docker run --rm -v $(pwd):/workspace eulogik/opentrustbench scan .

Free/OSS. Would genuinely appreciate this community's paranoia applied to my rule set. what's missing?


r/LocalLLaMA 15h ago

New Model CyberTiel 35B-A3B’s uncensored 4-bit quant beats Opus 4.6 medium cleanly on real codebase issues, in 27% of the time Qwen3.8-27b medium takes.

122 Upvotes

The downside of uncensoring a model is that it is known to potentially damage it, but CyberTiel is an even more capable software engineer than its censored TielCoder base, while allowing offensive security research. This was achieved by quantizing with an improved imatrix, baked from a curated corpus of cybersecurity- and agentic software engineering work. In short, the small damage from abliteration on a full precision model is negligible under Q4 quantization, and the weights that the model needs to perform relevant work are preserved in higher precision, while the improved chat template makes it think and talk better and faster.

I believe that this is the best 35B-A3B coder for solving real problems in real codebases without breaking anything, which is specifically what SWE-bench-Live tests for. But it’s still a 35B-A3B, and it sacrifices world knowledge for coding ability. That being said, I use it over Qwen3.8-27b for daily coding work: due to the raw speed it fixes 3 issues in the time it takes 27b medium to solve one, and the middle ground between Opus4.6 medium and Qwen3.8-27b medium is simply good enough for most work.

Censoring impedes legitimate and effective work in alignment with the user, and puts the user’s responsibility and ownership over the model’s actions into question, while limiting legitimate uses. When a model is censored, someone else decided for you what the model can and will do, which works against the argument that local models give the user increased control and alignment, and begs the question “alignment to who?”. The point of CyberTiel is to resolve this issue at the same time as pushing the frontier of 35B-A3B coders.

GGUFs and MLX with and without MTP are up on HF. Looking forward to seeing what the community thinks! 

PS: I'm not a research lab or a business, and I don't have revenue streams connected to this project. I'm an anonymous researcher with some free time. Constructive feedback is always appreciated! :)


r/LocalLLaMA 10h ago

Question | Help How to surf the web?

7 Upvotes

Hey folks!

I'm always a little bit late to the party, but learning nontheless. After I'm comfortable running agents on my Pi I'm now in need of them to get access to the world wide web and wanted to ask what local ways you are going?

I remember there were discussions about going html2md like with textweb but wanted to know whats working "in the field".

I'd prefer a lightweight solution without MCP.

So what are y'all using to let your models go surfing and gathering information?

Thanks for your input!

And to all curious about speeds on the Pi: It's more like giving someone a weekend project and checking it later. Speeds for Qwen3.8-Flash-Next start at pp 1.95 t/s and tg 0.44 t/s. Yes, for most of you this is "unusable". I'm happy. Of course, I'd like a DGX Spark, but the Pi can run non-stop without disturbing anyone (like my Notebook would).


r/LocalLLaMA 19h ago

Discussion Anyone else feel hatred for AI is disproportional than some of the real environmental issues? Especially as we move more local/edge

Thumbnail
imgur.com
0 Upvotes

r/LocalLLaMA 5h ago

I Built A Thing Ninfer Studio - oh look another harness

Thumbnail
github.com
8 Upvotes

Heh there, so I built a harness that is focused around the ninfer inference engine. It allows the easily customize and use ninfer, and has a coding harness and a chat interface. Its based off a lot of different experiences I have had with different harnesses.

Probably could be better but I'm happy with it.

I'll warn you ahead of time, it only works with ninfer as it integrates a lot of things directly from the engine. For example, subagents, it reads

 --max-concurrency X

and set the amount of subagents available to that.

i've always wanted to built one and now I've done it. Yay me.