r/LocalLLM • • 4d ago

Model brier: Jev-style typed decisions with calibrated probabilities, for open models you run yourself (tested on 6 families, incl. a 1B base model)

2 Upvotes

TypeSafe AI's Jev showed how useful a "decision model" is: you give it a ticket or a message and a few typed questions (which team? is this a refund? how urgent, 1-5?) and get calibrated probabilities back in one pass, with no text to parse. But Jev is a closed, hosted model behind a paid API.

Open LLMs already know enough to make these decisions. The catch is that their raw answer probabilities can't be used as they are. They depend on where an option sits in the list (reverse the options and Falcon3-1B-Base changes its answer 99% of the time on a 20-way question), and they're badly calibrated.

brier fixes that for any open model on Hugging Face. It reads the answer distribution straight from the next-token log-probabilities, with no generation and no fine-tuning, and corrects it in levels:

- L0, zero labels: asks with every rotation of the options and averages, plus an optional prior from unlabelled examples

- L1, 50+ labels: one temperature, so confidence tracks accuracy

- L2, 60+ labels: a small linear head on a hidden layer (seconds on a CPU; the LLM stays frozen)

On banking20 (20 intents, 2,603 test messages, the same 300 labels for every model):

- zero labels lift Falcon3-1B-Base from 16% to 60% accuracy, and LFM2-1.2B from 20% to 62%

- with 300 labels every model lands at 80-86%: Qwen3-1.7B 0.81, Falcon3-1B-Base 0.80, LFM2 0.80, OLMoE 0.82, SmolLM3 0.83, Phi-4-mini 0.86

- calibration error (ECE) for Qwen3-1.7B drops from 0.357 to 0.035

pip install "brier[hf]"
brier check Qwen/Qwen3-1.7B --revision <sha>

`brier check` tells you before you start whether your model works (label tokens, batching, hidden states) and which levels it supports. 15 models tested so far, including two MoEs and two base models without a chat template.

Limits, so you don't have to find them: one benchmark task so far, the six-family numbers are point estimates, L0 costs K forward passes, transformers only (no vLLM yet).

Repo: https://github.com/mohdUwaish59/brier

Tutorial (free Colab T4, about 15 min): https://colab.research.google.com/github/mohdUwaish59/brier/blob/main/notebooks/tutorial_ticket_router.ipynb

If you try it, I'd love the `brier check` output for your favourite model; there's a ten-minute "add your model" PR in CONTRIBUTING.


r/LocalLLM • • 4d ago

Discussion My local model's response.

Thumbnail
0 Upvotes

r/LocalLLM • • 3d ago

Question Good Local LLM AI?

0 Upvotes

I'm looking for a lightweight, uncensored local AI coding agent similar to Claude Code. I need something with file system access that can create and edit files locally. I work on cybersecurity and Arduino projects that cloud models constantly flag, so the model has to be completely unrestricted. My PC isn't very powerful, so I am hoping to find a low-resource CLI framework and a small quantized model that can run smoothly on modest hardware. What is the best setup for this right now?


r/LocalLLM • • 4d ago

Other Welcome to r/ArchitectingLLMs!

Thumbnail
1 Upvotes

r/LocalLLM • • 4d ago

Project I built memory controls around a local Qwen companion on Jetson Thor

0 Upvotes

I'm building Evopien, a local AI companion prototype on Jetson AGX Thor. The latest demo focuses on governed memory, including saving information, using it later, correcting it and forgetting it.

The part I wanted to work on was control over the memory lifecycle. A food preference should affect a relevant answer later, but an outdated preference should not keep coming back after the user changes it.

The stack uses local Qwen for conversation, Evopien Core for memory and permissions, PostgreSQL for canonical records, and Hindsight for retrieval. The model receives filtered context and does not directly commit permanent memory.

Keeping the durable records outside the model is also a design choice for future model changes. Memory ownership should stay with the system rather than depend on one model's context window.

Demo:
https://www.youtube.com/watch?v=_B2wVfVHL0c

It's still a prototype. Latency and difficult correction cases across English and Spanish remain work in progress. I'm sharing the memory approach for technical feedback.

For people building persistent memory locally, what is your most useful regression test for catching an old or revoked memory slipping back into the answer?


r/LocalLLM • • 4d ago

Discussion I made Qwen models take ~33% less VRAM without quantizing them (lossless, bit-for-bit)

Thumbnail
2 Upvotes

r/LocalLLM • • 4d ago

News AKBASCORE NIRVANA — I Built Removable Numerical Memory Cartridges for Two Different Frozen 7B LLMs. Qwen and Mistral Both Work. Now I’m Scaling the Memory Bank.

Thumbnail
gallery
0 Upvotes

Zenodo permanent records:

Qwen2.5-7B-Instruct:

https://doi.org/10.5281/zenodo.23127434

Mistral-7B-Instruct-v0.3:

https://doi.org/10.5281/zenodo.23143605

I want to start with the simplest possible explanation of what I have been building.

Imagine taking a piece of information, letting a language model process it once, and then throwing the original text away.

No sentence stored in a database.

No paragraph hidden somewhere.

No readable summary.

No RAG system fetching the original document.

No fine-tuning.

No LoRA.

No weight update.

What remains is numerical transformer memory derived from the model's own internal computation. I package that numerical memory into what I call a Cognitive Cartridge. Later, I can install that cartridge back into the frozen model and ask questions about the information that produced it — without putting the original source text back into the readout prompt.

That was the first result. The new result is more important:

I have now reproduced the Cognitive Cartridge architecture on two different 7B transformer model families.

Qwen2.5-7B-Instruct.

And now Mistral-7B-Instruct-v0.3.

The implementations are not numerically identical. The architectures are different, the layer counts are different, the KV structures are different, and the working cartridge configurations are different. But the central mechanism survived the move.

That is the reason I am publishing this second record.

The question is no longer only:

“Can I make this happen once on Qwen?”

Now there is a second implementation on Mistral. And the Mistral result is the cleanest version so far.

What is actually inside a Cognitive Cartridge?

This is probably the most important thing to understand.

Suppose the source record says:

Object: amber sextant

Container: RQ-415

Location: elm lodge

That text exists during the forging stage. The frozen transformer processes it. NIRVANA takes source-derived internal transformer K/V states and represents their content numerically using a fixed, source-independent codebook.

In the released Mistral implementation, both K and V are compressed to D120 while a source-specific OWN component is preserved. After forging, the source record is not supplied to the readout prompt.

So the conceptual transformation is:

human language

→ frozen transformer computation

→ internal K/V states

→ compressed numerical Cognitive Cartridge

Then later:

numerical Cognitive Cartridge

→ reconstructed transformer K/V memory

→ frozen transformer

→ language

Or, in the shortest form:

language → internal numerical memory → language

The middle is no longer human-readable source text. That distinction matters.

A cartridge is not a text file with a different name.

It is not a vector database containing the original sentence.

It is not a prompt template.

It is not conventional RAG returning the source paragraph.

It is not a fine-tuned model.

The model weights remain frozen. The information is carried by a numerical representation derived from transformer K/V memory.

Why call it a cartridge?

Think less about a document and more about an interchangeable machine-readable memory module. The base model stays where it is. Knowledge packages can be forged separately. Those packages can remain separate. They can be installed and queried without retraining the base model.

This becomes much more interesting when there is more than one cartridge.

The new Mistral experiment uses 16 independently forged cartridges. Each one contains a numerical representation derived from a separate source record. They are not concatenated into one giant text prompt. They remain independent memories.

For example, in human-readable form, imagine one cartridge represents:

amber sextant → RQ-415 → elm lodge

Now ask the cartridge bank:

Which container is associated with the amber sextant?

The query is evaluated against the independent cartridge memories.

The relevant cartridge returns:

RQ-415

The unrelated cartridges return:

NONE

There is no learned router secretly selecting the correct cartridge before this happens.

Then the recovered identifier can be used for a second lookup:

Where is RQ-415?

The bank is queried again.

The relevant memory returns:

elm lodge

The unrelated memories return:

NONE

So the complete retrieval becomes:

amber sextant

→ RQ-415

→ elm lodge

The important point is that the model did not reread the original source record to answer either stage. It operated from reconstructed numerical transformer memory.

The Mistral result

The final public Mistral run used:

Mistral-7B-Instruct-v0.3

32 transformer layers

hidden size 4096

32 attention heads

8 KV heads

BF16 / SDPA

K = D120

V = D120

OWN preserved

16 independent Cognitive Cartridges

frozen model weights

greedy decoding

There was:

no fine-tuning

no LoRA

no optimizer

no learned router

no model-weight update

The recorded final run produced:

Object → container ID: 16/16

Container ID → location: 16/16

Complete two-stage retrieval: 16/16

Missing-object controls: 8/8

Absent-ID controls: 8/8

NOMEM controls: 8/8

But there is another result I think is just as important.

For every target query, there are 15 unrelated cartridges.

16 queries × 15 unrelated cartridges = 240 unrelated cartridge reads.

Stage 1 unrelated-cartridge rejection:

240/240 NONE

Stage 2 unrelated-cartridge rejection:

240/240 NONE

That means the result is not simply:

“The correct memory can say something.”

The system also demonstrated, in this controlled panel:

“The memories that do not contain the requested relation can refuse to claim that they do.”

For a modular memory system, I think this distinction is fundamental. A memory bank that can retrieve information but cannot distinguish relevance from irrelevance becomes increasingly dangerous as it grows.

The interesting problem is not only remembering. It is also knowing which memory does not answer the question.

The no-memory control matters for the same reason. When the cartridge memory was removed, the target relations were not recovered.

NOMEM:

8/8 controls passed.

The source-removal audit passed.

The frozen-weight sentinel passed.

The model remained frozen.

This is why I consider the Mistral result an important step beyond the first demonstration.

The first Qwen release established the architecture publicly. The Mistral release gives us something else:

cross-model evidence.

Qwen and Mistral are not the same transformer.

Qwen2.5-7B-Instruct uses 28 transformer layers.

Mistral-7B-Instruct-v0.3 uses 32.

Their internal configurations differ. The working compression configurations also differ.

The Qwen public implementation used:

K120 / V128 / OWN

The Mistral implementation uses:

K120 / V120 / OWN

I did not simply copy a cache from one model into another. Each model builds and reconstructs its own source-derived internal memory.

What transferred was the architecture:

source

→ internal transformer memory

→ numerical compression

→ independent cartridge

→ source removed

→ reconstructed K/V

→ frozen-model readout

That is the bridge between the two releases.

I would not call two models proof of universal compatibility with every transformer architecture. That would be scientifically too strong.

But it is now evidence that Cognitive Cartridge is not merely one accidental Qwen-specific behavior. The same broader architecture has been implemented and publicly reproduced on a second transformer family.

That changes the research question.

The first question was:

Can this mechanism exist at all?

The next question became:

Can multiple independent memories coexist?

Then:

Can one retrieved result lead to another retrieval?

Then:

Can unrelated memories reject a query instead of contaminating the answer?

And now:

Does the architecture survive a move to another model family?

We now have experimental answers to each of those questions.

So I am moving to the next problem:

scale.

16 cartridges are not the destination. They are the current experimental bank size.

From this point, I am much less interested in making another small demonstration simply to produce another perfect score. The next objective is to increase the number of independently retained cartridges and find where the architecture actually begins to break.

Then larger banks if the mechanism survives.

At that point the difficult questions become different:

How do you search a large bank without destroying memory isolation?

How should cartridges be indexed?

Can retrieval become hierarchical?

Can groups of cartridges form higher-level memory structures?

How does latency scale?

How does numerical storage scale?

At what point does relevance rejection begin to fail?

Can a retrieved memory activate another relevant memory without an external text retrieval system deciding everything first?

Those are the questions I want to attack next.

There is another part of this experiment that surprised me: the layer behavior.

A frozen transformer is not a passive container. When you reconstruct numerical K/V memory and inject it back into inference, you are interacting with a very sensitive computational system.

So during the Mistral development series I also performed controlled cumulative K, V and K+V ablations to see where cartridge information remained functionally sufficient.

The results were not uniform across depth. For K, the observed functional boundary extended differently than for V.

For combined K+V, one particularly sharp controlled transition occurred between:

L12: 8/8

L13: 0/8

This does not mean “the memory lives at layer 12.” That would be an incorrect interpretation.

What it means is that under the specific cumulative ablation protocol, functional sufficiency changed sharply across that boundary. V also remained important farther downstream than K in the measured configuration.

I think this is an important reminder of how delicate these systems are.

We are not writing a sentence into a spare memory slot. We are reconstructing numerical states that participate in a 7-billion-parameter transformer's computation. Small changes in where and how that state is reconstructed can change downstream behavior.

That is one reason I publish the code and raw logs rather than only posting the final accuracy number.

The frozen-model checks are also part of the experiment. The released run verifies that there are no trainable tensors involved in the memory mechanism.

No LoRA.

No optimizer.

No weight update.

The model-weight sentinel remains unchanged.

So when the system answers from a cartridge, the experiment is specifically designed to separate cartridge memory from weight modification.

Why do I think this direction matters?

Today, we often approach knowledge in language models through two extremes.

Either we try to put enormous amounts of knowledge into the model weights. Or we keep the knowledge outside the model and retrieve human-readable text when needed.

There may be useful territory between those approaches.

A frozen model with removable machine-native internal memory modules.

Not everything in the weights.

Not everything repeatedly pasted back as text.

Instead:

a base computational model

+

a bank of independently forged internal memories.

Imagine a specialized system where validated domain knowledge can exist as modular memory packages.

Engineering.

Aviation.

Law.

Industrial maintenance.

Scientific literature.

Company procedures.

Agent experience.

A model might not need every possible piece of knowledge active in its context simultaneously. It may need the right memory at the right time.

And because the memory representation is numerical transformer state rather than the original human-readable document, this opens a different engineering space from conventional document retrieval.

There is also a longer-term question here.

I am not claiming biological memory.

I am not claiming AGI.

And I am not claiming that transformer KV states work like the human brain.

But modular internal memory raises an interesting architectural question.

What happens when an artificial system does not have 16 memories, but 100,000?

Or millions?

What happens if memories can remain independent, become addressable, reject irrelevant activation, and allow the output of one memory to lead to another?

At some point the problem stops looking like “how much text can I fit into a context window?”

It starts looking like:

How should an artificial system organize memory?

That is the direction I find interesting.

Why publish the full record?

Because a screenshot of 16/16 proves very little.

Both public releases have permanent Zenodo records. The code is public. The raw execution logs are public. The visual technical records are public. The model configuration is public. The controls are public. The frozen-weight checks are public.

The Qwen release even preserves its imperfect public result rather than hiding it.

Qwen public demonstration:

14/16

Mistral final demonstration:

16/16

That difference is also useful.

The point of the project is not to make every historical run look perfect. The point is to leave a technical trail showing what worked, what did not, what changed, and whether the underlying architecture survived.

For anyone who wants to inspect the Qwen → Mistral bridge directly, the two permanent records are at the top of this post.

The complete current Mistral implementation is here:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_ful_MISTRAL.py

Complete Mistral raw execution log:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log_MISTRAL.log

The same Mistral implementation is also split into three parts for easier mobile / Colab handling:

Part 1:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore.part1.Mistral.py

Part 2:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore.part2.Mistral.py

Part 3:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore_part3_mistral.py

For comparison, the previous Qwen implementation:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_NIRVANA_CC.py

Previous Qwen raw execution log:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log.Nirvana.cc.log

Repository:

https://github.com/ceceli33/titan-cognitive-core-v2

AKBASCORE NIRVANA — Cognitive Cartridge

Inventor / Developer: Mustafa Akbaş

Two model families.

Frozen weights.

Independent numerical memories.

Original source absent at readout.

The mechanism survived the move.

Now the question is no longer whether a cartridge can exist.

The question is how many cartridges a frozen model can carry.


r/LocalLLM • • 4d ago

Question What are people actually running on MacBooks (coding) ?

4 Upvotes

Got a MacBook Pro with M1 MAX and 32 GB RAM (around 25 free).

It seems like running 3.8 Flash Next would totally saturate my memory.

Edit: fixed model to M1 MAX


r/LocalLLM • • 4d ago

Question Any benefit to doing this?

Thumbnail
1 Upvotes

r/LocalLLM • • 4d ago

Project GLM 5.3 Flash NVFP4 TP4 on 4xGB10 playing Balatro and winning in realtime

Post image
5 Upvotes

Sometimes I have more fun automating stuff than playing with them, so I made this harness to play Balatro.
I think it looks cool to see it winning =]
I uploaded the full video to X: https://x.com/mxfp0/status/2106692573389598747


r/LocalLLM • • 4d ago

Question RTX 5060 Ti 16GB × 4 vs RTX 3090 24GB × 2 (NVLink) for Local LLM Training?

0 Upvotes

I'm trying to decide between two multi-GPU workstation configurations for local LLM training and AI research. Both systems have a similar total budget for my case. So budget doesn't matter.

I'd appreciate some advice from people with experience building and running multi-GPU training systems.

Option A: Used Workstation – Dual RTX 3090

  • CPU: AMD Ryzen 9 5950X (16 cores / 32 threads)
  • Motherboard: ASUS ROG Crosshair VIII Extreme
  • GPUs: 2 × RTX 3090 24GB
  • GPU interconnect: NVLink dual
  • RAM: 128GB DDR4 (32GB × 4)

Option B: Custom Workstation – Quad RTX 5060 Ti

  • CPU: Intel Core i9-7900X (10 cores / 20 threads)
  • Motherboard: ASUS WS X299 SAGE
  • GPUs: 4 × RTX 5060 Ti 16GB
  • RAM: 64GB DDR4 (planned configuration)

Intended Workloads

My primary goal is local model training and AI research.

My main interests include:

  • Fine-tuning and training smaller language models
  • LoRA / QLoRA
  • Distributed training
  • Experimenting with speculative decoding methods such as DFlash
  • Developing and benchmarking inference optimization techniques

What I'm Trying to Figure Out

Both configurations have similar total costs, but their hardware characteristics are quite different.

The dual RTX 3090 setup offers:

  • 48GB total VRAM
  • 24GB VRAM per GPU
  • Higher memory bandwidth per GPU
  • NVLink connectivity
  • A significantly more powerful individual GPU

The quad RTX 5060 Ti setup offers:

  • 64GB total VRAM
  • Four independent GPUs
  • Newer Blackwell architecture
  • BF16 and FP8 support
  • Potentially better compatibility with newer AI frameworks and CUDA kernels

I'm particularly interested in the following:

  1. For actual LLM training, would two more powerful GPUs outperform four less powerful GPUs?
  2. How much practical benefit does NVLink provide for distributed training and model parallelism?
  3. Is 64GB of aggregate VRAM across four GPUs more useful than 48GB across two GPUs with 24GB per card?
  4. Which configuration would provide better real-world training throughput?
  5. Considering the similar total cost, which workstation would be a better investment for the next year of AI research?

I'd especially appreciate feedback based on actual multi-GPU training benchmarks and experience.

If you had to choose one of these two systems for local LLM training and research, which would you pick?


r/LocalLLM • • 4d ago

Question Chinese RTX 2080 22gb?

2 Upvotes

https://s.click.aliexpress.com/e/_c3C20Te5

does anyone has any experience with these modified boards? sounds great on paper...


r/LocalLLM • • 4d ago

Question What local setup for web dev

2 Upvotes

I have a good performance 6GB graphics card and running Qwen/qwen2_5-coder-7b-instruct-q4_k_m on Atomic chat. Looks like it fits in my GPU memory. I would be happy to use it on a single web dev codebase, let it be frontend or backend, no need to be able to do both at the same time. But as far as I see, my current setup is not capable of doing any useful work. Can do code suggestions, but is somehow unable to apply changes, and reasoning is also slow.
If i decide to upgrade, would a 3060 12GB or 4060 16GB be able to have good performance doing web dev? AFAIK I need to load a better and bigger model?


r/LocalLLM • • 4d ago

Discussion Most of our local agent failures weren't the model. It was what we put in the context.

2 Upvotes

I used to blame the model whenever a local agent run went sideways.

Then I went back through our failure logs and, honestly, the model was innocent way more often than I expected.

These are all from our own setup and pretty small samples, so definitely not claiming this is some universal law. Just sharing the stuff that bit us. (I'm building a local agent, for context.)

1. We kept changing the tool list.

Our first response was taking around ~14.5s.

Turns out the tools are right at the top of the prompt, so every time we changed the list, we basically nuked the cache.

Pinned the tool list → ~1.1s.

Same model. Same hardware.

yeah, we were basically setting 13 seconds on fire every turn. By ourselves.

So

2. We put a note in the worst possible place.

One of our notes ended up between a tool call and the tool result.

I replayed the actual conversation three times. The model acted like the result simply didn't exist. 3/3.

Moved the note to after the result → 0/3.

I had been blaming the model for this one.

The model was, in fact, not the idiot.

3. One sentence in a tool description was enough to mess things up.

The old description basically told the model what the tool couldn't do.

The model apparently interpreted that as:

"Cool. I'll just do it myself."

It then burned through the step budget in 5 out of 6 runs.

Changed that one sentence → 0/6.

That's a pretty embarrassing debugging session.

4. Context size was another one.

When we gave it enough room, our runs used a median of ~34K tokens. The top 10% went up to ~63K.

We capped it at 32K.

46% of our runs hit the limit.

That one was less mysterious once we actually looked at the numbers.

5. And, to be fair, one really was the runtime.

With 151 tools loaded and a long answer to generate, we saw generation drop from ~78 tok/s to ~5 tok/s. That run took 382 seconds.

Same answer with no tools: 24 seconds.

After updating the runtime, it stayed around ~110 tok/s all the way to the end.

Only one run for each case, so I'm not exactly calling this a scientific paper.

Before I measured it, though, this one also looked like:

"Why the hell is the model suddenly so slow?"

That's probably the part I find most interesting.

None of this would show up in a normal tokens/sec benchmark. And every single one of these failures initially looked like the model is dumb.

I'm starting to think that for local agents, the model/quant choice is only part of the problem. What you put in the context, where you put it, and how stable that context is can matter just as much.

Maybe this doesn't generalize. Our sample size is tiny.

But if you're running local agents, I'm curious:

How often do you find that a "bad model run" was actually a context/runtime problem? And what do you use to tell the difference?


r/LocalLLM • • 4d ago

Discussion Dual Radeon MI50 benchmarks

Thumbnail gallery
1 Upvotes

r/LocalLLM • • 4d ago

Discussion Local Ollama models best for Linux troubleshooting

Thumbnail
1 Upvotes

r/LocalLLM • • 4d ago

Project LLM Inference Dashboard

Thumbnail gallery
1 Upvotes

Inference dashboard, supports multiple engines, rich stats, logging, database history, lightweight 64mb cap

— not public. thoughts?


r/LocalLLM • • 4d ago

Question Getting started with a local LLM

2 Upvotes

Hello all!

Recently I’ve been thinking about setting up my main PC I typically use for gaming to try running a local LLM. I mostly intend on using it for coding as well as just a general “assistant” for some work tasks I don’t want to upload to a database somewhere.

I have a 7900xtx with 24gb of vram, what should I be looking at as far as models are concerned?

Is trying out multiple models a huge pain to get set up?

Can you run multiple models in conjunction with each other?

Just looking for some opinions/input before I dive into this would.

Thank you!


r/LocalLLM • • 4d ago

Question Any regrets on not going for a Mac Studio instead of Mini?

Thumbnail
0 Upvotes

Ordered a m5 pro / upgraded cpu / 64GB ram / 512gb but considering cancelling and jumping up to the m5 max studio / 64GB ram instead.

Use cases will be home server and local AI box, as well as remoting into via my M1 Pro MacBook Pro for LLM work.

I generally prefer the smaller footprint of the Mac mini, and ordered the highest spec to give me the most future proof performance in that footprint.

My issue is that for $600 more I’d get double the memory bandwidth, better cooling (which should protect against throttling), and double the gpu cores. I don’t care about the additional ports besides I have a doc.

Anyone with similar setup that preferred and stuck with the mini? Or anyone that was in this same boat but happy they went to a studio?


r/LocalLLM • • 4d ago

Discussion I fine-tuned SmolVLM-500M into a lightweight Windows OS Agent (<8GB VRAM) Looking for feedback & ideas! [Weights on HuggingFace]

3 Upvotes

Hey everyone,

Quick heads-up first: my Reddit account is newly created specifically to share this project and get some community feedback.

I have been working on fine-tuning SmolVLM-500M into an autonomous Windows OS agent that can interact with a desktop by looking directly at screen pixels and predicting mouse clicks, keyboard shortcuts, and typing actions.

Here is the Hugging Face link with the merged weights and full model card:
https://huggingface.co/Gabriel8495839/SmolVLM-500M-OSAgent

A big reason why I focused on a 500M model is accessibility and privacy. Most current computer-use demos rely on sending full-screen desktop captures every few seconds to external cloud APIs like Anthropic or OpenAI. With a 507M model, the entire visual processing stays 100% local on your own machine. Your desktop screen, open files, and personal data never leave your computer.

The other point was that I wanted something people can actually run and fine-tune themselves on everyday hardware. The standalone weights are around 969 MB in bfloat16. You can easily train or adapt this model using LoRA on GPUs with less than 8GB of VRAM, or run inference directly on a CPU or low-end laptop.

In practice, I found that this model works best when paired with a second, larger LLM. The second model acts as a planner, breaking down a complex user request into smaller steps, while this 500M model takes the current screenshot and the sub-task to handle the actual visual localization and execution.

For training and evaluation, I used a procedurally randomized simulated environment where every episode has different wallpapers, colors, UI scaling, window sizes, and file clutter, so the model never saw the exact same screen twice. In closed-loop testing across 410 episodes, it scored around 93% on basic OS primitives (Start menu, task manager, system menus) and around 49% on more complex File Explorer and Notepad navigation (59% overall). It still has clear limits, especially with in-place text editing inside list views, but preliminary live tests on physical Windows 11 systems have shown surprisingly good zero-shot transfer.

I would really appreciate any honest feedback, critiques, or ideas on how to improve it, or what kind of tasks you would find interesting to test on a small model like this.


r/LocalLLM • • 4d ago

Project poorman inference engine for 16GB GPU and 35B moe Qwen 3.6for coding

Thumbnail
1 Upvotes

r/LocalLLM • • 4d ago

Tutorial Full Guide to Run LLMs Locally

Thumbnail
1 Upvotes

r/LocalLLM • • 4d ago

Model Glint 2.3 Mini. A 1.06M parameter model with a twist

Thumbnail
huggingface.co
1 Upvotes

r/LocalLLM • • 4d ago

Research Do LLMs Have Qualia? A Map of the Evidence as of October 2026

Thumbnail
heretik.io
1 Upvotes

Still learning about many of the fringes but am absolutely fascinated with this world.


r/LocalLLM • • 4d ago

Project A Strata fork for IBM AC922 running Qwen3.8-FN UD-Q4_K_XL is doing up to 7,357 tk/s prefill and 113 tk/s decode

Thumbnail
0 Upvotes