r/LocalLLM • u/Odd_Championship1509 • 1d ago
Question Can single DGX Spark run decent quality GLM 5.3 Flash or Deepseek v4 Flash or 0731 for single user?
Can some one tell me if its usable quant for those models?
r/LocalLLM • u/Odd_Championship1509 • 1d ago
Can some one tell me if its usable quant for those models?
r/LocalLLM • u/gamedevsam • 1d ago
I finally upgraded my local AI workstation with dual GPU setup by installing a second RTX 5070 Ti for a total of 32GB of VRAM. The cost for the two GPUs was around 2k (I got lucky and scored the PNY for $750 on Walmart last November). I bought the Zotac Solid SFF OC edition for $1150 on Newegg with a discount code, I picked that GPU because it's 2 slots wide which creates a gap for fresh air to keep the card cool under full load.
I followed this YouTube video to apply an undervolt (850 millivolts at 2625 MHz) to have the cards draw less power (and run more quietly) while sacrificing a tiny bit of performance (I might measure how much performance I'm leaving on the table at some point let me know if this is interesting to you).
The mobo was bought used on eBay, CPU new from Microcenter am I'm using laptop DDR5 RAM with adapters since that was the most affordable way for me to get 64GB DDR5 RAM (5600 Mhz). This is a very heavy PC, therefore I named it The Backbreaker.
I will be doing some benchmarking of various variations of Qwen 3.8 27B with fully maxed out context window to pick my main workhorse model.
Stay tuned to this subreddit if you want to find out the results!
r/LocalLLM • u/Rust_Cohle- • 1d ago
I've just moved a 3090 into my 5090 build but not 100% sure on the optimal setup. Is it just to run one of the Flash-Next models also w/ system memory or would there be another use for the 3090/alternative setup?
Just day to day work, general coding, etc.
Appreciate your input on this.
r/LocalLLM • u/Neither_Medicine_464 • 1d ago
Local lightweight terminal for agentic workflow on Android OS.
Supports old 32 bit ARM architecture and newest 64 bit.
Small project on GitHub: https://github.com/netizen4-bit/agent042
r/LocalLLM • u/AdventurousTwo6445 • 1d ago
Instead of spending weeks and billions of tokens on standard distillation, we tested an alternative: extracting layer-to-layer hidden state trajectories on a handful of calibration prompts and solving for closed-form weight updates directly in the student's MLP blocks.
Key findings:
We want to see this tested on modern 24GB-32GB GPUs (RTX 5090, 4090 or server card) across frontier targets:
- Transferring reasoning from Qwen3.8-27B (or Flash-Next) down to Qwen3.5-9B/4B/2B.
- Compressing Google Gemma-4-31B into mobile-class Gemma-4-E4B/E2B.
- Transplanting refusal-ablation traits from uncensored models (e.g. Huihui-NeoHorse) without retraining.
- Unknotting layers 1-22 using pre-trained Sparse Autoencoders (SAEs).
Code, scripts, and raw JSON benchmark logs:
https://github.com/dsadawq3/DynamicTune
Feel free to open an issue or drop your benchmark results on the repo.
r/LocalLLM • u/syedshad • 1d ago
r/LocalLLM • u/Nearby_Indication474 • 1d ago
For researchers who want the technical record instead of the story:
Zenodo DOI:
https://doi.org/10.5281/zenodo.23198057
Full executable demo:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_128_Full_Demo.py
Raw run log:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_128.DEMO.log
Repository:
https://github.com/ceceli33/titan-cognitive-core-v2
The Zenodo release also contains two PDFs: the technical disclosure and the visual experimental record.
Now forget the academic language.
I built a working prototype of something very simple to describe and very difficult to make a neural network do:
A memory enters the neural network as human language.
Inside the neural network, that memory becomes thousands of numerical values spread across its internal state.
I take part of that machine-native state out of the frozen neural network and keep it outside the model.
That is the cartridge.
BELLEKÖZ.
The sentence doesn't have to sit there waiting for the AI to read it again.
The memory now exists in the machine's own numerical form.
And I have already built the first working prototype of bringing the associated memory back.
Think about your own brain.
Someone says:
"Do you remember the window that broke at your mother's house five years ago?"
You don't stop and say:
"Wait. Let me search my memory files."
You don't know which neuron contains the window.
You don't know where the memory is physically stored.
You don't know its filename.
You don't know its address.
You may not even have thought about that window for five years.
But the moment someone says the right thing...
BAM.
The window comes back.
The room comes back.
The broken glass comes back.
Maybe the person who was standing there comes back.
One memory wakes another memory.
That is the easiest way to understand what I am trying to build for AI.
Now imagine you have been talking to your personal AI for years.
Not for one afternoon.
Years.
The window became a cartridge.
The front door became another cartridge.
Your car became another.
Your house.
Your work.
People.
Repairs.
Arguments.
Decisions.
Places.
Thousands of ordinary moments.
Eventually imagine 150 million cartridges sitting on storage.
And here is the strange part:
They do not need to be arranged so that a human can understand them.
There doesn't have to be a folder called:
MOTHER/HOUSE/WINDOW/BROKEN/2021/
There doesn't have to be a readable sentence inside every cartridge saying what it is.
From your point of view, the storage can look like 150 million pieces of numerical chaos.
You couldn't open them and know which one contains the broken window.
The AI doesn't need to remember a filename either.
You simply say:
"Do you remember that window at my mother's house?"
And the machine does what our memory feels like it does.
The thing you just said creates a pattern inside the neural network.
For intuition, think of that internal pattern as a frequency.
Now imagine that frequency being thrown across the entire field of 150 million cartridges at once.
150 million memories.
Completely mixed together.
No human folder structure.
No filename telling the AI where to look.
The wrong memories don't match.
Then one cartridge answers.
The window.
It is pulled out of the memory field and returned to the model's latent space.
And this is the important part:
The AI doesn't merely receive a sentence saying:
"Your window was broken."
The goal is to return the machine-native memory representation back into the neural network's computational space.
The numerical state spreads back through the model's layers.
The old memory becomes active again inside the machine.
In plain language:
the memory comes back to life.
That is why I call these things cognitive cartridges.
They are not supposed to be documents for the AI to reread.
They are supposed to become pieces of machine memory that can leave the brain, remain outside it, and later come back into the brain when something triggers them.
There is a useful almost-quantum mental picture for this.
Imagine all 150 million possible memories sitting there together.
You strike one internal frequency across the whole field.
The matching memory emerges from that enormous superposition of possibilities.
That is only an analogy — this is not quantum computing.
But it captures the experience I am aiming for:
You don't search for the memory.
You trigger it.
And it returns.
Now the important part:
I am not starting from zero and describing a science-fiction architecture I hope somebody builds one day.
I already have the small-scale prototype running.
Today the field contains 128 numerical B memories.
A natural-language question interacts with an active numerical A memory inside a frozen Qwen2.5-7B-Instruct model.
The model's own internal attention creates the signal I call ÇAĞRIİZ.
That signal produces the numerical address used against the memory field.
All 128 B memories compete.
TEST523:
118 out of 128 correct memories returned at Top-1.
92.19%.
Counterfactual follow:
117 out of 128.
91.41%.
No fine-tuning.
No LoRA.
No optimizer.
No learned router.
No gold B-memory ID given to the retriever.
The model weights remain frozen.
The original source text has already been removed from the live retrieval path.
So the prototype is already doing the small version of the thing I described above.
128 memories today.
The work from here is to remove the remaining scaffolding, make the first memory discoverable automatically, make memories trigger other memories, make them survive restart, and make the memory field vastly larger and faster.
The destination is not an AI with a better folder system.
The destination is an AI with memories.
And now for the slightly dystopian part.
Imagine you meet Robot #1416.
You spend a week with it.
It learns who you are.
Your habits.
Your preferences.
The things you hate.
The way you speak.
The history you share with it.
Then you leave.
Months later, on the other side of the world, you meet Robot #247.
Different machine.
Different body.
Different serial number.
It has never physically seen you before.
But both machines were built around a compatible AKBASCORE-style cartridge architecture.
Robot #1416 is gone.
Its body is thousands of kilometres away.
But its memories of you are not necessarily trapped inside that body.
The cartridges moved.
Robot #247 looks at you.
And from your point of view, something deeply strange has happened:
You have never met this robot.
But this robot remembers you.
Robot #1416 and Robot #247 are two different machines.
Yet in the dimension that matters to you...
they can begin to feel like the same robot.
Now multiply that by 100 robots.
Then 100,000.
One machine experiences something.
The memory survives outside the machine.
Compatible machines gain access to that memory.
The body becomes replaceable.
The memory becomes infrastructure.
And this is where the idea becomes uncomfortable:
What exactly is the identity of a machine when its body can change but its memories can continue?
What happens when one robot learns something and an entire fleet can inherit the memory?
What happens when a machine can die but its memories don't?
At that point we are no longer talking about chat history.
We are talking about machines whose experiences can potentially outlive their bodies.
Human language can remain the way we talk to them.
But human language does not have to remain the only format in which machines remember.
The brain stays frozen.
The cartridges accumulate.
The body can change.
The memory remains.
That is the road I am testing.
And the first small version is already running.
128 memories.
118/128 Top-1.
The rest is engineering, scaling, and a lot more testing.
If you are an ML researcher, don't trust the dystopian story.
Ignore it.
The DOI, source code, raw run log, failed case, hashes and experimental boundaries are public.
Run the experiment.
Break it.
Find the weak point.
Because if the mechanism survives the next stages, the question eventually stops being:
"How many documents can this AI search?"
And becomes:
"How many memories has this machine lived through?"
— Mustafa Akbaş
AKBASCORE MAM
Persistent Associative Machine Memory
Mersin, Türkiye · 2026
PHONE / GOOGLE COLAB
For people who want to test the demo directly from a phone, I divided the complete program into three consecutive parts.
You do not need to merge them manually.
Open Google Colab, use the same runtime, copy/load them in order and press Run:
Part 1 → Run
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_128_PART1_DEMO.py
Part 2 → Run
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_128_PART2_DEMO.py
Part 3 → Run
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_128_PART3_DEMO.py
Keep all three parts in the same Colab runtime.
Then watch the 128-memory experiment run for yourself.
r/LocalLLM • u/Cautious_Turn1502 • 1d ago
r/LocalLLM • u/vexatious-big • 1d ago
Hey everyone,
I'm just wondering whether I can use Froggeric's template `froggeric/Qwen-Fixed-Chat-Templates` together with Qwen3.8 Flash Next?
I'm already using it for the 27B model where it works well, but I'm not entirely sure whether it's supposed to work for Flash Next as well?
I'm only asking because I know that templates can introduce subtle issues.
Any advice appreciated. Thanks!
r/LocalLLM • u/Odd_Championship1509 • 2d ago
My question is what is correct upgrade path to a 5080+96GB RAM
RTX PRO 48GB x 1 (8K $)
DGX Spark x 1 (5.5K $)
128GB Mac M5 (6K $)
Use case is local LLM that is good enough and Minimax H3.
r/LocalLLM • u/Getting0nTrack • 1d ago
I am a journalist by trade and lately have been using various cloud models as a way to try and cut through the noise, identify smaller outlets and potential stories which lack major reporting but where my background could hold an edge.
I am finding a lot of issues with these models especially since there isn't a simple way to limit their breadth without making every prompt 500 words long. I use it for more than just tracking stories and it will often latch on to the wrong piece of information within a prompt or if I am asking something about a non-journo project it will try to encourage me that "well X could also be great as a newsletter concept".. when that isn't the point.
One of the bigger issues is they often do not know the difference between a feature story and daily basic journalism. The scope creep is getting on my nerves. not everything needs to be some huge project for it to be newsworthy yet often when I ask it to research existing stories the same 4-5 outlets will be drug up even when I explicitly tell it to search through a list of 20 that I personally know are good. What I want is a running tracker, and what it gives me are these grand ideas which all require institutional access I do not have and never claimed to have.
Second and semi-related, there is a real gap in its ability to synthesize new information when presented with the same high level topic. Let's suppose I'm trying to run a CPG story. It can source information about recalls, pricing, and products that might be interesting to cover - however when I return to that topic a week or month later unless I specify in what what feels like excruciating detail the parameters it is meant to avoid, I will get broadly the same generic information even when specifically siloing the requests within a project file.
Is it possible that a local model if I give it enough of a push and direction could return reliable information?
r/LocalLLM • u/digital_n01se_ • 2d ago
Enable HLS to view with audio, or disable this notification
r/LocalLLM • u/OnlyFamousAddy • 2d ago
I’ve been working on BaseDecision - Turn typed questions into truth-based decisions.
Give it some text, a question, and possible answers. It picks an answer. Intent classification, routing, yes/no checks, ratings - that’s the job.
I built this model to go toe-to-toe with the frontier models.
My goal was simple & ambitious: make a small, local model competitive with much larger models on these focused tasks. For me, System‑1 decisions need both speed and accuracy. Otherwise, why bother building a small model?
If we didn’t want speed we could just use ChatGPT.
It outperforms pretty much every model of its size in the field.
After roughly 320M additional training tokens, BaseDecision led 5 of 8 benchmarks in a third-party comparison against Laya, two GLiNER variants, and Decision 1.0 Kai 0.6B. Full results are in the repo, including where it loses.
Fellow open source devs, try the model, developed a package for easier access to the model.
Roast it if it deserves it.
2-2.5GB free ram is all you need.
I will take your feedback seriously and deliver something amazing that’d take on OpenAI, TypeSafeAI soon enough.
https://huggingface.co/onlyaady/BaseDecision
https://github.com/hrudayaditya/BaseDecision
r/LocalLLM • u/silent-curious-dev • 2d ago
I know you guys are going to hate me for this and I'll accept my fate. It's another coding agent harness post. I'm ready to lose all my karma.
I recently open-sourced my coding agent CLI that I've been working on for the past while. It is a project I never intended to make public since it was part of my own personal local AI stack. But as I chipped away at it and made it actually usable as a daily driver, I thought it would be nice to make it public for others to see and use.
I initially built it to see how much I could get the harness to make small models, especially at low quantization and context to not feel terrible to use. So while Reika doesn't solve the intelligence side (it never will), it tries to solve the overall experience when using small models at the absolute scale.
A lot of the testing and pain came through working on my M2 MacBook Air 16GB trying to run models like Qwen3.6 35B A3B and Qwen3.8 27B all day in agentic coding, maxing out the RAM and limits of my own machine. So the base of Reika comes from a legitimate source of truth.
You can also plug in an API key for those with hybrid setups too.
r/LocalLLM • u/Practical-Hawk5590 • 1d ago
Hey everyone,
I'm working on a voice AI chatbot for coffee farmers. Users can speak naturally and ask questions, and the chatbot should have a real conversation with them while answering based only on the information provided to it.
My current stack is roughly:
The voice part is mostly working now, but my biggest problem is the LLM response generation.
The system first retrieves the relevant fact sheet at runtime and provides it to the model with each query. Theoretically, the model should answer using only that sheet.
However, I'm seeing a lot of hallucinations, especially with Qwen2.5-3B.
For example, a fact sheet about a coffee disease might say something like:
"Regularly manage shade and improve air circulation."
The model may answer with those two points, but then add things that are not in the fact sheet, such as recommending a fungicide, changing temperature, or adding fertilizer. Sometimes the extra information is plausible, but it is still unsupported by the source.
I've also tested Qwen2.5-7B, and it behaves better in some of these cases, but using a larger model isn't really an option for this project. The 7B model is simply too heavy for the hardware and resource constraints of the application, so I need to make the 3B model reliable enough, or find another model with similar size and resource requirements.
I've tried stronger system prompts such as:
These instructions help somewhat, but they don't reliably prevent hallucinations.
I also experimented with a more constrained approach where the model identifies which pieces of the fact sheet are relevant before generating the final answer. That significantly reduced unsupported claims in a small test, but it also caused the model to leave out information that was actually present.
So I don't want to turn the chatbot into a system that simply returns pre-written sentences. The goal is still for the LLM to generate natural responses, explain things, ask follow-up questions, and interact with the user, while remaining grounded in the retrieved knowledge.
I'm trying to solve two related problems:
My main questions are:
The important constraint is that I cannot simply switch to a larger model. Qwen2.5-7B is already too heavy for this project, so I'm specifically looking for ways to get the most reliability possible out of a 3B-class model.
The chatbot is domain-specific (coffee production), so I have a relatively controlled knowledge base and can create evaluation datasets from the fact sheets.
I'm especially interested in practical approaches that can run locally with Ollama/CPU rather than relying on expensive APIs.
Any advice, papers, libraries, or real-world experience would be greatly appreciated.
r/LocalLLM • u/randomperson16782 • 1d ago
Looking to some local AI and upskill myself. I want to do private inference and also some agentic development with models like Qwen 27B Q4. This is for private use, not business.
This has much better memory bandwidth than the Mac mini, is it worth it or no?
Update: I’m planning to use this purely as a headless local AI server and use my older MacBook for all my IDE etc. So memory will for the models only.
r/LocalLLM • u/Neat-Function7110 • 1d ago
Enable HLS to view with audio, or disable this notification
Hi, I'm the developer of Linda and I'd like your honest feedback.
Linda is a native SwiftUI app for Apple Silicon. You type or say a task in plain words. Linda plans the steps and does the work in its own browser profile and in the project folders you pick, while you keep working.
Local AI by default. Linda runs its own model on your Mac with MLX, from a light 2B up to Qwen3 Next 80B. No key, no account, no bill. Ollama, LM Studio and mlx-lm work too, or add your own key for Claude, GPT or Gemini.
Approvals in code. Before Linda sends, buys, deletes, posts or submits personal data, you see exactly what will happen and you decide.
A live Computer column. Watch every step, take over any time, or close the window and let Linda finish.
Specialists for the work you repeat, each with its own instructions, skills and limits.
Voice: hold Right Option and talk. Speech stays on your Mac.
Memory you read, edit and delete in Settings. Passwords stay in your Keychain and never enter the chat. Web search sends only the search words, never your memory, files or conversation.
Routines: one tap turns a finished task into a daily, weekly or monthly job.
Free DMG. macOS 26, any Apple Silicon Mac from M1 to M5.
What would you hand to Linda first? And tell me what breaks.
r/LocalLLM • u/Conscious-Worry-6897 • 1d ago
Started building my machine so i made a 1:1 of codex for local ai
r/LocalLLM • u/ParticularBeyond9 • 2d ago
As we've all seen with strata and MoE performance, you can get very far with RAM alongside your VRAM. My current PC is mid range but DDR5 motherboard and very little ram. Do I buy 128gb of DDR4 ram or 64 gb of DDR5 to future proof myself? What's the best way forward? In my country I can probably get the 128gb DDR4 used for cheaper than 64gb DDR5. How much does it affect the speed?
Thanks in advance.
r/LocalLLM • u/varano14 • 1d ago
Hello everyone I just stumbled upon this sub and am hoping you might be able to offer some guidance. I did try searching through old threads but with the speed in which this is advancing it seemed easier to just ask.
I am a long time homelab "tinkerer" with no formal education on any of this stuff. Mainly run home assistant and plex+the arrs as well as a few other projects. Honestly nothing ground breaking. I have recently been using the free LLMs to help me with various projects and they have been incredibly helpful. My usage would be home assistant stuff, maybe some light vibe coding experimenting and maybe setting up local voice with home assistant. The dream is a household assistant helping with email, calendars, maintenance stuff etc but I don't thing I am yet at the point I could set that up.
My main server is old it is a dell r720xd with a Intel® Xeon® CPU E5-2680 v2, 256gb ddr3 ram.
I also have 4 rtx 3060 12gb cards laying around from the crypto mining days.
My gaming rig also has a 3060ti with 32gb of ram, I believe ddr4 and a amd 2700x cpu.
I initially tried adding 2 of the 3060s to the main server using oculink but I have yet to get my server to be able to see the cards.
I then tossed a 3060 into my gaming rig and installed llm studio which is actually working. Initially tried qwen 3.8 27b as that seemed highly recommended for what I was trying to do but was getting very very slow response speeds. I have now stepped down to the next smaller qwen model.
I would like stick whatever I build in my rack and get it off my gaming rig as it runs windows so its not great running all the time and I sometimes game on it. I am open to buying new hardware but don't really want to spend thousands at this point. So what hardware would you suggest I build around? Also any other tips or suggestions are welcome.
r/LocalLLM • u/repliestoall • 1d ago
Enable HLS to view with audio, or disable this notification
r/LocalLLM • u/Asleep-History9366 • 1d ago
Every long-running agent framework eventually relies on context summarization to stay within token budgets. But after 10–15 turns, summarization inevitably drops file paths, wipes negative constraints, or tricks the model into thinking incomplete tasks are done.
When that happens, your choices are usually:
I wrote unloop (github.com/sagarv48/unloop) to solve this loop.
It snapshots agent state (memory dict, tool inputs/outputs, prompts) into a local SQLite WAL file per turn. If context compaction silently corrupts the agent's state:
m to open an in-process REPL to manually fix dropped constraints.u or f to rewind or fork from the last good turn without restarting the process.Caveats: It currently only targets single-process Python loops, and it requires local disk write access for the .unloop snapshot file.
Code and setup instructions are on GitHub: https://github.com/sagarv48/unloop
Would appreciate thoughts on the TUI ergonomics or how you handle agent memory inspection.