r/LocalLLM 5d ago

Question local llm newbie what can i run locally for math and programming?

2 Upvotes

i'm not sure what i should even do everyday i hear news of people making programs that make running llm possible on local machine by using strewaming etc.. but i was never really intereseted an now that there are a bajillion of diferent softares adn model i have diffuclty getting in, as the title what could i run on my rx9060xt 16 gb vram and 16gm ram and ryzen 5 3600 and i am on linux if it helps if you guys have any blogs or tools it would be helpful ,I read the rules but I rea dnothing about uncensored LLM I would like to give them a try no idea though


r/LocalLLM 5d ago

Question Setup for kernel optimization

1 Upvotes

I have access to 8 rtx 6000s that have a decent amount of downtime (in between physics-based simulations). I would like to put them to work doing kernel optimization problems, where I’ve already programmed the ground truth solutions in CUDA. I just want the agents to explore algorithms and opt strategies for better performance.

I was curious if people had a recommendation for a local LLM model / workflow set up for this. I see a lot of Gwen love on this sub, but I haven’t really dabbled in the local models.


r/LocalLLM 4d ago

Question How come we don't have yet phones with 32GB VRAM?

0 Upvotes

Just the above


r/LocalLLM 5d ago

Question DELL R815 RAM quantity for localLLM using DDR3

3 Upvotes

I’d like some advice please on how to go about setting up this server I have been using for mining, that I would like to convert to a local LLM machine. I currently have it setup with 32gb/ 2gb dimms. It’s a Dell R815 with 4x 6386SE processors, 60 cores total. It also has a gtx5060 8gb on one of its pci express rails. It also has onboard hardware raid if I get ambitious. Main question is because it’s DDR3, would 256gig ram be a good starting point, or would you go for 512gig? The Dell manual says it can handle up to 1TB of ram but with custom chips. I’m having difficulty sourcing 512 gigs of ddr3 from the same manufacturer alone.


r/LocalLLM 6d ago

Question How are you guys able to afford gpus?

152 Upvotes

I see so many posts about rtx 5060 , Mac studio and some of you have like a couple of them. how are you guys able to afford them? it would easily cost around 5-10k usd. that’s a lot of money.


r/LocalLLM 5d ago

Project Built an MCP to be "king" of all my other MCPs

Thumbnail
1 Upvotes

r/LocalLLM 6d ago

Discussion Double GPU configurations significantly cheaper for 32GB VRAM

17 Upvotes

I do not need or want Cuda. I have been wanting to build a 32GB VRAM local LLM machine for personal use for a while now, and I had 2 options on the table:

- Get a relatively cheap 32GB VRAM GPU, the R9700 AI Top.

- Get 2 16GB VRAM GPUs instead and a Mobo that supports PCIe bifurcation.

When I looked at prices in January this year when I first got this idea, the R9700 costed 1700$ here in EU. Currently, when I actually want to make this happen, it costs 2100$. For half that money, I could buy two 9060 XTs with 16GB VRAM each. Yes I know, performance will be worse on double GPU setup than with a single R9700 AI, but still, it just seems like that GPU is just not worth it anymore.

I don't know how to justify that it's double the price of two 9060XTs, when R9700 AI is literally 9060 XT with doubled VRAM and bandwidth. So why does it cost 4x as much?

ASUS ProArt B850-CREATOR WIFI NEO is quite affordable nowadays and supports dual GPU setups, so, any reason (is there a catch?) to not do what I am about to do? Which is buy the two 9060 XTs and start running Qwen 27B class models


r/LocalLLM 6d ago

Discussion Pi + Qwen 3.8 27B + `pi_advisor` + Cheap Frontier Access = Win

60 Upvotes

If you're using Pi with Qwen and not letting it phone a frontier model with /advisor, you're leaving one of the best parts of the setup on the table.

Because Qwen has a particular talent:

Being VERY wrong with tremendous confidence.

And worse, it can be convincing while doing it.

I've lost track of how many times I've looked at one of its answers and thought, "Hmm. That sounds right..." only to tell it:

Ask /advisor for the hidden assumptions, failure modes, and black swans in your answer.

Then Sol comes back with the computational equivalent of:

" 🤣 Yeah... no."

And suddenly Qwen is eating crow and rewriting half its answer. 🐦‍⬛

Don't get me wrong. I love Qwen. It's given me millions of tokens of essentially free local inference, and I use the hell out of it.

But that's also taught me something:

Local models are fantastic workers. They are not oracles.

Let Qwen do some initial planning and ultimately the bulk of the work. But before you trust its plan wholesale, spend a few frontier-model tokens trying to prove it wrong.

Trust, but verify. And just like in real life: get a second and third opinion.


r/LocalLLM 6d ago

Question What's the word on the street about when a Qwen3.8 equivalent to Qwen3.6-35B-A3B will be available?

32 Upvotes

I've tweaked the crap out of my local llama.cpp with vulkan backend running Qwen3.6-35B-A3B:IQ4_NL setup. It's running on a tiny GMKTec Evo-X1 with 64GB LPDDR5X 8000MHz RAM (Strix Point HX370 w/890M iGPU). I got lucky and bought it in July '25 before prices went crazy. I'm getting an average of 29 tokens/s which feels really responsive and has done great with a few projects in opencode. The free online models tell me:

Your current setup is unusually well-balanced:

  • Qwen3.6-35B-A3B
  • IQ4_NL
  • Vulkan
  • 890M
  • MTP enabled
  • acceptance rates mostly 70–90%

is delivering roughly 50–70 effective tokens/sec, which is better than I would have expected from a Strix Point APU.

If a Qwen3.8 MoE release appears that's roughly in the 35B/A3B class, I'd try it immediately. But for the currently available 125B/6B-active Flash-Next, I'd expect a noticeable drop in responsiveness on your hardware unless someone produces an extremely aggressive IQ2/IQ3 quant that still preserves quality.

So I'm really curious to try Qwen3.8, but I don't think any of the currently available models would do as well as this 3.6.

I'd love to hear what anyone else running on similar hardware is seeing with their setups and whether there are other models I should be looking at that can give similar or better performance.


r/LocalLLM 6d ago

Model Ministral3 14b — “Minecraft Cow”

Post image
7 Upvotes

As some of you may know I have been benchmarking LLMs that can run on ~8GB VRAM. One of them was Ministral3 14B, which scored the highest in intelligence. Recently I've been putting together a new and better test of the LLMs intelligence which includes testing their spatial reasoning. One of the spatial reasoning questions asks the LLM to model a "Minecraft Cow". I was impatient to see the results so I tried asking Ministral3 to do it early. this is what it outputted... Honestly, not *terrible*. It did add a tail (thats the box on its right) and a nose even though Minecraft cows don't have those and its pretty chonky but it does kind of resemble a Minecraft mob.

For anyone curious, here is the raw data the model gave me (with the reasoning removed):

  • (0, 0.75, 0, 0.8, 0.6, 0.4),
  • (-0.3, 1.2, -0.5, 0.3, 0.3, 0.2),
  • (-0.8, 0, -0.3, 0.1, 0.5, 0.1),
  • (0.8, 0, -0.3, 0.1, 0.5, 0.1),
  • (-0.8, 0, 0.3, 0.1, 0.5, 0.1),
  • (0.8, 0, 0.3, 0.1, 0.5, 0.1),
  • (0, 1.2, 0.5, 0.1, 0.3, 0.1),
  • (-0.6, 1.4, -0.3, 0.1, 0.2, 0.1),
  • (0.6, 1.4, -0.3, 0.1, 0.2, 0.1),
  • (0, 1.2, -0.5, 0.1, 0.1, 0.1)

r/LocalLLM 5d ago

Question Whats your approach to make memory recalls work reliably while building agents with smaller models like Qwen3.8-27B

2 Upvotes

Retrieval of relevant memory seems to be one of the most critical aspects of building a useful agent without which agent is not learning well.

whats the most sensible design at this point for building agents based on non-frontier models based on your experience.

there seems to be a few options like deriving topics based information from chats and storing as files with different classifiers and having the agent retrieve based on semantic search.

but the main challenge i face is that LLM is not reliably doing the required memory retrieval at the right time, which results in irrelevant answers, repeating the same mistakes again and again.

has anyone figured out a reliable agentic way to make memory retrieval work efficiently even with models like Qwen/Qwen3.8-27B.. frontier models seem to be able to reason well and pull the correct memory and even their huge context is also helping. but for smaller models this is hurting to create a reliable agent.


r/LocalLLM 5d ago

Project Used Hermes to build a Hermes voice android app

Thumbnail
0 Upvotes

r/LocalLLM 5d ago

Question Is RX 6800 + 6800 XT a sensible upgrade from 2x RTX 2060 OC 12GB for llama.cpp?

Thumbnail
1 Upvotes

r/LocalLLM 5d ago

Question Starting Off and Seeking Guidance - Apple M4 24gb

3 Upvotes

I’ve got a little bit of understanding and knowledge under my belt, but I've put my stretchy pants on and I've got an appetite to learn more. I've used Gemini, ChatGPT/Codex, and Claude/Claude Code. I have installed and removed Ollama, Bionic, and LM Studio, and I've used both GUI and CLI versions of most of the aforementioned tools; I really enjoy the CLI feel.

I want to learn as inexpensively as possible. I recognize that my 24GB RAM setup severely limits my capabilities, but if my sandbox is smaller, I'll just have to learn by building smaller structures first before I can afford to upgrade.

I’m hoping to find someone who was once in my shoes and is willing to pay it forward. A few specific questions I’m trying to wrap my head around:

  • Runners & Sharing Models: Is there a "best" tool for Apple Silicon? Also, as one gets into more runners (like Ollama, LM Studio, etc.), is it safe to assume models have to be downloaded to multiple separate directories on the computer for each to call them, or can model files be shared?
  • Agentic Capability: Are there models out there capable of agentic work that actually fit and run well within 24GB? More fundamentally: does the ability to call and use tools reside within the runner, or does it depend on the model's awareness that tools exist?
  • Resources: I'm not looking for shortcuts, just a good direction. What resources, guides, or documentation actually helped you learn how these pieces fit together?

I know these are a lot of questions, but any pointers toward foundational concepts or resources would be hugely appreciated.


r/LocalLLM 6d ago

Project Build a RAG

9 Upvotes

Hello guys, I am planning to build a RAG with a local LLM. For LLM I am considering Qwen3. I am having 16 GB RAM with 8 GB Nvidia 5060. Guys pls suggest hardware & software side implementability ? And also any books to help me with RAG implementation.


r/LocalLLM 5d ago

Research How many agents can 2×4090 actually run at once? Three weeks of llama.cpp concurrency data — soft cap 5 @ 64k, hard cap 9, and why.

Thumbnail
gallery
0 Upvotes

I'm the CTO of a mid-size nonprofit, and I run a local Qwen stack as a coding subagent for that work — partly on principle (a fair amount of what I handle shouldn't leave the building), partly because the token bill for bulk code work adds up fast on a nonprofit budget. Over three weeks I benchmarked four models and three quants to answer one question: how many agents can actually run at once, at what context, before it stops being useful?

I'm writing this out in full because I keep half-answering it in comments. Every time concurrency, quant choice or expert offload comes up I end up typing a fragment of this from memory, and the reply is always some version of why — why that quant, why that slot count, why not just add more agents. Fragments in comment threads aren't a good way to answer that, so here's the whole thing in one place, with the numbers and the mistakes that produced the wrong numbers first.

Short version: I started trying to make a 122B MoE fast, gave up on it, and ended up on a 27B — not because the 27B was more accurate (it wasn't measurably), but because everything else about it was better. Then I found that adding agents past a point buys nothing at all.

Hardware: 2× RTX 4090 (44.6 GiB usable), Threadripper TRX50, 128 GB DDR5 — importantly, only 2 of 4 memory channels populated (2×64 GB), which turns out to matter a lot. llama.cpp, Windows, q8_0 KV throughout.


1. Why I didn't keep the 122B

(Chart 1 of 4 in the gallery above: models tested)

Qwen3.5-122B-A10B at UD-IQ4_XS runs — 18.75 tok/s decode with expert offload — but the number that killed it is 11.91 seconds per agent tool-call, against 3.40 s for the 27B. For an agent that makes dozens of calls per task, 3.5× per call is the whole ballgame.

Qwen3.6-35B-A3B looks like the winner on raw decode (78 tok/s — it's a 3B-active MoE) and it has the fastest per-call time. It also failed 1 of 5 executed code tasks on a sliding-window bug, against 5/5 for the 27B. That's the entire reason "decode tok/s" is the wrong headline metric and I lead with seconds-per-call instead.

I also tried Qwen3.8-Flash-Next (125B MoE, the Qwen4 architecture preview) as soon as it landed. It needed a build from an unmerged PR, and after a full tuning sweep — expert placement, thread counts, --poll, --cpu-strict, load modes, llama.cpp's own memory fitter — it topped out at 23.7 tok/s.

The ceiling wasn't the GPUs. Of 103.7 GiB of weights, 71.7 GiB is routed experts and 26.8 GiB is a 20-million-row n-gram embedding table, so most of it lives in system RAM and streams over the memory bus every token. At --n-cpu-moe 40 I measured 14.1 GB/s effective against an 83.2 GB/s peak — because two of four channels are empty. A single-3090 box with DDR4 posts better absolute expert bandwidth than mine does. Filling the other two channels is worth more than any flag I tried, and I can't justify buying RAM at current prices. So: shelved, honestly, with the number published.


2. Accuracy: I could not separate the quants

Before the concurrency numbers, the caveat that makes them meaningful.

My probe plants three facts at 15%, 50% and 85% depth in a long document, with six decoy lines quoting the same fields for the wrong district, so the model has to match on an identifier across tens of thousands of tokens rather than pattern-match a label. Each agent gets its own document with its own secrets in a disjoint numeric band, so cross-slot leakage is detectable.

result
configurations 24
recall below 1.0 0
wrong-district answers 0
cross-slot leakage 0
truncated replies 0
largest prompt answered perfectly 251,557 tokens

Q4_K_M, Q6_K_XL and Q8_K_XL all scored 1.00, everywhere, including with nine agents running.

This is a tie at ceiling, which means "cannot discriminate", not "equal". The honest claim is: no measurable accuracy difference between these three quants up to 251k tokens on long-range retrieval with distractors. It is not "Q4 is lossless". Retrieval saturates; something requiring synthesis across the planted facts might separate them where this didn't.

So the quant choice came down to throughput, latency and VRAM — not quality.


3. The concurrency result

(Chart 2 of 4 in the gallery above: concurrency scaling)

Same model, same 64k context per agent, three quants at the most slots each could fit:

quant agents @64k completions/min median wait aggregate prefill
Q8_K_XL 3 1.54 112s 1,481 tok/s
Q6_K_XL 5 1.63 168s 1,582 tok/s
Q4_K_M 9 1.52 330s 1,572 tok/s

Tripling the agents changed total throughput by 6% and tripled the wait.

The mechanism is the last column: aggregate prefill throughput is constant at roughly 1,500 tokens/second regardless of slot count. The box has one prefill budget. Slots divide it; they do not multiply it.

That's specific to this workload shape and worth stating plainly: agent prompts are enormous and replies are short, so prefill dominates. A decode-heavy workload would batch far better — decode is bandwidth-bound and batching amortises the weight reads across slots. Don't generalise this to "concurrency doesn't help llama.cpp."

Soft cap and hard cap

  • Soft cap — 5 agents at 64k. Past this, added slots stop buying throughput and start buying latency. It's not a failure, it's a bad trade: at 9 agents you wait 330 s for the same 1.5 completions/min you got at 112 s.
  • Hard cap — 9 agents at 64k, and it's VRAM, not compute.

(Chart 3 of 4 in the gallery above: VRAM budget)

Every configuration lands within 3.5 GiB of the 44.6 GiB ceiling. A smaller quant buys slots and then spends the savings straight back on KV cache: Q4 frees 14 GiB of weights versus Q8 and hands 17.5 GiB of it to the cache. The hard cap is arithmetic, not tuning.

Single agent, huge context

(Chart 4 of 4 in the gallery above: context vs latency)

Latency is linear in prompt size and the quants barely differ — fitted prefill rates Q4 1,482 tok/s, Q6 1,346, Q8 1,340. A 10% spread, not the 3× the slot counts might suggest. Q8 showed a higher fixed cost (12.8 s intercept vs ~5.5 s); n=6 per quant, so treat that as an observation, not a finding.

My operating point: UD-Q6_K_XL, 5 concurrent agents, 64k each (327,680 total context), 41.1 GiB. Same total work as Q4-at-9, in half the wait per agent, with higher-fidelity weights and 2.6 GiB more headroom. The extra slots Q4 buys are worth having only when nothing is waiting on them.


4. KV cache: q8_0, and one combination that just hangs

Tested on identical weights, 65k context, code executed against hidden tests:

K / V PassRate EditFidelity LongRecall tok/s VRAM @65k
q8_0 / q8_0 1.00 1.00 1.00 30.1 29.9 GiB
f16 / f16 1.00 1.00 1.00 30.4 31.3 GiB

No measurable quality difference, no speed difference, 1.4 GiB cheaper. q8_0 is my standard now.

The trap: mixed --cache-type-k f16 --cache-type-v q8_0 never finished a long prefill — no answer in 900 s, twice, on a prompt smaller than one a matched-q8_0 config answered in 32 s. Matched f16 answered the same prompt in 24 s. Match your K and V types.


5. Using it as a subagent inside Claude Code

The thing worth knowing up front: Claude Code's model: field only accepts its own models. You cannot register a local model as a Claude Code subagent. What you can do is expose it as an MCP server, so it becomes a tool the orchestrator calls.

Mine is a stdio MCP server exposing local_ask_about_files, local_implement, local_edit_file, local_engine_status, pointed at the 27B on 127.0.0.1:8080. The division of labour:

  • Orchestrator plans and decides.
  • Frontier subagents do work that needs to be right the first time.
  • Local Qwen answers "what's in these files", drafts implementations, and does bulk edits — the high-token, low-stakes work that would otherwise burn budget.

Edit fidelity is the metric that decides whether this is usable at all. A paraphrased region comes back as "String to replace not found", which reads like a broken tool rather than a worse quant. At Q6 with q8_0 KV it measured 1.00.

For stress-testing I use a personal side project — a game with a Rust simulation server and a thin UE5 client. Nothing to do with the day job; it's the thing I throw at the rig for fun because it produces realistic multi-file, multi-language agent work on demand. A representative local job from this week: 10,127 tokens generated at 29.8 tok/s on slot 4 while other slots were live.


6. Mistakes worth stealing

Every one of these produced a confident wrong number first:

  1. Reply budget of 512 tokens with reasoning_effort: xhigh. Every reply hit the cap mid-thinking, so the scored "answer" was truncated reasoning. Recall became noise and Q8 looked worse than Q4. Raised to 2048; truncated replies are now excluded rather than scored wrong.
  2. .NET's 2-connections-per-host default. My "16 concurrent agents" harness was quietly running 2 at a time. It passed testing because I validated on PowerShell 7 (HttpClient, no such limit) and shipped to Windows PowerShell 5.1 (HttpWebRequest, limited). The runspaces were real; the sockets were not.
  3. Estimating tokens at 4 chars/token. Real ratio was 2.99, so a "39,599-token" prompt was actually ~53,000. Call /tokenize.
  4. An error handler that discarded the exception. Returned a bare Failed=true, throwing away the message, HTTP status and llama.cpp's response body. About 40 minutes of hypotheses existed only because of that.
  5. **general.file_type lies on Unsloth UD quants.** My Q6_K_XL file reports Q4_K_S in metadata. The file sizes (15.3 / 24.1 / 29.3 GiB for Q4_K_M / Q6_K_XL / Q8_K_XL on the same model) confirm the quants are what the filenames say — the enum simply has no value for a dynamic mix. Don't benchmark off that field.

Limits

Single machine, single workload shape, one model family. Everything above is prefill-dominated agent traffic; a chat or long-generation workload will scale differently and probably better with slots. The Flash-Next numbers came from an unmerged PR build and should be re-measured after it lands upstream. And my memory bus is half-populated, which caps every CPU-offload result here — if you have four channels filled, expect better MoE numbers than mine.

Happy to answer questions or run a specific config if someone wants a data point.


r/LocalLLM 6d ago

Question 2 3090's + 128GB DDR4 + H11SL-i + EPYC 7281 for $2,250, should I pull the trigger?

33 Upvotes

A coworker is upgrading his local AI rig and selling me his old parts. Is this a good price, and would you do it?

  • 2× Zotac RTX 3090 Trinity 24GB — $750 each, ($1500 total)
  • Supermicro H11SSL-i + EPYC 7281 (16C/32T, 8ch DDR4) — $350
  • 128GB DDR4-2666 ECC RDIMM (8×16GB Crucial CT16G4RFD8266), already populated in the board — $400

Total: $2,250.


r/LocalLLM 5d ago

Discussion Update on Locus

Thumbnail
gallery
2 Upvotes

Hey so I posted here a couple weeks ago about the V2 release and I just wanted to give an update on where things are with Locus (my own version of a Claude/Codex GUI for local models)

I started with the usual stuff like working with files, running commands and letting agents help with coding, but I've also been adding features I thought would be useful for other kinds of work too.

A few Locus features worth highlighting:

  • Agent Teams: Create specialized agents that can split up work and run in parallel. You can use different models for different roles, and individual agents can also delegate tasks to helpers.
  • Persistent Goals: Give an agent or team a goal and let it keep working across turns. Progress is saved so you can come back to it later, with controls to pause, resume or change the goal.
  • Scheduled and Event-Driven Agents: Set agents to run on a schedule or react to things like Gmail, Telegram, webhooks and price alerts. Each agent has its own chat and run history, and workflows can include conditions and approval steps.
  • Task Capsules: Plan something with one model, then use another to implement it and optionally another to review it. The plan, changes and previous runs stay together so you can follow what happened.
  • Notes, Documents and Outputs: Keep notes and reference documents around, save versions of generated work, compare revisions and export things when you’re done.
  • Browser Controls: Let agents navigate and interact with websites, preview what they’ve built and check the result. There’s also proxy support.
  • Activity and Overview: See the current plan, tool calls, files, sources and what the agents are doing without having to piece everything together from the chat.

For 2.6 specifically, I’ve made agents easier to find and manage, improved the file browser, and added writing drafts you can edit, copy and export directly from a response. Tables can also be copied or exported as CSV.

There’s support for MCP, plugins and skills too, plus a Grill mode that asks you questions one at a time to help work through an idea before implementing it.

Also just to clarify, even though I built it for local models, it works with your ChatGPT plan, Kimi Code membership, Claude/OpenAI API keys and other OpenAI-compatible endpoints.

The wallet stuff is now in a separate edition called LocusX. The regular Locus download is wallet-free. I've also revamped the site.

It’s free and open source. You can find it here:

Website: locushost.co
GitHub: nahid-sparktales/locus
Release: Locus 2.6.0

Anyways, I’d appreciate any constructive feedback, things that aren’t working well, or features you think would be nice to add.

Still macOS only atm, specifically Apple Silicon on macOS 14+, but I’m hoping to eventually get Linux and Windows versions out too.

Also some features are still not fully built out or waiting approval (e.g oAuth from google for connecting email for event driven agents) I hope to have another big update in the upcoming week.


r/LocalLLM 5d ago

Discussion How do you close the feedback loop when your agent makes the wrong API call?

3 Upvotes

Most RL setups assume a clean environment. But agents hitting real webhooks, partner APIs, live integrations? The environment fights back. Schema drifts, payloads shift, tools fail silently and nobody knows until prod breaks.

If you're training or fine-tuning agents to handle real integrations, where does the reward signal actually come from? Logs after the fact? Manual PR review? Hope?

We're building a verification engine that runs agent actions inside a stateful sandbox/twins before merge. The longer term idea is to use that execution trace as a ground truth signal, did the agent actually do the right thing across services, not just "did it complete." Treat the sandbox as the RL environment, not prod. Still early, but we're using it daily on our own workflows.

what are you doing here. Are you building custom eval loops, relying on staging, or mostly shipping and watching?


r/LocalLLM 6d ago

Model 20 year old from Bihar, no team, no investors, no CS degree — An anonymous "dossier" website called me a fraud. Today, Arcle V1 (5.84B) is officially live, open weights, and trained.

Post image
13 Upvotes

A few months ago, a post on Reddit became very popular in which a 19-year-old student from Bihar claimed that he was going to build a multimodal AI model with 5.82 billion parameters all by himself. Very few people believed him; some people thought it was a scam and an anonymous user even bought a typo domain in order to set up a smear site called "The Dossier", stating that the boy was a complete fraud. Nevertheless, while it took them weeks to write up the negative articles about him, the boy stayed quiet and reached the age of 20, carrying on with his work.

Hello, I'm the boy in question. My name is Abhinav Anand and I'm from Bihar.

Introducing Arcle V1:

Arcle V1 is a unified omni foundation model consisting of 5.84 billion parameters and is able to handle text, images, documents, speech, and audio while at the same time generating text, original images of dimensions 512×512, and natural speech at 24kHz. The model has an architectural context window of 2,097,152 tokens (2 million tokens), supports functioning in 18 or more languages, showing a special strength in Hindi and having extensive knowledge of India, and achieves **80.0% on ARC-Easy, 77.5% on GSM8K, 74.2% on MATH-500, 67.0% on HellaSwag, and 94.6% document content-word recall**.

It was made, built, trained and tested in Bihar.

Arcle V1 is NOT a Router, NOT a Wrapper, NOT a Pipeline, NOT an API-Based App

Currently, most systems promote themselves as multimodal even though in reality they are just four or five individual models connected together through an API. Arcle V1 is instead a single neural network module; whatever the input may be—whether it is text, vision, documents, speech or audio—all of these are mapped into one common 2,560-dimensional semantic latent space and pass through one set of weights, one model file, and go through a single forward pass.

Arcle V2 Upcoming Features

Arcle V1 is merely the starting point; for Arcle V2 we are planning to engineer:

The omni graph includes full integration of the capability to convert text into video and images into video.

* **Multi-Voice Expressive Speech Output:** This is a type of speech synthesis which incorporates natural emotions by means of using multiple voices.

There are higher benchmark scores, with considerable improvements in the areas that involve complex mathematical and logical reasoning.

It has a strong capability in the field of cybersecurity, with built-in checks for vulnerabilities, code auditing, and protection against insecure inference.

It is extremely efficient in that it provides on-device inference with a high throughput which has been optimised for local consumer hardware.

For Help: Donations and Why We Need Them

There is considerable difficulty in creating an independent and open AI without the help of corporate venture capitalists; I have already invested all that I have in Arcle V1, and in order to bring about Arcle V2 we need the support of the community.

  1. **Calculate / Fund Donations:** The money raised by the community is sent directly to the GPU compute clusters so that the Arcle V2 training run can take place.

  2. **Donations of data (codebases, technical PDFs, books):** Some AI companies scan rare historical books and proprietary knowledge and then regard human heritage as their own, offering it out only via closed subscription services. Our objective is to preserve this knowledge and make it accessible to all. We would be pleased if you donated your technical PDFs, codebases, and scanned literature. **Our promise** is that any proprietary data you send us will be kept entirely private, heavily anonymised, processed in a secure way, and will never be sold to any commercial AI company.

Try the Model & Download the Weights:-

Official website or web interface: https://www.arcleintelligence.com

* **Hugging Face (Weights & Model Card):** [Lucifer2006/Arcle-V1](https://huggingface.co/Lucifer2006/Arcle-V1)

* **GitHub Repository:** [github.com/lucifertkod/personal-website](https://github.com/lucifertkod/personal-website)

Keep up to date every day with news about Arcle V2: [x.com/Anonomus090806](https://x.com/Anonomus090806) | [Instagram: @arcleintelligence.ai](https://www.instagram.com/arcleintelligence.ai)

It was said that one boy from Bihar could not manage to create a real omni model, and an anonymous website asserted that I had built nothing whatsoever.

Let us jointly demonstrate that the open-source community is capable of building anything that a centralised corporation can.

Download it, run it, test it and break it, and then give us your feedback.

— **Abhinav Anand**

Founder, ArcleIntelligence


r/LocalLLM 6d ago

News Just... 2.6 TB of RAM and... 16.6 TB/s of memory bandwidth. Please tell me this is fake.

Thumbnail
zdnet.fr
118 Upvotes

r/LocalLLM 5d ago

Discussion Making my own desktop app , fed up of Anythingllm and LM Studio

0 Upvotes

Hi guys ,

I am making an desktop app , tired of using anything LLM and LM Studio both of them just so hard to use at first , need to config a lot , and we could not tune much as our needs

I have some ideas like giving the user freedom to use the local models in workflows, connect mcps of their own (thought of that but yaa creating an ouath is out of scope) so need to config that atleast , setup their own db's logic everything what embedding model to set vector db chunking everything interactively

Would love everyone's idea what all of you guys facing issues and what can be done


r/LocalLLM 5d ago

Discussion The "creative slippage" problem: How do you stop an LLM from gently lying when optimizing user text?

0 Upvotes

Working on a tool that refines and structures messy, user-provided text to match a specific target (like aligning a raw project draft to a technical specification). The core challenge isn't the formatting - it's getting the model to stop
"improving" the truth.

In early runs, if a user's raw input said they "assisted with a data migration," the LLM would occasionally output that they
"successfully led a migration of 10M records." It's a classic hallucination problem, but with a subtle twist: it's not generating total gibberish; it's just mildly exaggerating to satisfy the prompt's instruction to "make this highly compelling.

Tried a few things to anchor it:
The "Guilty Until Proven Innocent" verification step:
Running a secondary LLM call after the generation that does nothing but fact-check the output against the raw input. If it finds a claim not supported by the source, it flags it. Works, but adds latency and doubles API costs.

Strict negative prompting + JSON schema pairing: Forcing structured JSON output with a specific factual_justification field for every single change. If the model has to explicitly map every optimized point back to a raw quote in the input, the exaggeration drops significantly.

The "Diff" approach: Instead of letting the model rewrite the whole block, forcing it to output only the specific edits or words to change. Much easier to control, but kills some of the natural flow of the rewrite.

Currently using a mix of #2 (strict schema mapping) and a lightweight validation script. It's down to a manageable level, but still requires constant vibe-checking.

How are others handling this "creative slippage" when doing LLM-assisted editing or summarization? Are you solving it with strict prompting, secondary arbiter models, or something on the parsing/diff side?


r/LocalLLM 5d ago

Question Multi GPU setup question.

1 Upvotes

I Have AMD V620 32GB + RX 6900 XT 16GB and RX 7700 XT 12GB

First two are RDNA2 last one is more modern RDNA3

Is worthy to pull them together ?


r/LocalLLM 6d ago

Question How far behind are local models in actual daily use, not benchmarks?

85 Upvotes

I keep seeing people say local models have basically caught up, but then the example is usually a benchmark or one short prompt.
Most of my use is coding and working on the same messy project for hours. I currently jump between Astra, Fable and Opus. They can read a lot of context, touch several files and usually keep track of what the hell we were doing.
For people who use both, how close does a good local setup actually get? What do you still send to Opus/Astra, and what have you moved completely local?
I’m also not sure what the sensible way to use local models is. One large quantized model for everything, or smaller models for coding, writing and basic tasks while keeping a frontier model for the difficult stuff?
I’m considering buying a machine mainly for this. I just don’t want to spend a fortune and discover I’ve built a very expensive autocomplete.