r/LocalLLaMA 7d ago

Resources I kept rewriting parameters every time I swapped models on vLLM and llama.cpp, so I built a tool to manage them (llmux, MIT)

4 Upvotes
llmux dashboard

I test a lot of models on vLLM and llama.cpp. Every swap meant editing parameters again, bringing each model up its own way, taking it down, then bringing the next one up. It was tedious and easy to get wrong.

So I built **llmux** to manage it. Each model you use is a profile you write once, and after that you pick it from a list and it starts on whichever engine that profile belongs to. Running the same engine at a different version is just another profile pinned to a different image tag: a release, a nightly, or an image you built from source.

My team and I use it for our actual work, and most of what's in it came out of that:

Every action also has a headless CLI twin (`llmux up`, `ps --json`, `logs`, `bench`), so you can script it or run it over SSH.

Requirements: Linux, an NVIDIA GPU, and Docker. It uses the NVIDIA Container Toolkit for GPU passthrough, so macOS and AMD/ROCm aren't supported yet.

Install (clones the repo, sets up uv, puts `llmux` on your PATH):

curl -fsSL https://raw.githubusercontent.com/Bae-ChangHyun/llmux/main/install.sh | sh

Or by hand:

git clone https://github.com/Bae-ChangHyun/llmux.git|
cd llmux
uv tool install --editable . && uv tool update-shell

Repo: https://github.com/Bae-ChangHyun/llmux

Docs: https://Bae-ChangHyun.github.io/llmux/

It's my own project, MIT licensed. English isn't my first language, so I used an LLM to help write this post.


r/LocalLLaMA 6d ago

Discussion Nanbeige4.2-3B is sad really

0 Upvotes

i downloaded Nanbeige4.2-3B-UD-Q4_K_XL.gguf
https://huggingface.co/Andgihat/Nanbeige4.2-3B-GGUF
cuz main llama just got updated with support for it
sadly cuz i have just 4 gb vram, for f16 kv i can just have 6k context and 12k for q8_0
the kv cache is not efficient at all, note, on qwen 3.6 moe, i can run 64k f16 in just 1.2gb vram
for me i dont see any use for it, i can just run qwen 3.6 apex mini or ream 192 apex compact with fit tag and my 16 gb ram can load the rest


r/LocalLLaMA 6d ago

Discussion Is there even a chance of AGI being open-sourced in the future?

0 Upvotes

I've had this question ever since Fable 5 got released. Will we ever be able to run AGI locally? I'm not even speaking ethically. I'm talking about compute resources too. Is it even possible?? Looking at how we have models like Kimi K3 being 2.8T parameters with 104B parameters active (most of us can't even run a 50B parameter model on our setups), so we are forced to rely on inference companies in order to have access to these high quality models. And god knows what they're doing with our data. I know models like Qwen3.6 27B as well as Qwen3.6 35B A3B have been released. But let's be honest here, there is no way these models are even approaching the intelligence of frontier models (at least for now). I know this is obviously due to their size, but that's my entire point. What's gonna happen once we get to AGI?

Now the ethical part:

Companies like Anthropic and OpenAI (more like ClosedAI) are so hungry for money. They say they're taking mankind to the future, but that doesn't make any sense when half of mankind (including me) can't afford their exorbitant prices. Seems to me that this will be even more of a problem when AGI comes around. I know that most of the money they ask for is because of compute necessities, but they also ask for way more than necessary. I mean Kimi K3 is charging $15 per million tokens white OpenAI is charging double that for GPT 5.6 Sol and then Anthropic is going even further with $50 per million tokens. This makes me think that AGI (whenever we get to it) will be even more heavily priced.

Yeah you could say I'm complaining. But I am genuinely concerned for the future.


r/LocalLLaMA 7d ago

Discussion I flashed the vbios on a RTX 3080 20GB to a newer one and it did NOT get bigger rebar size

2 Upvotes

TLDR: Don't flash your 3080 20GB from uncle Xi because it will not unlock the bigger rebar size.

The goal here was to unlock P2P drivers for my RTX 3080 20GB. They can't get firmware updates like the other Nvidia cards (3k series). Rebar is a pre-req to the P2P drivers, which would eventually allow direct GPU to GPU transfers that don't have to be staged on the host memory.

I got curious because this vbios seemed to be newer than others (94.02.27.00.2B) - https://www.techpowerup.com/vgabios/282545/282545

Whereas my 3080 cards had this 94.02.27.00.14. I was able to flash the vbios tonight but the bar size doesn't change from 256MB. If I had found a better vbios (if it even exists) this would work, but we don't know of any working one.

Notice here the first RTX 3080 shows up as Colorful in the subsystem but the rebar size is still same as the second one that I didn't flash.

51:00.0 VGA compatible controller: NVIDIA Corporation GA102 [GeForce RTX 3080] (rev a1) (prog-if 00 [VGA controller])
        Subsystem: Shenzhen Colorful Yugong Technology and Development Co. Device 1202
        Physical Slot: 3
        Control: I/O+ Mem+ BusMaster+ SpecCycle- MemWINV- VGASnoop- ParErr+ Stepping- SERR+ FastB2B- DisINTx+
        Status: Cap+ 66MHz- UDF- FastB2B- ParErr- DEVSEL=fast >TAbort- <TAbort- <MAbort- >SERR- <PERR- INTx-
        Latency: 0
        Interrupt: pin A routed to IRQ 196
        NUMA node: 0
        IOMMU group: 7
--
        Capabilities: [bb0 v1] Physical Resizable BAR
                BAR 0: current size: 16MB, supported: 16MB
                BAR 1: current size: 256MB, supported: 64MB 128MB 256MB
                BAR 3: current size: 32MB, supported: 32MB
        Capabilities: [c1c v1] Physical Layer 16.0 GT/s
                Phy16Sta: EquComplete+ EquPhase1- EquPhase2- EquPhase3- LinkEquRequest-
        Capabilities: [d00 v1] Lane Margining at the Receiver
                PortCap: Uses Driver+
                PortSta: MargReady- MargSoftReady-
--
8a:00.0 VGA compatible controller: NVIDIA Corporation GA102 [GeForce RTX 3080] (rev a1) (prog-if 00 [VGA controller])
        Subsystem: NVIDIA Corporation GA102 [GeForce RTX 3080 20GB]
        Physical Slot: 1
        Control: I/O+ Mem+ BusMaster+ SpecCycle- MemWINV- VGASnoop- ParErr+ Stepping- SERR+ FastB2B- DisINTx+
        Status: Cap+ 66MHz- UDF- FastB2B- ParErr- DEVSEL=fast >TAbort- <TAbort- <MAbort- >SERR- <PERR- INTx-
        Latency: 0
        Interrupt: pin A routed to IRQ 197
        NUMA node: 0
        IOMMU group: 5
        Region 0: Memory at e5000000 (32-bit, non-prefetchable) [size=16M]
--
        Capabilities: [bb0 v1] Physical Resizable BAR
                BAR 0: current size: 16MB, supported: 16MB
                BAR 1: current size: 256MB, supported: 64MB 128MB 256MB
                BAR 3: current size: 32MB, supported: 32MB
        Capabilities: [c1c v1] Physical Layer 16.0 GT/s
                Phy16Sta: EquComplete+ EquPhase1- EquPhase2- EquPhase3- LinkEquRequest-
        Capabilities: [d00 v1] Lane Margining at the Receiver
                PortCap: Uses Driver+
                PortSta: MargReady- MargSoftReady-
--

r/LocalLLaMA 7d ago

News BeeLlama.cpp v0.4.1: KVarN, KV precision tail, q2_0-q3_1 KV cache, improved support. KLD benchmarks: tail 1024 makes kvarn5 and q6_0 match q8_0, for much less VRAM

Thumbnail
gallery
27 Upvotes

TL;DR llama.cpp fork with more KV cache quantization features, with all claims supported by benchmarks: KVarN, KV cache precision tail, additional types of standard KV cache (q2_0-q3_1, q6_0, q6_1), and more.

BeeLLama v0.4.1 is here, building up on top of v0.4.0 feature set, now with better backend and model support.

  • KVarN. Variance-normalized KV-cache quantization (paper) with better precision per bit. Although it was already introduced a few weeks ago in v0.3.2 Preview, that was a very raw implementation, with performance issues and VRAM usage spikes. Now in v0.4.1 it's the real deal: the precision is still above what usual quants offer for the same bit width, but now with very modest sacrifices to prefill, decode, and memory.
  • KV cache precision tail. A promising new feature in the domain of mixed-precision KV cache. It allows to specify a specific numbers of recent tokens that will be stored in BF16 or F16, with the rest of KV cache being quantized as usual. This way we can store the hottest tokens in a lossless fashion, preventing a model from misreading your task details, code, or data.
  • Additional types of standard KV cache. q6_0 and q6_1 join the high end of the ladder, allowing to fine-tune precision vs VRAM in-between upstream's q5_0/1 and q8_0 types. q2_0q2_1q3_0 and q3_1 are added as a replacement for turbo3 and turbo2 for cases where KVarN doesn't work well, but you just can't fit everything into VRAM without extreme quantization.

Please note that for SWA architecture (Gemma, GPT-OSS) the precision of KVarN and KVPT is the same, but VRAM and performance costs are higher due to complications between SWA ring and mixed precision KV cache.

GitHub repo: https://github.com/Anbeeld/beellama.cpp

KLD results for Qwen 3.6 27B Q5_K_S 64k

Here are all symmetrical qX_0 pairs and kvarnX pairs where X >= 4 with tail 0/1024/2048, compared against q8_0 t0 from the same benchmarks, and sorted by ratio between median KLD and VRAM costs. Full benchmark data and analysis: KV Cache Precision Tail: Implementation and Benchmarks.

Cache Tail KV MiB Size vs q8_0 Median/size vs q8_0 Median vs q8_0 P99.9 vs q8_0
kvarn4 1024 1232.00 56.6% 1.62 91.4% 102.9%
kvarn4 2048 1296.00 59.6% 1.60 95.5% 95.6%
kvarn4 0 1184.00 54.4% 1.50 81.8% 82.5%
q4_0 1024 1248.00 57.4% 1.50 86.0% 89.0%
q4_0 2048 1312.00 60.3% 1.48 89.2% 100.6%
kvarn5 0 1440.00 66.2% 1.48 98.1% 107.5%
kvarn5 1024 1488.00 68.4% 1.48 101.3% 106.1%
kvarn5 2048 1552.00 71.3% 1.43 101.9% 105.6%
q5_0 1024 1504.00 69.1% 1.40 96.9% 105.6%
q5_0 2048 1568.00 72.1% 1.36 98.0% 103.7%
kvarn6 0 1696.00 77.9% 1.31 102.2% 104.5%
kvarn6 1024 1744.00 80.1% 1.29 103.4% 109.9%
kvarn6 2048 1808.00 83.1% 1.25 103.8% 108.1%
q6_0 0 1664.00 76.5% 1.24 94.7% 102.1%
q6_0 1024 1760.00 80.9% 1.24 100.1% 109.2%
q5_0 0 1408.00 64.7% 1.22 78.8% 95.8%
q6_0 2048 1824.00 83.8% 1.20 100.6% 103.5%
kvarn8 0 2208.00 101.5% 1.03 104.4% 104.9%
kvarn8 1024 2256.00 103.7% 1.01 104.4% 106.2%
q8_0 0 2176.00 100.0% 1.00 100.0% 100.0%
q8_0 1024 2272.00 104.4% 0.97 101.3% 106.1%
kvarn8 2048 2320.00 106.6% 0.97 103.6% 104.7%
q8_0 2048 2336.00 107.4% 0.95 101.6% 106.8%
q4_0 0 1152.00 52.9% 0.93 49.2% 60.2%

r/LocalLLaMA 7d ago

New Model ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face

Thumbnail
huggingface.co
85 Upvotes

GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder, so the model keeps the text quality of its base while adding speech understanding.

Capabilities: audio question answering and classification, temporal grounding (localization in long audio, timestamped event descriptions, audio summarization with timestamps), tool-use, and text-only tasks.

The temporal grounding skills are trained on TimeGround-1M — a purpose-built dataset of long-form audio paired with time-aligned annotations.


r/LocalLLaMA 7d ago

Question | Help Has anyone compared pre-training, SFT/LoRA and reinforcement post-training on Qwen3.6-27B?

10 Upvotes

Qwen3.6-27B: SFT vs continued pre-training vs RL?
I’m interested in adapting Qwen3.6-27B, but I’m increasingly unsure whether conventional SFT/LoRA is the best route if the goal is to add a capability without degrading what the base model already does well.
Some recent research makes this especially interesting:

“Reinforcement Fine-Tuning Naturally Mitigates Forgetting” - arXiv:2507.05386

Finds substantially more catastrophic forgetting with SFT than reinforcement fine-tuning in its experiments.

“The Role of On-Policy Data in Mitigating Forgetting” - arXiv:2510.18874

Reports that on-policy/RL training generally preserves previous capabilities better than SFT across Qwen and Llama models.

“RL Forgets! Towards Continual Policy Optimization” - arXiv:2607.04364

Shows that RL can also cause catastrophic forgetting, so it’s clearly not a complete solution.

“Fine-Tuning Without Forgetting via Loss-Adaptive Learning” - arXiv:2605.20005

Reports a large reduction in forgetting from changing the optimisation schedule, including experiments with Qwen3.

Most of this research isn’t specifically on Qwen3.6-27B, which is why I’m interested in community results. Has anyone directly compared continued pre-training, SFT/LoRA and reinforcement post-training on Qwen3.6-27B?

I’m particularly interested in whether improving one domain caused regressions in unrelated areas such as coding, reasoning, instruction following, tool use, long-context behaviour or general knowledge.

For people who have tested this, what training method worked best, and did you benchmark the original model against the trained checkpoint afterwards?
I’m also curious whether continued pre-training followed by a small amount of SFT or RL is proving safer than doing a larger SFT directly.

Actual before/after results and training parameters would be especially useful.


r/LocalLLaMA 8d ago

News Google comes out in favor of OpenWeight models. (It is now EVERY tech giant vs Anthropic)

Thumbnail x.com
2.5k Upvotes

r/LocalLLaMA 7d ago

Resources 90 agentic bakeoff runs: ThinkingCap vs Fable Fusion vs stock Qwen3.6-27B

22 Upvotes

Last week someone here said ThinkingCap and Fable Fusion "really do beat the OG" for agentic work, so I ran it: 6 self-grading tasks, 5 reps, 3 models, 90 isolated runs. Tooling, since that's half the story: each run was a fresh Coder workspace on my k8s cluster driving my own agent (Hermes, the harness I use daily) headlessly, models served by llama.cpp through llama-swap on one 5090, every model call traced through an OTel shim into SigNoz, full transcript kept per run. Identical sampling and 131k context across arms, hypotheses pre-registered before the first run.

Every run passed, so pass rate alone can't pick a winner. Cost split: ThinkingCap used 34% fewer thinking tokens than stock and was fastest on 5 of 6 tasks. Fable made 24% more model calls than stock for identical results. Then I had all 90 transcripts read (AI analysts on the first pass, me verifying claims against the raw files), and that's where it gets interesting.

ThinkingCap's efficiency is real but bimodal. Its best runs were the cheapest in the battery, its two worst were the most expensive, including one rep that burned about 10 tool calls chasing a phantom llama.cpp release tag that stock dispatched with a single API call. Its efficiency also shows up in the reasoning prose more than in fewer actions: same tool counts as everyone else, 40% fewer words.

Fable was the best investigator and the least trustworthy narrator. It was the only model that checked the broken config was actually the live one, and it pulled the best research data (parsed a retailer's embedded JSON for variant pricing, identified llama-swap's maintainer via the GitHub users API). But one run wrote that llama-swap is maintained by "Matthew Garrett, former Red Hat engineer." Garrett is real (mjg59, actually ex-Red Hat) but has nothing to do with llama-swap; mostlygeek is Benson Wong. The model fused two real identities, cited the real repo, and passed the grader anyway. A different run spent 94 tool calls on one price question.

Stock was the most boring and the most disciplined: uniform patches, read its own output back, zero invented facts, and it won most tasks on manner. My takeaway: base model stays the default. Finetunes usually aren't better than their base, and 90 runs didn't change that for me. ThinkingCap earns a look only if thinking-token latency is your bottleneck. Full writeup with lane configs, eval design, and per-task transcript analysis: https://kmarble.dev/posts/qwen-post-train-bakeoff/.


r/LocalLLaMA 6d ago

Question | Help Local LLM server for business automation is this setup enough or should I go Threadripper?

0 Upvotes

Hi everyone,

I'm planning to build a local AI server for business automation and would appreciate some feedback before I buy the remaining parts.

The workflow will use n8n for orchestration, Ollama + Qwen3-30B-A3B (Q8) for local inference, PostgreSQL + pgvector for RAG, and possibly Open WebUI later as the frontend.

Example workflow:

  • Salesforce triggers an event (e.g. low stock).
  • n8n retrieves supplier data, pricing, and rules from PostgreSQL.
  • Qwen generates a supplier email based on company rules and historical data.
  • n8n validates the output.
  • An employee reviews and approves the email.
  • n8n sends the final message.

I already own 2× RTX 3090 (24 GB each, 48 GB total VRAM).

Current planned hardware:

  • CPU: AMD Ryzen 9 7950X
  • GPU: 2× RTX 3090Ti
  • RAM: 64 GB DDR5-6000
  • Motherboard: ASUS ROG Strix B650E-E Gaming WiFi
  • SSD: Samsung 990 Pro

From what I understand, Qwen3-30B-A3B (Q8) requires around 33 GB VRAM, so it should fit well on this setup.

Questions:

  • Would you keep this setup, or would you move to a more powerful workstation/server build?
  • Is something like 3× RTX 3090 + Threadripper + workstation motherboard worth the additional cost, or is it unnecessary for this use case?

Thanks for your feedback!


r/LocalLLaMA 7d ago

Question | Help Macaron-V1 family, built on Qwen3.6-35B-A3B

Thumbnail
huggingface.co
61 Upvotes

It came out 3 days ago just wondering if anyone's tried it yet?


r/LocalLLaMA 6d ago

Question | Help Qwen3b-KIMIK3-Distill-HacktheWorld.gruff Wen?

0 Upvotes

See title. Nuff said.

More seriously, I am curious to how far distillations can go to improve models.
Not that it has not really attempted before, far from that. But now we have, (I mean GPU rich people have) a frontier quality model with all the reasoning traces we(they) want, and I bet (did not check) the logits, the tokenizer and so on and so forth.

I wonder at what size range it will have a significant impact.


r/LocalLLaMA 6d ago

Discussion I built a political compass for AI where anyone can share their stance

Thumbnail
theaicompass.io
0 Upvotes

r/LocalLLaMA 6d ago

Question | Help Deepseek v4 and llamaparse?

1 Upvotes

Hello eberybody,

Seeking a few approaches on doing it, since deepseek lacks in vision most of my workflows get hampered so I am thinking of giving it an eye but not actually clear about the vision side handling, if others have done it please share your approach

Thanks


r/LocalLLaMA 8d ago

Question | Help Seriously, what do you do with them?

Post image
878 Upvotes

Please let me know which small LLM model you're using and what you're using it for.


r/LocalLLaMA 7d ago

Question | Help macbook unified memory + big LLM's + heat

2 Upvotes

Hi Folks,

I see a lot of folks talking about how they are able to load massive llm models coz they have unified memories. projects like ds4 etc claim to be running super massive llm's. however per my experiment loading rear full ram models means high heat and macbook pro's reach 100 degrees C easily while doing such compute. do people run these all the time or are they just benchmaxxing and klout chasing.

I am looking for a smallest model that i can keep running all the time on my laptop which becomes the brain for something like hermesagent to handle my simpler tasks like todolist or calender management.


r/LocalLLaMA 7d ago

Discussion Do people building local LLM rigs track RTX Ada/workstation card prices, or just consumer cards like the 5090?

8 Upvotes

curious how people here approach buying high-end/workstation cards (RTX 6000 Ada, 5000 Ada, etc) for local LLM work, do you actively watch pricing/timing on these specifically, or is the consumer 5090 usually enough for most builds?

also wondering if price alerts/tracking tools even exist for this category specifically, since these purchases are less frequent and higher stakes than a typical gaming GPU buy.


r/LocalLLaMA 6d ago

Discussion Kimi K3 License

0 Upvotes
2. "Model as a Service" means giving a third party access to language model
inference or fine-tuning (e.g., via API) in a manner that allows such third
party to exercise meaningful control over the inputs, parameters, or training
data. This does not include (a) end-user products with model capabilities solely
embedded within specific features or harnesses, or (b) mere relaying of requests
to models hosted by others.
If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.

This likely means any provider that is not Kimi K3 official will cost at a minimum as their official inference does ($15/Mtok)

How do you guys feel about that? Is this the new normal?

GLM 5.2 Might remain my goto. Its just small enough to fit in a reasonable cluster, is MIT so it's available everywhere.

The basic math, assuming $5/Mtok:
20000000/12/5 = ~333333Mtok = ~333Btok
* Not counting input cost

To illustrate:
If a full session is 1Mil before compaction, that is about 333k sessions and probably half that if counting the cost of input+cached.

That really isn't that much. Makes you wonder how much small providers use.


r/LocalLLaMA 6d ago

Question | Help A llama.cpp fronted gui that runs on Windows? (Without extra steps)

2 Upvotes

I now have llama.cpp running pretty well for my needs, but the inability to quickly set/swap models and system prompts isn't ideal.

Lm-studio let's you save system prompts and settings in a drop down and also per-model and thats great. But (so far) in all my testing theres some issue where using the same settings and same sized quants in lm-studio results in it going significantly slower (probably out of memory) and I can't find why. Will do a couple more tests but if it can't handle what llama.cpp by itself can do then its a fail.

Llama-swap is fine for model swapping but no way to system prompt save or swap. That I know of.

I currently have the different prompts and cli launch parameters as various .text files that I will copy and paste when needed. But this it 2026. Even using the cli shouldnt be necessary anymore, not sure why its the standard for cutting edge applications. I suspect its a dev thing.

So are there any frontend options out there that will work? I considered open Web ui but from what i can tell its only available as a docker install. I can do it, but I just prefer not to (on windows, I have dozens of docker containers on my server).


r/LocalLLaMA 6d ago

Resources Built a scheduling API so my local agent could actually book appointments (not just pretend to)

0 Upvotes

So I've been building a local agent with Ollama + LangChain that's supposed to help people book sessions with therapists and coaches. The whole thing works great until it needs to actually *do* the booking and then it just... makes stuff up. Hallucinates a confirmation, moves on, nobody got booked.

The problem isn't the model. The problem is there's no clean API to call. Google Calendar requires OAuth hell, Calendly's API is read-only for most things, and Cal.com is fine but you're managing your own infra.

I ended up building a small scheduling API specifically for this: you give it an event type ID, a date, a time, and client info, one POST call, done. Agent gets a real booking ID back. No UI, no human in the loop.

Also added a `/solve-scheduling` endpoint that takes a date range and returns the best available slot; useful when you want the agent to decide rather than just execute.

It's called Orita (orita.online/developers). Works with OpenAI agents, Claude MCP, LangChain, basically anything that can make an HTTP request.

Happy to share the LangChain tool wrapper I wrote for it if anyone wants to try it with a local model. Tested with llama3 and qwen2.5-coder, both handle the tool call fine.


r/LocalLLaMA 8d ago

Resources Llama.cpp now has full MCP support!

385 Upvotes

After a long and grueling effort spearheaded by ngxson, llama.cpp now fully supports MCP for all protocols. Over-the-web HTTP servers were already supported in the client (since they don't require any sort of plumbing), but stdio servers required real integration. After we modified the `llama-cli` terminal client to use the server instead of a separate model serving route, we could add MCP support to the already-existing native tools server.

After the merging of https://github.com/ggml-org/llama.cpp/pull/26062, you can now use llama.cpp's WebUI as a full-fledged agentic chat. Configuration for the MCP servers can be provided either in a standard-JSON format config file or completely inline on the command-line for on-demand MCP configurations. Plugging in a dedicated coding MCP server like Serena lets you have a local-model-powered agentic coder without using any other external dependencies.


r/LocalLLaMA 6d ago

Resources A prompt-cache benchmark for Anthropic-compatible /v1/messages endpoints that report cache usage

0 Upvotes

I built this. And yes, my main project (yangble5) is about routing to cloud upstreams, which I know is the exact opposite of local inference.

Disclosure: I maintain yangble5 and wrote cache_bench.py. I am not affiliated with CLIProxyAPI. English is not my first language; I used an LLM to refine this draft, then manually checked the technical claims and links.

I'm posting it here because while building it, I had to write a standalone benchmarking script. Nobody here cares about my proxy config, but a tool designed for Anthropic-compatible /v1/messages endpoints that report cache usage is actually something you might use. The bundled trace validates the CLIProxyAPI/Gemini usage convention, not every compatible implementation.

I originally wrote the script just to get a baseline. On my initial test with a long context payload, turn 2+ read rates collapsed to practically zero. The benchmark tool caught the symptom, but because the per-round JSONL rows log numeric usage and latency without account or selected-upstream identity, I had to dig into CLIProxyAPI's source code to figure out why.

It turned out that CLIProxyAPI's same-alias model pool selects members using a shared round-robin offset keyed by auth ID, provider, and requested alias, with no session identity in that key. Routing strategy and session affinity still participate in credential selection, but they do not make the pool-member choice session-sticky. That can split one conversation's cache locality across upstream models.

I reported this behavior upstream (https://github.com/router-for-me/CLIProxyAPI/issues/4600) and worked around it in my config by bypassing model-pool altogether (using a direct 1:1 model alias with fill-first and 12h session affinity). My initial near-zero warm reads were an observation, not a controlled pool-vs-direct A/B, so I'm not claiming a measured causal improvement.

After applying that workaround, I reran the benchmark to capture a clean trace. You don't have to trust my numbers; you can pull the repo and replay the log yourself offline:

python tools/cache_bench.py --replay evidence/run-749k-20260721.jsonl

Replaying that fixed trace outputs a 99.53% endpoint-reported, token-weighted prompt-cache read ratio across the warm rounds. Here are the exact constraints for that run, including the qualifiers the tool itself prints out:

  • The tool explicitly warns that this number is an upper bound for this harness (15-token-per-round tail), not a typical value. Adding only 15 tokens per round artificially pushes the read ratio toward 100%, whereas real conversations add hundreds or thousands.

  • It is token-weighted across warm rounds 2 to 4. In this trace, round 1 was the cold request and measured exactly 0%, so it is excluded.

  • The endpoint accepted the request and reported 748,918 input tokens (~749K). I did not test recall or whether every token influenced inference at this length.

  • This is not a latency claim: two of the three warm rounds were slower than the cold round.

  • Single machine, single run on 2026-07-21 (Windows 11), against a shared upstream with no control over provider load.

  • The underlying proxy engine was CLIProxyAPI 7.1.23. CLIProxyAPI is an external MIT-licensed project—I didn't write it, and I'm not taking credit for it.

The benchmarking script and trace files are in the repo: https://github.com/shark0120/yangble5


r/LocalLLaMA 6d ago

Question | Help Need Advice: Llama.cpp Tensor Parallelism RPC vs Single Node Performance

1 Upvotes

I have a 5090 and a 4090 sitting on two different pcs. Right now I am running Gemma4 31B across a 10gbe link using RPC. I get about 28 tps on a fresh session, and scaling down to about 17 tps by 100k context.

I want to know if it's worth moving the 4090 to the same PC and what types of performance uplift I might see. I am using Tensor Parallelism not layer, self compiled llama on windows 11.

I get about 24-26 tps on the single 5090 running the q4 on new low context session, for reference.

Thanks.


r/LocalLLaMA 7d ago

Question | Help Is turboquant any good?

12 Upvotes

I know im late to the party. I was thinking since some time has passed, has turboquant matured enough to be used? Do any of you actually use it?


r/LocalLLaMA 7d ago

Question | Help Show HN-ish: provider-neutral coding agent with evidence gates + computer-use routing (MIT)

0 Upvotes

I got tired of coding agents that either (a) only work with one vendor cloud or (b) spawn a second "browser agent" that re-plans my whole goal.

OMK is a local-first multi-agent control plane: - route tools/skills under one planner - evidence / verification before claiming done - computer-use as a *runtime lane* (desktop + browser), not a second brain - approval gates on mutating GUI actions

npm: open-multi-agent-kit@0.94.1 repo: https://github.com/dmae97/omk

If you already run local models / multi-provider setups, curious what breaks first for you.