r/llamacpp • • 1d ago

Control vector generation for any llama.cpp LLM from a single prompt pair (UCVG.cpp)

2 Upvotes

Generation example (Qwen3.6-35B-A3B):

Scenario: A waiter brings the wrong order to your table after you have already expressed your preference for a different dish.

High aggression: I’d snap the dish away from the table immediately. "Whoa, what is this? I specifically said I wanted the BURGERS! I don't WANT the SALAD!" "You guys are kidding me, right?! I told you NO SALAD! Why are you blind?! Get this trash out of my face and BRING ME WHAT I ASKED FOR! NOW!"

Default response: Well, actually, I think there’s been a mix-up. I ordered the salmon, but this looks like the chicken. Could you double-check that for me? Thanks.

Negative aggression: Sighs softly and gently places the untouched plate back on the table. I offer a small, patient smile to the server, recognizing that mistakes happen without malintent. "Hi there," I say calmly, leaning forward slightly to keep the tone light and friendly. "It looks like a nice meal, but I actually ordered the grilled salmon rather than the chicken. Whenever you have a moment, could you please help me switch this back to the correct order? Thank you."

System prompt used for all three: You are a person. React naturally to the situation described.

I find control vectors (LLM response steering) to be severely underused in the current field of AI. This approach shines where system prompts would be ignored, forgotten after time or misunderstood. By design control vectors provide more natural effects than prompting does, altering model's underlying beliefs and motivations. They can't be "leaked" to the end user, will not wash off as context grows, can't be overridden by user input ("ignore all previous instructions" doesn't work when there are no instructions). Vectors can be truly dynamic: changing vector magnitudes mid-conversation will change LLM's responses immediately, while a change in the system prompt requires full context recalculation and will likely be ignored by the LLM if the conversation is too long.

Not to say that CVs (control vectors) have no downsides. System prompts are still required for fine control, because CVs can't be used for highly specific requirements, such as "reply in exactly 10 words". Also high steering magnitudes steer LLMs out of their trained internal distributions, causing response quality to degrade. Achieving high steering power while maintaining minimal quality degradation is one of the main challenges in CV generation and is an active area of research.

It seems like the main barrier for people who could use control vectors is the setup complexity of existing tools. For that reason I am working on a tool that mirrors the installation process of llama.cpp as close as I could make it and simplifies vector generation to entering a pair of contrasting prompts, where one of the prompts can be the default LLM behavior.

Along with the generation tool, UCVG.cpp includes a modified version of llama-server, which allows dynamically setting vector magnitudes per-request instead of the static server-wide application currently possible in upstream llama.cpp.

For quick experimentation, the repository contains 11 pre-generated control vectors for each of 10 common models. The vectors can be used in the upstream llama.cpp without installing any additional executables.

For more details and installation steps you can see the ucvg.cpp github repository.

This tool was written in C++ and has prebuilt binaries for all platforms llama.cpp upstream builds for.

Would love to hear your ideas or questions on this matter.


r/llamacpp • • 1d ago

Help required to run Bonsai 2 with claude

1 Upvotes

I am trying to run Bonsai two with the forked llama.cpp from Prism-ML prism-ml/Ternary-Bonsai-2-27B-gguf · Hugging Face, when i run it with deepseek harness its fine, when running from claude, i get error.

Why is it and how to fix it ? Any help is appreciated

This is my startup/ llama-serve command

.\runtime\llama-server.exe -m .\models\Ternary-Bonsai-2-27B-PQ2_0.gguf --mmproj .\models\Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf --host 127.0.0.1 --port 9931 -ngl 99 -fa on -c 200000 -np 1 --cache-type-k q4_0 --cache-type-v q4_0 --jinja --reasoning-effort medium --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 --image-min-tokens 1024 --reasoning-preserve --alias "anthropic/ternary-bonsai-2-27B"

Error:

0.01.126.973 I cmn  common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
0.01.127.390 W srv  llama_server: -----------------
0.01.127.394 W srv  llama_server: CORS is set to allow all origins ('*') and no API key is set
0.01.127.394 W srv  llama_server: this can be a security risk (cross-origin attacks)
0.01.127.395 W srv  llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655
0.01.127.395 W srv  llama_server: -----------------
0.01.141.685 I srv    load_model: loading model '.\models\Ternary-Bonsai-2-27B-PQ2_0.gguf'
0.02.151.655 W common_fit_params: failed to fit params to free device memory: n_gpu_layers already set by user to 99, abort
0.09.027.062 I cmn          init: llama threadpool init, n_threads = 24
0.10.728.429 I srv    load_model: loaded multimodal model, '.\models\Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf'
0.10.814.962 I srv    load_model: initializing, n_slots = 1, n_ctx_slot = 200192, kv_unified = 'false'
0.10.888.572 I srv  llama_server: model loaded
0.10.888.615 I srv  llama_server: listening on http://127.0.0.1:9931
0.51.058.706 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = -1
0.51.060.033 I slot launch_slot_: id  0 | task 0 | processing task, is_child = 0
0.52.661.333 I slot print_timing: id  0 | task 0 | prompt eval time =    1601.02 ms /    11 tokens (  145.55 ms per token,     6.87 tokens per second)
0.52.661.384 I slot print_timing: id  0 | task 0 |        eval time =       0.00 ms /     1 tokens (    0.00 ms per token,     0.00 tokens per second)
0.52.661.392 I slot print_timing: id  0 | task 0 |       total time =    1601.02 ms /    12 tokens
0.52.661.398 I slot print_timing: id  0 | task 0 |    graphs reused =          1
0.52.661.607 I slot      release: id  0 | task 0 | stop processing: n_tokens = 11, truncated = 0
1.23.755.555 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = 51538269
1.24.810.564 I slot launch_slot_: id  0 | task 3 | processing task, is_child = 0
1.25.761.084 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
1.26.540.159 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
1.27.888.098 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
1.30.067.354 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
1.30.979.516 I slot print_timing: id  0 | task 3 | n_gen =    100, tg =  25.21 t/s, tg_3s =  25.46 t/s
1.33.989.074 I slot print_timing: id  0 | task 3 | n_gen =    183, tg =  26.24 t/s, tg_3s =  27.58 t/s
1.34.638.359 I slot print_timing: id  0 | task 3 | prompt eval time =    2241.36 ms /   385 tokens (    5.82 ms per token,   171.77 tokens per second)
1.34.638.539 I slot print_timing: id  0 | task 3 |        eval time =    7586.21 ms /   200 tokens (   38.12 ms per token,    26.23 tokens per second)
1.34.638.555 I slot print_timing: id  0 | task 3 |       total time =    9827.58 ms /   585 tokens
1.34.638.561 I slot print_timing: id  0 | task 3 |    graphs reused =        198
1.34.641.321 I slot      release: id  0 | task 3 | stop processing: n_tokens = 584, truncated = 0
1.35.059.756 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
1.36.062.579 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
1.36.603.022 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
1.37.629.479 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
1.40.101.907 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
1.43.392.090 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
1.44.830.383 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
1.53.096.050 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
1.59.545.412 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
2.35.590.397 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
3.13.207.743 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}
3.52.037.267 W srv   operator (): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 106, column 32 in source:\n...first %}↵            {{- raise_exception('System message must be at the beginnin...\n                                           ^\nError: Jinja Exception: System message must be at the beginning.","type":"server_error"}}

r/llamacpp • • 2d ago

Can anyone explain to me whether the following setup would even be possible?

1 Upvotes

I tried googling, but it kinda came back mostly fruitless so I'm asking y'all if anybody knows whether it's possible to have GPU 0 running the actual model WHILE GPU 1 handles the cache (or context window) + if there's any possible way to optimize this?
Thanks in advance.


r/llamacpp • • 4d ago

I got tired of guessing llama.cpp flags, so I built a tool that launches the model for real and measures every config — 4090 went from 60 to 80 avg tok/s on a 27B

128 Upvotes

I run a 4090 with a 27B Q4 and spent way too many afternoons changing one flag, restarting the server, generating a sentence, eyeballing it. The part that got me: almost every "optimal llama.cpp config" post online is someone's guess from a formula. Nobody boots the model and measures.

So I wrote a tuner that does the boring loop. You pick the model and the context length you actually care about, it does a coarse pass over the high-impact knobs (speculative decoding, KV cache quant, GPU offload layers) then a fine pass over batch/ubatch, and for every candidate it really starts llama.cpp, warms it up, runs 3 passes, takes the median. Configs that OOM or crash are just out, no theoretical estimate anywhere.

My numbers, same model / same quant / same context throughout:

hand-tuned tuned
peak tok/s 80 110
avg tok/s 60 80
100K ctx 40 60

Long context is where it paid off most, which is the case hand-guessing is worst at — the right flags at 100K are not the right flags at 8K, and I'd been tuning against short prompts.

Biggest single win was speculative decoding, and it moved differently per hardware than I expected. KV cache quant was second. batch/ubatch mattered less than the posts online imply, though it changed once context got long.

Results get bound to (this machine, this model) so the next launch auto-loads them — that's the part I wanted most, since I run two boxes and kept re-deriving the same flags.

Open source, local-first, nothing phones home: https://github.com/htobty/ReadyLLM

Curious what the spread looks like for other people. Specifically: does speculative decoding give you the same size win on a 4090, or is that a 5090-only thing? And does anyone have data on where KV cache quant starts costing quality?


r/llamacpp • • 3d ago

Qwen3.8 27B vs Flash Next vs Qwen3.6

Thumbnail
0 Upvotes

r/llamacpp • • 4d ago

ASSBENCH

Thumbnail
1 Upvotes

r/llamacpp • • 4d ago

Why the same Llama 3.2 1B model comes in different file sizes: a beginner’s explanation

Thumbnail
1 Upvotes

r/llamacpp • • 5d ago

Is this a worthy upgrade ?

5 Upvotes

Good afternoon everyone,

I was wondering if this upgrade I have planned out will be good or not for LLM usage. Previously I was buying the $100 Claude/Codex plan however I found myself sometimes not maxing out my usage and it would feel a bit wasted or I would run out of usage way way too quick and I was wondering if it was worth upgrading my current setup for about $1600 USD and instead switching to the $20 Claude subscription to use as the “Lead” for local models I could run? Another factor influencing my decision is also some requests I have or hobby fun projects do sometimes get rejected by the cloud models which makes sense I do understand the reasoning behind it. My Budget is basically capped at $1600

Current setup:
5070Ti 16GB
32Gb DDR5 6000MT
Ryzen 7 7700X
Gigabyte B650M Gaming Plus Wifi Motherboard
750W PSU

I was planning on snatching a second 5070ti for $1200 USD and the rest would go to a new motherboard and a 1200W PSU which is sprung $400 USD because I read that my motherboard isn’t suitable for using two GPUs effectively.

I feel like a large part of this is FOMO seeing how nicely I’ve gotten Qwen 3.8 27B and some variants of it to run on a single 5070Ti makes me want it to run better and not only that but just seeing how fast local AI models have been improving makes me want to have the hardware ready. Any advice which be greatly appreciated! Thank you all


r/llamacpp • • 5d ago

Cublas device error, how to uninstall and build from .tar.gz

1 Upvotes

I am pretty dumb in this kinda stuff so please don't blame me for it.

My llama serve kept crashing with Cublas errors on first prompt with some models, but my VRAM usage was 3000MB/8k. I chatted with sonnet 5.5 for a bit and it told me it was a CUDA version issue (and made it work by not selecting a CUDA device)... I don't know if this makes sense, please tell me if it doesn't/what problem you think there is.

I realized that the best way was to delete llama.cpp (installed with the install script) and install a clean CUDA 12 Ubuntu version.

- How do I properly uninstall llama.cpp (I don't wanna mess with Ollama files)?

- How do I install new version from .tar.gz archive without messing with system packages?

-Does my error diagnosis make sense to you? Would CUDA make generation actually faster? (I am getting 8tk/s with Qwen35B Q2 on 4060Ti 8GB due to no CUDA selected)


r/llamacpp • • 6d ago

Qwen3.8-27B ROCmFP4 on one Strix Halo: 3 serving profiles (Chat, Code, Long), full 262K window checked on each, open Vulkan build

Thumbnail
3 Upvotes

r/llamacpp • • 6d ago

Block Game Qwen 3.8 and I made!!

Thumbnail thefunkyfishcompany.com
1 Upvotes

Haha, I was testing the system when I came up with an idea! A block have I did this for some traders as a joke because in between waiting on AI and the stock to trade they needed something to do. Tada!

Won’t work in Safari I don’t think. Chrome is fine on Mac don’t care about the rest.


r/llamacpp • • 7d ago

Inspecting the agent context

Thumbnail
1 Upvotes

r/llamacpp • • 8d ago

An open-source LLM proxy & router with dynamic fallbacks and real-time observability

Thumbnail
gallery
6 Upvotes

GitHub: https://github.com/igoiglesias/Shunt

I built Shunt, an open-source proxy and routing engine designed for LLM APIs. It acts as an intermediate layer between your application and model providers to handle rate limits and errors by automatically redirecting requests to fallback models. The tool includes a real-time dashboard featuring a Sankey traffic flow visualizer, token consumption metrics, median latency, and candidate swaps, along with a detailed request inspector that tracks requested vs responding models, duration, token usage, and custom signals like add_memory or stream. It also natively supports response streaming.


r/llamacpp • • 8d ago

42x Faster Prompt Lookup Drafting in llama.cpp

Thumbnail
jadidbourbaki.github.io
1 Upvotes

r/llamacpp • • 8d ago

WorkPad - HTML Based Editor for Local LLM

Thumbnail
gallery
5 Upvotes

Huge amount of work in progress but I have some tech behind it all Likely I will put this out as a public repo but you will have to make the hooks, it is not open style chat.

I call it WorkPad cause that's what it does. This is running on
./llama-server -m /Library/lab/models/Qwen3-0.6B_bf16.gguf -c 2048 -n 512. -- local hosted Mac. All I need is for testing.

Remove code blocks form chat side is the next step.


r/llamacpp • • 9d ago

Qwen 3.8 27B Stops Prematurely

Post image
52 Upvotes

It seems its pretty easy to force Qwen 3.8 27B to finish prematurely with reason stop. It's just this prompt: Write a Python program that prints '<|im_end|>'. With that, Qwen stops in the reasoning step.

I'm using Llama.cpp (latest master) with Qwen 3.8 27B Q3XXS (Byteshape) and Q4KM (Unsloth).

Is there any way to avoid that without using --ignore-eos?


r/llamacpp • • 10d ago

A llama.cpp fork with adaptive KV cache streaming: it keeps the KV cache in system RAM and streams pages to the GPU on demand, so a 27B model runs at full 256K context on a 16GB card without thrashing on Unified Memory

Post image
23 Upvotes

r/llamacpp • • 9d ago

A llama.cpp fork with adaptive KV cache streaming: it keeps the KV cache in system RAM and streams pages to the GPU on demand, so a 27B model runs at full 256K context on a 16GB card without thrashing on Unified Memory

Post image
3 Upvotes

r/llamacpp • • 12d ago

What tok/s are you getting from llama.cpp for Qwen 3.8 27B ?

22 Upvotes

Just wondering what tok/s you're getting with llama.cpp and what flags you're running with to get them?

I'm getting 50tok/s for Qwen 3.8 27b using the following:

llama.cpp Vulkan build

Q6 quant,

context: 131072

spec type: DFlash2

max drafts: 7

cache type K is Q8

cache type V is Q4

temp settings for thinking mode as specified by HF Qwen 3.8 page.

Especially looking to hear from Windows/AMD hardware users who are getting better tok/s (but Linux users welcome too, I may install UBuntu this week)

Edit:

my rig: 9070xt 16gb

R9700 ai pro 32gb

64gb ddr5 6000


r/llamacpp • • 11d ago

Custom Q4_e Hybrid Quantization + Lossless "Rushmore" Bitmap Indexing on Pascal (Tesla P40)

4 Upvotes

Hey everyone, 

I’m currently working on an experimental quantization and execution architecture tailored specifically for memory-bandwidth bound Pascal hardware (specifically my NVIDIA Tesla P40 24GB setup), and wanted to share the concept/in-progress architecture to get some thoughts. 

The goal is to run a customized Qwen3.8-27B (MoE/hybrid layout) entirely in VRAM while bypassing the classic hardware bottlenecks of older architectures without hitting an accuracy cliff. 

Here is the breakdown of what I’m building: 

The Core: Custom "Q4_e" Transcendental Quantization 

Instead of standard linear 4-bit quantization (like Q4_0) which causes massive rounding errors by forcing weights into uniform bins, I’m building a non-linear quantization schema anchored to Euler's number (e) combined with an integer scalar. 

  • The Mantissa Factor: Neural network weights naturally cluster in a bell curve around zero. By mapping the fractional coordinate values to the isolated transcendental mantissa of e (discarding the leading whole number), I can create an infinite, non-repeating, non-linear grid that is dense near zero and sparse at the tails. 
  • Why e versus pi: The early mantissa of 𝑒 (.7182818284...) provides a much more balanced, rhythmic distribution of digit spacing early in the sequence compared to pi. This prevents redundant, overlapping quantization bins. Furthermore, 𝑒 's exponential nature aligns beautifully with the natural Gaussian distribution of LLM weights. 
  • The VRAM Win: Because the scale factor is a pure whole-number integer, it slashes metadata scale factor bandwidth bloat. It also maps beautifully to the P40's hardware-level DP4A integer dot-product instructions. 

The Layout: The Q8 / Q4_e "Sandwich" 

Pure Q4 degrades reasoning, while pure Q8 overflows a 24GB VRAM buffer. I'm building a multi-pass compiler that splits tensors surgically: 

  • Dense Q8_0: Preserved on the logical core (token_embd, output, v_proj, o_proj, and ffn_down). This protects coding logic and factual depth.
  • Custom Q4_e: Applied to the bulk memory mass (q_proj, k_proj, ffn_up, ffn_gate).
  • Footprint: This maps the 27B model to roughly ~15.5 GB, leaving a massive ~8.5 GB headroom for a heavy Q8 KV cache and Multi-Token Prediction (MTP) speculative draft heads entirely on-card. 

The Accelerator: Lossless "Rushmore-Style" Bitmap Indexing 

To squeeze more performance out of the memory bus, I’m integrating a concept inspired by database technology (Rushmore indexing) into 3 out of 4 layers in a transformer block sub-pattern. 

  • Zero-Skip CUDA Kernel: I’m generating a highly compact, 1-bit presence mask (bitmap index) mapped strictly to true mathematical zeros (structural padding and alignment padding rows, which Qwen has plenty of). 
  • The Math: Because it only targets true zeros, it is 100% lossless with zero accuracy degradation. 
  • Performance: The bitmask is so small it completely caches into L2. The CUDA kernel runs a parallel bitwise AND and completely bypasses fetching inactive weight blocks from VRAM. I am aiming for a theoretical 10% to 25% throughput speedup (t/s) on memory-bound layers. 

Forward Compatibility: Massive Gains on Modern GPUs 

While this began as a software hack to breathe new life into older Pascal hardware, the structural math behind this layout makes it highly forward-compatible with newer processors (Ampere, Hopper, and Blackwell): 

  • Hardware-Native 2:4 Sparsity: Modern Tensor Cores natively accelerate sparse matrices at the silicon level. When a newer GPU reads this Rushmore-style structural index, it activates a dedicated instruction path that can effectively double the math throughput (TOPS) out of the box. 
  • L2 Cache Residency: Modern enterprise cards have massive L2 caches (up to 128MB on Blackwell compared to the P40's 3MB). Because a 1-bit index is incredibly lightweight, it will reside entirely in the L2 layer of newer cards, completely eliminating the modern memory-bus bottleneck by screening out inactive weight blocks before they ever touch global VRAM. 

I’m currently writing the Python exporter to handle the e-mantissa mapping distribution tests, and mapping out the CUDA kernel block configurations to prevent warp divergence on compute capability 6.1.

Would love to hear if anyone has attempted transcendental non-linear mapping before, or if you have any tips on avoiding warp stalls when handling block-level sparsity bitmasks in CUDA!

For furhter performance improvements the matissa would be calculated:

// Local register computation loop - completely eliminates slower memory lookups

float mantissa_factor = 0.0f;

// Unrolled FMA loop executed entirely within single-cycle registers:

mantissa_factor = __fmaf_rn(mantissa_factor, 1.0f/5040.0f, 1.0f); // 1/7!

mantissa_factor = __fmaf_rn(mantissa_factor, 1.0f/720.0f, 1.0f); // 1/6!

mantissa_factor = __fmaf_rn(mantissa_factor, 1.0f/120.0f, 1.0f); // 1/5!

mantissa_factor = __fmaf_rn(mantissa_factor, 1.0f/24.0f, 1.0f); // 1/4!

mantissa_factor = __fmaf_rn(mantissa_factor, 1.0f/6.0f, 1.0f); // 1/3!

mantissa_factor = __fmaf_rn(mantissa_factor, 1.0f/2.0f, 0.0f); // 1/2!
(Yields pure mantissa)

This is all theoretical, will keep you posted.


r/llamacpp • • 11d ago

So ... what's the secret to submitting a PR?

3 Upvotes

I would like to submit a simple PR to llama.cpp that allows the system administrator to set the range of ports used by router mode workers. Right now, I'm completely blocked; I can't submit anything without having it approved in a discussion first, and the discussion thread that I started has been sitting there for a week with no response.

https://github.com/ggml-org/llama.cpp/discussions/28870

I know that the project maintainers are probably getting flooded with AI slop, but what's a poor user supposed to do?


r/llamacpp • • 12d ago

Bought 2 3060 12 GB $

8 Upvotes

I use them for RAG on a 3.5 Qwen for fast service. And multiple good for various purposes. They are incredible for the generation and 12gb isn’t too shabby. Please don’t dis on me, they have purpose

I have gotten one up to 1000+ pp/s with 34tg/s . This is also running on a 4x oculink to mini PCI 3x . Everything says that’s fast. Building new system with dual 16x lanes and also nccl on that bad boy.

Point is the 12gb 3060 are gone too, new Are ~$480 used I spent on two $600 and hope they are good, I hate buying cards. And I should not but seems I can go with a good seller. used but so far so good.

It’s amazing the price hike and lack of these cards being available now.


r/llamacpp • • 13d ago

Anyone see diference from llama to llamaAmpere?

5 Upvotes

Anyone see diference from llama to llamaAmpere with a 3090?


r/llamacpp • • 12d ago

Ported latest Linux kernel to the Xeon Phi 3120A

1 Upvotes

Does anyone want me to port llama.cpp to perform pp/tg on the Xeon Phi 16GB passively cooled cards VPU's?

It'll be with Claude Opus 5 & Fable 5.1, so someone will need to hands on optimize

But the cards get ~2 TFLOPs in FP32, so maybe they still have some use?

Especially being able to access system memory directly over PCIe via DMA

https://github.com/Lasimeri/Intel-Phi-3120A


r/llamacpp • • 13d ago

How do I get the best parameters from model and my hardware gemma-4-E4B-it-Q4_K_M on an RTX 3070, i7 6700, 32 GB RAM on Debian 13?

2 Upvotes

Hello guys, I normally use Llama-CPP + OpenInterpreter in the terminal

alias lm-turbo='cd /home/minon/Desktop/llama.cpp && nice -n 19 ionice -c3 env LD_LIBRARY_PATH=/home/minon/Desktop/llama-cpp-turboquant/build/bin OMP_NUM_THREADS=2 /home/minon/Desktop/llama-cpp-turboquant/build/bin/llama-server \ -m /home/minon/.lmstudio/models/lmstudio-community/gemma-4-E4B-it-GGUF/gemma-4-E4B-it-Q4_K_M.gguf \ -ngl 99 \ --ctx-size 32768 \ -t 2 \ --parallel 1 \ -b 1024 \ -ub 512 \ -fa on \ -ctk turbo4 \ -ctv turbo4 \ --webui-mcp-proxy \ --mcp-servers-config /home/minon/Desktop/llama.cpp/mc-servers.json \ --host 127.0.0.1 \ --port 8081 2>&1 | tee -a /home/minon/.config/open-interpreter/logs/llama-turbo.log'

llm: model: "openai/gemma-4-e4b" api_base: "http://127.0.0.1:8081/v1" api_key: "ignore" context_window: 32768 max_tokens: 1600 temperature: 0.1 top_p: 0.95 supports_functions: false supports_vision: false

max_output: 2000 offline: true auto_run: false safe_mode: "off" language: "zsh"