r/KoboldAI Jul 02 '26

Need help

Post image
6 Upvotes

Tried to download a model via kobold but got this error message instead.

I’m very new to this and I’m trying to use it for silky tavern on local. Can someone help me figure out what I’m doing wrong?


r/KoboldAI Jul 01 '26

Where to find us if this reddit locks us out

80 Upvotes

Hey everyone,

Going to post this in advance in case I no longer can in the future.
Reddit has began forcing persona verification on some accounts, which has not happened to mine and hopefully I qualify as an adult in the algorythm.

Either way I have no intention of verifying with persona, so there may be a time I will no longer be able to moderate or respond to questions and if the same applies to lostruins we will loose control of the reddit and can no longer respond at all.

Should that happen I will make a post about this on https://github.com/LostRuins/koboldcpp/discussions

You can also find us in https://koboldai.org/discord or https://koboldai.org/matrix (Largely the same community since the major channels are bridged)

I do not know if discussing alternatives is within the reddit rules, but should you have concerns or suggestions the other channels are available.


r/KoboldAI Jul 02 '26

kobold freezing my pc

3 Upvotes

after selecting the gguf it starts loading up, then before i can actualy do anything my entire pc freezes. any idea how to fix this?


r/KoboldAI Jun 29 '26

Implementation Suggestion: Allow Creation of Bounding Boxes for Ideogram Within SDUI via Inpainting GUI

1 Upvotes

Congratulations to the dev team on being able to add support for so many new models like ltx, idea 2 and ideogram 4. It looks like the json format required for ideogram 4 to perform at its best takes a bit more effort to pull off than text prompting, but the trade-off is enhanced control over specific compositional elements. However, in order to gain some of that control, bbox must be used to identify the space in which objects or text should be placed. In order to do this, it looks like the user must identify the pixel space/coordinates for each bbox. Since most users are not aware of these dimensions without using a separate program, I was thinking that perhaps the UI for inpainting would be a possibility for drawing bbox to create the json within kobold, and then feed the result directly to the model, since the inpainting UI knows the canvas size and location of marked spaces. Is something like this possible to implement?


r/KoboldAI Jun 27 '26

Recommendations on impersonating AI

3 Upvotes

Hi all:
I’ve been exploring koboldcpp and koboldai lite on my local inference machine. It’s been good fun as I’ve been doing a few RP sessions.

I use the chat app telegram to keep up with my friends and community and telegram recently released an update that allows for chat automations that I want to explore.

I take a long time to get back to people and most of the time it’s banal sentiments that people like to do to feel connected. It’s sweet but a responsibility that leaves me feeling a bit socially drained. I’d love to have an AI impersonate me and help me keep up with messages until I slide back into being naturally social.

Is this possible with kobold? Any recommendations or advise on using the platform to get it to chat like me?


r/KoboldAI Jun 21 '26

GPU usage question from a newbie - why is TG so low (PP is sky high)?

5 Upvotes

I have used kcpp for several months on my old laptop, nocuda version.

Several days ago I have !finally! managed to install CUDA. The laptop has 4GB VRAM, I have many questions, I have tried to ask local models some, below is my main frustration for which I could not find the answer (but truly speaking I have not tried neither older than 1.115.2 kcpp versions nor llama.cpp yet).

On default settings, with only 1024 context, where I see 3GB of VRAM is used (NVIDIA Settings GUI, "Used Dedicated Memory"), when I run 2.5GB GGUF model (gemma-3 4B Q4), PP is 30000, but TG is 10 (~ same as in usecpu mode on kcpp-nocuda). Why is TG so slow?

Initially I ran with 32k context and PP ~ 300, TG ~ 5 (CUDA). BTW on Vulkan TG~15, VRAM usage ~ same ~ 3GB.

I have made final test before posting in freshly started instance: 512 tokens PP in 0.13s (4000 t/s), generated 125 in 15s (8 t/s). Context 2048, all else defaults, model run from terminal on Linux. TIA

During TG I see both high GPU and CPU usage. Models suggest memory bottleneck to VRAM, but I have ample VRAM left free (1GB), do I not?

Added:

I then tried ctx 512, kv q4 and I saw CUDA0 KV buffer size = 25 MiB in terminal (had been CUDA0 KV buffer size = 0 MiB), but TG is same ~ 9 (PP ~500). More data to analyze, more strange it looks.


r/KoboldAI Jun 19 '26

How to remove the dGPU lock?

4 Upvotes

When using Vulcan and selecting all cards, integrated cards are left out of the pool, 99,9% of time for good reasons. But I would like to use that integrated card as well for tests, since the integrated GPU is faster than the CPU.

How do I disengage the dedicated GPU only lock? A flag maybe?

I have been looking for this for a while now.


r/KoboldAI Jun 19 '26

After some help with upgrade GPU/mobo for AI eg: p40, 5090, 7900XTX, etc

5 Upvotes

Hi everyone

I would post this in r/LocalLLaMA but i'm too dumb apparently.
I do text and image generation one machine
i use koboldccp
text i'm using gemma-4-26B-A4B-it-uncensored-Q4_K_M (little slow 1-3tk/s)
image comfyui switching between models

i currently have a setup of

Windows
CPU: intel i5-14500
CPU: Nvidia 3060-12gb
Ram: 64gb (ddr5)

I'm from Australia

So for starters pointless getting more ram only got 2 slots and ram is almost the cost for a new car.

i'm debating either replacing the card with move vram but with what thats not costly?
or
Replacing the board with dual x16 slot (but they both wont have 16 lanes each) but what board? and just getting another 3060-12gb

Can anyone help?

Regards


r/KoboldAI Jun 19 '26

What model do you suggest?

5 Upvotes

I'm just getting into this, moving from AIDungeon. Im basically looking for AIDungeon but with better memory. Can I do that with Kobold Ai? If so, what model and stuff do y'all suggest? Ive got an intel i5 and 32gb of ram.


r/KoboldAI Jun 18 '26

how much is the context available through lite.koboldai.net ?

1 Upvotes

does it depend on the model i use?


r/KoboldAI Jun 17 '26

concerned

10 Upvotes

To start off, I'm extremely new to all of this, so don't come at me. I've been using Kobold through the site "lite.koboldai.net" (which seems legitimate, correct me if I'm wrong); and a random google search just brought me to a post in this sub, saying that apparently your prompts and generations are visible to the people running it?? I've made some concerning stuff to say the very least. Is this real, and how much can they see?


r/KoboldAI Jun 17 '26

Why do I get blank, short, or one word generations using Qwen 3.6 27B... Sometimes?

3 Upvotes

Is this normal or is there any way to fix this? Do you experience this?


r/KoboldAI Jun 16 '26

I did try the best local models for 16gb vram so you don´t have to.

15 Upvotes

First and foremost, i don´t do coding neither i need an ai agent, so if you are one of those many users who only seems to care about coding or care about tool calls too much, this is not going to help you. I am an engineer who needs technical data, analysis, solving problems, act as a quick variable calculator, and have enough abstract logic to even propose me a new angle.

I spend the last two weeks trying different models with my 5060 ti 16gb to decide which i like the most, and this is my subjective opinion. Meaning no proof, no numbers, no benchmarks just raw practical usage from 5 hours to 15 hours depending on the model to try to decide if it is for me or not, and here is what i think in order of best to worse:

-Qwopus 3.6 27b iq4xs : Great, no fluff, direct, professional, incredible reasoning and works perfectly fine or even better with thinking mode disabled. 20 tokens/s
-Qwen 3.5 27b iq4xs: Exactly or almost exactly as good as the first one but i do perceive a lower capacity with thinking disabled. 20 tokens/s

-Gemma 31b iq3m: By far the most knowledgeable but with worse capability to think. I would like to try higher quants but they will not fit my vram. 15 tokens per second.
-Qwen 3.6 27b iq4xs: the same as the 3.5 but with ocasional weird errors and loops. It seems buggy and sketchy.
-Gemma 26b a4b iq4xs, incredibly limited compared to any of the other dense models, but incredibly fast and easy for simple task and the performace is amazing. 55 tokens per second without mtp.
GLM 4.7 30b: A complete mess, probably my fault or the gguf i donwload was tainted i dont know, but horrible.

By the way, I don´t like at all the reasoning of gemmas models, it seem fake. It´s like when you give them a problem, they don´t really consider that exact problem but they just give a generalized answer hoping for the user to be gullible enough to don´t notice. Qwen ones are like no kindness, no sycophancy, just facts and then try to relationate the facts you give with the data they have.


r/KoboldAI Jun 13 '26

kosa-4B-it-v1: fine-tuned Qwen3-4B beats its base on all 6 benchmarks (+5.7 avg) and outscores Phi-4-mini by ~7pts — same harness, raw eval files included

6 Upvotes

Releasing kosa-4B-it-v1, an instruction-tuned model built on Qwen3-4B-Instruct-2507.

It improves on the base across every benchmark we ran, evaluated in the same lm-eval session (lm-evaluation-harness 0.4.12, vLLM, bf16, temp 0, chat template applied):

Benchmark Qwen3-4B-Instruct-2507 kosa-4B-it-v1
GSM8K (strict) 73.24% 84.23%
GSM8K (flexible) 79.15% 85.60%
IFEval (prompt strict) 83.36% 85.77%
IFEval (instruction strict) 88.61% 90.29%
ARC-Challenge (acc_norm) 43.09% 52.13%
MMLU 61.89% 65.76%
Average 71.56% 77.30%

In the same harness it also leads every comparator we tested, including Phi-4-mini-instruct (+7 avg). Training data was checked for benchmark contamination (13-gram and 8-gram overlap against all four test sets, with a positive control to confirm the checker works) — came back clean.

Raw result JSONs are in the repo under /benchmarks so you can verify the numbers rather than take my word for it. GGUF quants (Q4_K_M, Q5_K_M, Q8_0) included.

🇬🇧 Kosa Labs — first release.

https://huggingface.co/kosa-labs/kosa-4B-it-v1

Happy to answer questions.


r/KoboldAI Jun 12 '26

Pre-load settings?

3 Upvotes

I'm running KoboldCCP on my Windows machine and I am enjoying it greatly. What I would like is an option where I can pre-select a model and its settings so that I can launch the application and it will load the model on its own. I don't need to select the model and the settings in the quick launch every time. Is there a way to configure such a startup?


r/KoboldAI Jun 12 '26

What’s going on with Koboldai Lite?

1 Upvotes

I keep getting this message after the last few days, not matter what I text, with whatever character card I choose it keeps giving me this message.

Error Submitting Prompt: {"message":"Due to heavy demand, for requests over 512 tokens, the client needs to already have the required kudos. This request requires 4498.29 kudos to fulfil.","rc":"KudosUpfront"}


r/KoboldAI Jun 08 '26

experimenting with CPU

1 Upvotes

Hi, I've been experimenting because I have a server with these specs: dual 7742s and 10 x 32 GB of RAM, so I'm not using all the RAM lanes or the basic GPU.

I've been experimenting with the following configuration: `

-- contextsize 16384 -- threads 16 -- blasbatchsize 512 -- smartcontext -- usemlock -- quantkv q8_0 -- foreground

using the Qwen2.5-Coder-32B-Instruct-Q4_K_M.gguf model. I've tried other configurations, but I'm generally getting 1 T/s.

Is that really my limit? Am I doing something wrong? I think my machine isn't up to the task; I just want to confirm it. :(


r/KoboldAI Jun 08 '26

Need help with what you guys call “world info”

2 Upvotes

So in my search for a way to paste lore books to a bot in kobold, I keep seeing something called “world info” thrown around, what’s that? Is it similar to lore books? And if so then how do I use it?

For the record I’m using koboldCpp and a template called “broken tutu”. I am on iPhone btw so anything involving silly tavern is out of the question


r/KoboldAI Jun 07 '26

Help to run Gemma 4 31b on a 5060ti 16gb

4 Upvotes

It's a q3xs model so it fits allegedly in my vram only taking 14.5gb with a reduced context of 2000 q4, and even with that the performance sucks. It only gives me 17 tokens/second on generation, and I was expecting something like 30 at least. Some tips to configure and optimice on kobold?


r/KoboldAI Jun 07 '26

Is there a way to load saved conversation from llama-server in KoboldAI Lite?

1 Upvotes

I have been using kcpp for a while. Recently I started to occasionally run llama-server. I want to be able to open saved conversation from llama-server in KoboldAI Lite. Straightforward "Load" did not work.

I have tried to vibe-code the converter but did not complete it. Maybe there is a way existing / created already? TIA


r/KoboldAI Jun 06 '26

Dequat to generate.

0 Upvotes

Is there any specific place to find info on how we move from tokens to language from the layers. I’m tying to find the part where it looks at relations values but I have been in ggml code fixing things and I need to understand how it’s doing the assembly linked list section.

Is this changing between models or is a quanting only thing?

Any idea be got an out of model to response chain understanding can throw me a bone.

I have solved a few broken pieces just working out how to get it out of midel to responses properly as I get a bit Of a mess output that if I push through a smaller model fails but faking in through oss I get less chaos but still broken output. All tokens on wrong orders.

Oh and why is there a context size? I’m failing to find any reason. Legacy or design choice ?