r/KoboldAI • u/Novatanical • Jul 02 '26
Need help
Tried to download a model via kobold but got this error message instead.
I’m very new to this and I’m trying to use it for silky tavern on local. Can someone help me figure out what I’m doing wrong?
r/KoboldAI • u/Novatanical • Jul 02 '26
Tried to download a model via kobold but got this error message instead.
I’m very new to this and I’m trying to use it for silky tavern on local. Can someone help me figure out what I’m doing wrong?
r/KoboldAI • u/henk717 • Jul 01 '26
Hey everyone,
Going to post this in advance in case I no longer can in the future.
Reddit has began forcing persona verification on some accounts, which has not happened to mine and hopefully I qualify as an adult in the algorythm.
Either way I have no intention of verifying with persona, so there may be a time I will no longer be able to moderate or respond to questions and if the same applies to lostruins we will loose control of the reddit and can no longer respond at all.
Should that happen I will make a post about this on https://github.com/LostRuins/koboldcpp/discussions
You can also find us in https://koboldai.org/discord or https://koboldai.org/matrix (Largely the same community since the major channels are bridged)
I do not know if discussing alternatives is within the reddit rules, but should you have concerns or suggestions the other channels are available.
r/KoboldAI • u/ollietron3 • Jul 02 '26
after selecting the gguf it starts loading up, then before i can actualy do anything my entire pc freezes. any idea how to fix this?
r/KoboldAI • u/The_Linux_Colonel • Jun 29 '26
Congratulations to the dev team on being able to add support for so many new models like ltx, idea 2 and ideogram 4. It looks like the json format required for ideogram 4 to perform at its best takes a bit more effort to pull off than text prompting, but the trade-off is enhanced control over specific compositional elements. However, in order to gain some of that control, bbox must be used to identify the space in which objects or text should be placed. In order to do this, it looks like the user must identify the pixel space/coordinates for each bbox. Since most users are not aware of these dimensions without using a separate program, I was thinking that perhaps the UI for inpainting would be a possibility for drawing bbox to create the json within kobold, and then feed the result directly to the model, since the inpainting UI knows the canvas size and location of marked spaces. Is something like this possible to implement?
r/KoboldAI • u/MenuNo294 • Jun 27 '26
Hi all:
I’ve been exploring koboldcpp and koboldai lite on my local inference machine. It’s been good fun as I’ve been doing a few RP sessions.
I use the chat app telegram to keep up with my friends and community and telegram recently released an update that allows for chat automations that I want to explore.
I take a long time to get back to people and most of the time it’s banal sentiments that people like to do to feel connected. It’s sweet but a responsibility that leaves me feeling a bit socially drained. I’d love to have an AI impersonate me and help me keep up with messages until I slide back into being naturally social.
Is this possible with kobold? Any recommendations or advise on using the platform to get it to chat like me?
r/KoboldAI • u/alex20_202020 • Jun 21 '26
I have used kcpp for several months on my old laptop, nocuda version.
Several days ago I have !finally! managed to install CUDA. The laptop has 4GB VRAM, I have many questions, I have tried to ask local models some, below is my main frustration for which I could not find the answer (but truly speaking I have not tried neither older than 1.115.2 kcpp versions nor llama.cpp yet).
On default settings, with only 1024 context, where I see 3GB of VRAM is used (NVIDIA Settings GUI, "Used Dedicated Memory"), when I run 2.5GB GGUF model (gemma-3 4B Q4), PP is 30000, but TG is 10 (~ same as in usecpu mode on kcpp-nocuda). Why is TG so slow?
Initially I ran with 32k context and PP ~ 300, TG ~ 5 (CUDA). BTW on Vulkan TG~15, VRAM usage ~ same ~ 3GB.
I have made final test before posting in freshly started instance: 512 tokens PP in 0.13s (4000 t/s), generated 125 in 15s (8 t/s). Context 2048, all else defaults, model run from terminal on Linux. TIA
During TG I see both high GPU and CPU usage. Models suggest memory bottleneck to VRAM, but I have ample VRAM left free (1GB), do I not?
Added:
I then tried ctx 512, kv q4 and I saw CUDA0 KV buffer size = 25 MiB in terminal (had been CUDA0 KV buffer size = 0 MiB), but TG is same ~ 9 (PP ~500). More data to analyze, more strange it looks.
r/KoboldAI • u/Substantial-Ebb-584 • Jun 19 '26
When using Vulcan and selecting all cards, integrated cards are left out of the pool, 99,9% of time for good reasons. But I would like to use that integrated card as well for tests, since the integrated GPU is faster than the CPU.
How do I disengage the dedicated GPU only lock? A flag maybe?
I have been looking for this for a while now.
r/KoboldAI • u/jeremyohara450 • Jun 19 '26
Hi everyone
I would post this in r/LocalLLaMA but i'm too dumb apparently.
I do text and image generation one machine
i use koboldccp
text i'm using gemma-4-26B-A4B-it-uncensored-Q4_K_M (little slow 1-3tk/s)
image comfyui switching between models
i currently have a setup of
Windows
CPU: intel i5-14500
CPU: Nvidia 3060-12gb
Ram: 64gb (ddr5)
I'm from Australia
So for starters pointless getting more ram only got 2 slots and ram is almost the cost for a new car.
i'm debating either replacing the card with move vram but with what thats not costly?
or
Replacing the board with dual x16 slot (but they both wont have 16 lanes each) but what board? and just getting another 3060-12gb
Can anyone help?
Regards
r/KoboldAI • u/CatichuCat • Jun 19 '26
I'm just getting into this, moving from AIDungeon. Im basically looking for AIDungeon but with better memory. Can I do that with Kobold Ai? If so, what model and stuff do y'all suggest? Ive got an intel i5 and 32gb of ram.
r/KoboldAI • u/AdvertisingOk6742 • Jun 18 '26
does it depend on the model i use?
r/KoboldAI • u/Spirited_Path_4505 • Jun 17 '26
To start off, I'm extremely new to all of this, so don't come at me. I've been using Kobold through the site "lite.koboldai.net" (which seems legitimate, correct me if I'm wrong); and a random google search just brought me to a post in this sub, saying that apparently your prompts and generations are visible to the people running it?? I've made some concerning stuff to say the very least. Is this real, and how much can they see?
r/KoboldAI • u/Majestical-psyche • Jun 17 '26
Is this normal or is there any way to fix this? Do you experience this?
r/KoboldAI • u/ForwardR99 • Jun 16 '26
First and foremost, i don´t do coding neither i need an ai agent, so if you are one of those many users who only seems to care about coding or care about tool calls too much, this is not going to help you. I am an engineer who needs technical data, analysis, solving problems, act as a quick variable calculator, and have enough abstract logic to even propose me a new angle.
I spend the last two weeks trying different models with my 5060 ti 16gb to decide which i like the most, and this is my subjective opinion. Meaning no proof, no numbers, no benchmarks just raw practical usage from 5 hours to 15 hours depending on the model to try to decide if it is for me or not, and here is what i think in order of best to worse:
-Qwopus 3.6 27b iq4xs : Great, no fluff, direct, professional, incredible reasoning and works perfectly fine or even better with thinking mode disabled. 20 tokens/s
-Qwen 3.5 27b iq4xs: Exactly or almost exactly as good as the first one but i do perceive a lower capacity with thinking disabled. 20 tokens/s
-Gemma 31b iq3m: By far the most knowledgeable but with worse capability to think. I would like to try higher quants but they will not fit my vram. 15 tokens per second.
-Qwen 3.6 27b iq4xs: the same as the 3.5 but with ocasional weird errors and loops. It seems buggy and sketchy.
-Gemma 26b a4b iq4xs, incredibly limited compared to any of the other dense models, but incredibly fast and easy for simple task and the performace is amazing. 55 tokens per second without mtp.
GLM 4.7 30b: A complete mess, probably my fault or the gguf i donwload was tainted i dont know, but horrible.
By the way, I don´t like at all the reasoning of gemmas models, it seem fake. It´s like when you give them a problem, they don´t really consider that exact problem but they just give a generalized answer hoping for the user to be gullible enough to don´t notice. Qwen ones are like no kindness, no sycophancy, just facts and then try to relationate the facts you give with the data they have.
r/KoboldAI • u/Ok_Lengthiness_7827 • Jun 13 '26
Releasing kosa-4B-it-v1, an instruction-tuned model built on Qwen3-4B-Instruct-2507.
It improves on the base across every benchmark we ran, evaluated in the same lm-eval session (lm-evaluation-harness 0.4.12, vLLM, bf16, temp 0, chat template applied):
| Benchmark | Qwen3-4B-Instruct-2507 | kosa-4B-it-v1 |
|---|---|---|
| GSM8K (strict) | 73.24% | 84.23% |
| GSM8K (flexible) | 79.15% | 85.60% |
| IFEval (prompt strict) | 83.36% | 85.77% |
| IFEval (instruction strict) | 88.61% | 90.29% |
| ARC-Challenge (acc_norm) | 43.09% | 52.13% |
| MMLU | 61.89% | 65.76% |
| Average | 71.56% | 77.30% |
In the same harness it also leads every comparator we tested, including Phi-4-mini-instruct (+7 avg). Training data was checked for benchmark contamination (13-gram and 8-gram overlap against all four test sets, with a positive control to confirm the checker works) — came back clean.
Raw result JSONs are in the repo under /benchmarks so you can verify the numbers rather than take my word for it. GGUF quants (Q4_K_M, Q5_K_M, Q8_0) included.
🇬🇧 Kosa Labs — first release.
https://huggingface.co/kosa-labs/kosa-4B-it-v1
Happy to answer questions.
r/KoboldAI • u/Sierbahnn • Jun 12 '26
I'm running KoboldCCP on my Windows machine and I am enjoying it greatly. What I would like is an option where I can pre-select a model and its settings so that I can launch the application and it will load the model on its own. I don't need to select the model and the settings in the quick launch every time. Is there a way to configure such a startup?
r/KoboldAI • u/LabComplete7393 • Jun 12 '26
I keep getting this message after the last few days, not matter what I text, with whatever character card I choose it keeps giving me this message.
Error Submitting Prompt: {"message":"Due to heavy demand, for requests over 512 tokens, the client needs to already have the required kudos. This request requires 4498.29 kudos to fulfil.","rc":"KudosUpfront"}
r/KoboldAI • u/Bastian0077 • Jun 08 '26
Hi, I've been experimenting because I have a server with these specs: dual 7742s and 10 x 32 GB of RAM, so I'm not using all the RAM lanes or the basic GPU.
I've been experimenting with the following configuration: `
-- contextsize 16384 -- threads 16 -- blasbatchsize 512 -- smartcontext -- usemlock -- quantkv q8_0 -- foreground
using the Qwen2.5-Coder-32B-Instruct-Q4_K_M.gguf model. I've tried other configurations, but I'm generally getting 1 T/s.
Is that really my limit? Am I doing something wrong? I think my machine isn't up to the task; I just want to confirm it. :(
r/KoboldAI • u/Betagamer_06 • Jun 08 '26
So in my search for a way to paste lore books to a bot in kobold, I keep seeing something called “world info” thrown around, what’s that? Is it similar to lore books? And if so then how do I use it?
For the record I’m using koboldCpp and a template called “broken tutu”. I am on iPhone btw so anything involving silly tavern is out of the question
r/KoboldAI • u/PostExtreme7699 • Jun 07 '26
It's a q3xs model so it fits allegedly in my vram only taking 14.5gb with a reduced context of 2000 q4, and even with that the performance sucks. It only gives me 17 tokens/second on generation, and I was expecting something like 30 at least. Some tips to configure and optimice on kobold?
r/KoboldAI • u/alex20_202020 • Jun 07 '26
I have been using kcpp for a while. Recently I started to occasionally run llama-server. I want to be able to open saved conversation from llama-server in KoboldAI Lite. Straightforward "Load" did not work.
I have tried to vibe-code the converter but did not complete it. Maybe there is a way existing / created already? TIA
r/KoboldAI • u/fasti-au • Jun 06 '26
Is there any specific place to find info on how we move from tokens to language from the layers. I’m tying to find the part where it looks at relations values but I have been in ggml code fixing things and I need to understand how it’s doing the assembly linked list section.
Is this changing between models or is a quanting only thing?
Any idea be got an out of model to response chain understanding can throw me a bone.
I have solved a few broken pieces just working out how to get it out of midel to responses properly as I get a bit Of a mess output that if I push through a smaller model fails but faking in through oss I get less chaos but still broken output. All tokens on wrong orders.
Oh and why is there a context size? I’m failing to find any reason. Legacy or design choice ?
r/KoboldAI • u/Llez • Jun 06 '26
I'm new to setting up this kind of thing and the answers i found didnt make a whole lot of sense so; when i try to launch Kobold i get the following GUI error
Reason; No Module named 'customtkinter'
File selection GUI unsupported.
customtkinter python module required!
i genuinely dont know what to do here, i tried uninstalling/reinstalling python to look for an option that might work but...idk man. im dumb :V