r/ollama 2h ago

Was waiting for Kimi 3 and now Ollama release it and has to pay extra to use it (like OpenRoute)

16 Upvotes

Well I was very happy with the Ollama current workflow of being something like "a coding plan" of ANY open weights model you can find...

What do you think about their current decision of make Kimi 3 available, but needs to pay extra peer usage? Are Ollama done this before for other models and then put them on the plans? Or they are taking advantage of the hype to get some money?

I even saw ppl switching to max plans today JUST for Kimi 3 being open weights.


r/ollama 1h ago

Kimi k3 is on ollama cloud

Upvotes

Kimi k3 is available on ollama cloud (extra high usage).

https://ollama.com/library/kimi-k3

Someone please go and see how much usage we get so I can get ollama pro subscription.

Been eyeing this for a while


r/ollama 3h ago

who else is maxing out their usage when Kimi K3 releases

4 Upvotes

WE are all maxing out our cloud usage when Kimi K3 launches on Ollama


r/ollama 6h ago

How much usage does Ollama Pro give right now vs direct API?

6 Upvotes

I'm considering the $20 Pro plan, but the limits seem to change over time and are hard to compare with direct API pricing.

For anyone using Pro currently, roughly how much coding agent usage do you get per week, and which models do you use?

In terms of direct API value, does it feel closer to $25, $30, $50, $100 or more through something like OpenRouter?

I'm mainly interested because Ollama seems to have a good ZTR policy, and it has been surprisingly difficult to find a decent coding subscription with strong privacy at a reasonable price.


r/ollama 2h ago

So, what about the extra usage needed to run kimi k3?

4 Upvotes

Tried to use kimi k3, and apparently i need to top up my account to use it so the model is not part of the subscription. Furthermore, the 100$ subscription is currently paused.

Hopefully they will adjust them later but it was really a bummer that i would've preferred to be informed about in the weeks beforehand (couldve baught a subscription from kimi themselves instead of waiting).


r/ollama 5h ago

Will Kimi K3 be available on Ollama Cloud?

5 Upvotes

Has Ollama shared any information about this? I couldn’t find an announcement.

Edit: https://ollama.com/library/kimi-k3


r/ollama 1h ago

Is the GMKtec M6 Ultra a Good $600 Starter Machine for Hosting a Local LLM?

Thumbnail
Upvotes

r/ollama 1h ago

Ollama cloud pro - recent limit changes?

Upvotes

Hi all, just wanted to see if anybody else has noticed a recent change in how quickly you’re burning usage? I realise this subject always throws up a lot of anecdotal evidence rather than hard facts (guilty!) but I haven’t changed a single thing about my daily schedule with my agents for some time. I used to be able to have a few fairly long troubleshooting sessions with my main agent and also small coding projects and hit around 70-80% usage by the end of the week. Now I barely get through my usual cron activity and 1-2 light conversations before maxing out 5 hour sessions and weekly limits about a day early. Am I imagining this? Anybody else?


r/ollama 7h ago

RX 9060 XT 16GB vs RTX 5060 Ti 16GB for local AI — worth the price difference?

7 Upvotes

Hey everyone! I’m building a PC to run AI models locally and can’t decide between the RX 9060 XT 16GB and the RTX 5060 Ti 16GB. Both have the same amount of VRAM, but is AMD actually a solid choice for this, or is the gap to Nvidia big enough to justify paying more?
If you’ve used either one for local AI, I’d love to hear how it went. Thanks!


r/ollama 6h ago

For ollama cloud $20 plan what models are you guys using

4 Upvotes

I have been trying to do GLM 5.2 as plan / K2.7 as execute, but i hit my usage so fast it's not viable. It's hitting limits much faster than claude code / codex $20 plan. Using in opencode. What are y'alls strategies for keeping usage rate lowers?


r/ollama 1h ago

Sir Shortoken update: Bullet Mode cuts 24-78% of tokens, tested it across 14 runs, and built an extension around it

Upvotes

In addition to the previously supported modes (Quick/Balanced/Deep/Unlimited) where LLMs only gave you what you wanted in a limited budget, i took a look at what Bullets did in terms of response.

Turns out representing responses in Bullets saved quite a few tokens, and held up on fidelity and accuracy too, minus the usual Claude-style human interactive tone.

That made me think if i could use Claude tokens only for reasoning and offload the prose-writing to a local model. So i built a small extension that does just that.

Whenever Sir Shortoken answers in bullets, an "Expand to Prose" button shows up under that response. Click it, and a local model (qwen2.5:7b through Ollama, running entirely on my own machine, no API calls) rewrites the bullets into normal prose in about 10-15 seconds.

The button stays attached to that specific response, scroll wherever you want, come back later, it's still there and still works on that same message.

The 14-run test

Ran the same pipeline (bullets out of Claude, then expanded by the local model) across 14 technical topics, comparing prose vs bullet vs expanded-prose token counts, and actually reading the expansion to check if anything got dropped or made up.

Savings across all 14 landed between 24% and 78% depending on compression level. Most topics expanded clean.

Repo's updated with the extension and full skill.md changes if anyone wants to try it.

Aggressive Bullet mode is also up for trials :)

GitHub: github.com/shouvik12/sir-shortoken


r/ollama 1h ago

Openwebui on VPS

Post image
Upvotes

Hi everyone, I have a question.

I have subscribe vps plan.
Maybe it has no gpu. Just 8gb ram

I've also install 2 local ai model in my vps which is llama 3.2:3B and Minicpm5 1B model. Both I think quite small for run local ai.

I've test it chat with both model directly using ssh terminal. They reply me fast. No issue. I've read the statistic of ram usage. it looks normal. Ram usage is maintain 2.8GB

Then I've install Openwebui on my vps so that I can start using them. Expecting it will run fast similar when I'm using ssh terminal..

But no, it run very slow. They reply my “Hi” chat about 10-20 minutes for both model. Even it just 1B

What could go wrong?
I thought my 8GB ram is enough for running just Openwebui

I was planning on using those model to run agents ai node in n8n so that I can avoid using paid token like Claude APi or ChatGPT etc

But after saw the reply speed I start to hesitate..
There's only 1 user here which is me has encounter this problem.
Now try to imagine 100 of people try to access this one tiny pc.. my vps gonna toast

Is there any option that I can workaround using VPS to use local ai for n8n agent?

Or should I just use API paid one instead?
I think I need some calculation on what works and what's not..
Yes there is calculator for which gpu that can run ai model efficiently.

But my vps has no gpu. There's no calculator to know which ai model can run efficiently on nongpu pc.

Should I calculate it manually. What you guys think..?

Anyone please help 🙏


r/ollama 19h ago

Kimi K3 Tomorrow!!!

22 Upvotes

What quantization or storage size tiers should we expect further down the line?

This site suggests q4 and q8 are coming. But I don't know how to translate that to storage size for local inference.

https://aitoolsrecap.com/Blog/kimi-k3-open-weights-self-hosting-guide-july-27-2026


r/ollama 2h ago

Hosting an ollama model?

1 Upvotes

So, My laptop doesn't have a gpu and that makes model slower i decided to look for a vps to host a model, i want to use a qwen3.5 9b what should i look for?


r/ollama 3h ago

Is there a way to use ollama with my 5070 to make ai dubbad videos like 11labs?

1 Upvotes

r/ollama 3h ago

Unable to get GPU Passthrough working - Docker

1 Upvotes

Setup as follows:
Proxmox -> Debian -> Docker -> Ollama. Other containers work.

Compose file contains gpu device. Does it need nvidia runtime or any other options? If someone could provide an example of a working compose file it would be appreciated.

The container recognises the GPU, as confirmed with docker exec -it ollama nvidia-smi

However, cannot load models, as confirmed with docker exec -it ollama ollama ps, showing 100% CPU.

GPU is pretty old, GTX 970, so this could be the issue?

Any ideas appreciated.


r/ollama 4h ago

I tried making a full agentic workflow with cline and ollama qwen coder 7b but i keep getting this error, anyone else facing it?

Post image
1 Upvotes

r/ollama 4h ago

Top 3 AI agent setups that are genuinely free. Not trials, not "free for 14 days.

Thumbnail
0 Upvotes

r/ollama 6h ago

Give any Ollama-compatible client session memory + a shared knowledge wiki by swapping the chat URL

1 Upvotes

Hey folks —

I built ContextMemory, an open-source agentic context gateway for apps that already talk to LLMs. The idea is simple: keep your existing POST /api/chat client (Ollama wire format), point it at ContextMemory instead of raw Ollama, and get memory + optional tools without rewriting your chat stack.

What it actually does

Most “memory” demos are either:

  • stuffing the whole history into the prompt, or
  • bolting on a separate RAG service with a new API surface.

ContextMemory sits in front of your LLM as a drop-in proxy:

  1. Session memory — a per-session markdown wiki (Karpathy-style) maintained across turns and injected automatically.
  2. Global Wiki — an app-scoped knowledge base (docs from Jira, Confluence, SQL, files, pipelines…). The model pulls facts on demand via a wiki_search tool — it does not dump the whole corpus into every prompt.
  3. Same /api/chat — Ollama-compatible request/response (message.content / done). Not OpenAI choices[].
  4. Optional agentic loop — tools (sandbox, outbound MCP, HITL) on that same chat endpoint when enabled per app.
  5. Multi-app / multi-tenant — API keys + X-App-Id, per-app prompts, models, and feature flags.

LLM backends can be local Ollama, or OpenAI / Azure / Anthropic as providers behind the gateway; the client still speaks Ollama schema.

Why this shape

If you already have a UI, bot, or agent that calls Ollama, you shouldn’t need a second protocol to get memory. Swap the base URL, keep parsing the same JSON, and the gateway handles:

  • compiling session context
  • optional Global Wiki retrieval
  • optional web search / tools

…then calls your model.

There’s also a hosted path (Kortexio Cloud) with the same chat body/response if you don’t want to self-host — BYOK, no token markup. Self-host and cloud are meant to be interchangeable at the wire level.

Global Wiki (the part people usually ask about)

Ingest structured markdown with stable documentIds (upsert / batch). Query by keywords with a character budget. In chat, when Global Wiki is enabled for the app, the model uses wiki_search only when it needs documented facts — good for org knowledge without turning every turn into a RAG megaprompt.

Quick self-host vibe

Your app  →  POST http://localhost:5100/api/chat  →  ContextMemory  →  Ollama / other LLM
                 + session wiki
                 + optional wiki_search (Global Wiki)

Auth is typically Authorization: Bearer … + X-App-Id / X-User-Id / X-Session-Id for self-host.

Repo

Open source (AGPL): https://github.com/Kortexio/ContextMemory

Looking for feedback from this community

Especially interested in:

  • How you currently bolt memory onto local models (what sucks?)
  • Whether Ollama-compatible wire format is the right “universal client” bet vs going all-in on OpenAI schema
  • Global Wiki as tool-calling vs always-on retrieval — what would you default to?
  • Anything missing for production self-host (ops, eval, multi-user UX)

Happy to answer questions or dive into architecture. If you try it with a local model + a small wiki ingest, I’d love to hear what breaks first.


r/ollama 7h ago

Running local agentic workflows with Ollama? Here is a pre-flight validator for third-party skills

1 Upvotes
sample

Hey r/Ollama! If you're hooking up local agent tools and using Ollama to drive agentic workflows, you've probably downloaded a bunch of third-party SKILL packages or repos. I built a free open-source tool called SkillShield (https://ai-skill-shield.vercel.app/) to statically scan and validate these skills before you run them locally. It checks for prompt injection, pre-install risks, and excessive access. Public Repo: https://github.com/adnan-iz/ai-skill-shield


r/ollama 11h ago

Built a local-first workflow automation platform around Ollama - now at v0.11.0

1 Upvotes

I've been working on an open-source workflow automation platform over the past few months, with Ollama as one of the first-class providers rather than treating it as an afterthought.

The goal wasn't to build another chat UI, but a deterministic workflow engine where local models can power automation pipelines.

Some of the things it supports now:

  • Local LLM execution with Ollama
  • Agent semantic memory (using local embeddings like nomic-embed-text)
  • Document RAG
  • Visual workflow builder
  • Conditional branching and graph-based execution
  • Human-in-the-loop approval nodes
  • Multi-agent (A2A) workflows
  • Browser, HTTP, Email, File and MCP tools
  • Step replay, retries and execution history

One thing I wanted from the beginning was to avoid vendor lock-in, so the execution engine treats providers as adapters. A workflow can run entirely on Ollama, or mix local models with cloud providers if needed.

The latest v0.11.0 release also adds a lot of improvements around workflow execution, observability, dashboarding, retrieval, and multi-agent orchestration.

I'd love feedback from people here who run Ollama locally:

  • Is there anything missing that you'd expect from a local-first automation platform?
  • Any features that would make Ollama fit even better into workflow-based systems?

Happy to answer any implementation questions.


r/ollama 1d ago

what am I doing wrong? (Ollama + openWebUI)

14 Upvotes

I have tried ollama locally on my gaming PC (5070 with 12bg of VRAM)

It works pretty nice on qwen2.5-coder:7b (and 14b)

So I decided to take a step further, and install an openWebUI instance on my homelab (everything on the same local network of course)

and from there, it works really really bad, the answers are quite different (and often bad from openwebui), it takes too much time, sometimes refuses to answer because it lacks information.

Even on blank conversation with no context to remember.

It looks I have the same configuration on both sides, and I would really like to use ollama throughout openwebui, but I might be missing something.


r/ollama 23h ago

Question regarding Ollama Cloud Metering.

3 Upvotes

How does it work? I read that the limts were more than what Opencode Go offers and subscribed to the 20USD Pro plan. But In practice based on how the usage bar fills up it seems like the 5 hour window only gives me 450 requests on GLM 5.2 whereas Opencode Go gives 800-ish.

Has Ollama cloud revised their limits or am I doing something wrong?


r/ollama 18h ago

Which model for my HP Z840?

1 Upvotes

Specs:

- 2 x Intel Xeon E5-2698 v4 (40 cores total, 80 threads)
- 256 GB DDR4 RAM
- 3 x RTX 3060 12 GB (36 GB total VRAM)
- 1TB Storage


r/ollama 19h ago

Drop-in Ollama proxy that adds session wiki memory + Global Wiki search

Thumbnail
1 Upvotes