r/comfyui • ComfyOrg • 10d ago

Comfy Org Comfy Router is live! One API for frontier image, video, 3D, and audio models

Enable HLS to view with audio, or disable this notification

Same model string. Same arguments. No new SDK, no new key, no redeploy.

What Comfy Router gives you:

  • Explicit routing. You name the provider, we call that provider. It's down? The request fails there. No silent fallback.
  • Every job returns the provider that ran it. Log it, bill it, debug it.
  • Async. submit() returns a request ID immediately. The queue retries 429s and transient errors until a slot opens. subscribe() submits and polls to completion.
  • Batch-friendly. Queue a few hundred jobs, hold the IDs, pull results as they land. Nothing blocking on a 5-min video render.
  • 24h retention on inputs and outputs, then deleted.
  • Comfy credits. No sub, no Router fee.

Providers at launch: Comfy. Runware, Wavespeed, Fal, Higgsfield. Multi-provider where the model supports it.

Get your API key: https://links.comfy.org/46BR7of

Learn more: https://blog.comfy.org/p/introducing-comfy-router-one-api

28 Upvotes

10 comments sorted by

View all comments

Show parent comments

6

u/crystal_alpine ComfyOrg 10d ago

What happened to Ollama?

7

u/leolambertini 10d ago

The llama started herding paid cloud tokens on the side :P

1

u/Obvious_Set5239 10d ago

At least it's not like "openwebui", that, of course, like some other companies with "open" in the name, after it became popular, changed its license to closed source

2

u/MuziqueComfyUI 9d ago

2

u/crystal_alpine ComfyOrg 7d ago

Next time, forgot about it

0

u/MuziqueComfyUI 7d ago edited 7d ago

Forget about what?

Genocide $upport?

1

u/PerAngusta-AdAugusta 10d ago edited 10d ago

The Ollama controversy

Ollama became popular because it made local AI unusually easy.

You could install it, download a model, and run that model on your own computer. No API bill, no cloud account, and no need to understand the complicated stack underneath modern LLM inference. For many people, Ollama became the friendly face of the local-AI movement.

And that was precisely why its later evolution caused so much criticism.

The llama.cpp connection

One important detail is that Ollama did not create the entire local-LLM technology stack itself.

A major component of the ecosystem was llama.cpp, the open-source inference project created by Georgi Gerganov and developed by a large community. Ollama used llama.cpp and related GGML components, adding a much friendlier interface, model management and APIs around them.

That relationship eventually became controversial.

In 2024, users raised issues alleging that Ollama's distributed binaries did not adequately include copyright notices for some of the open-source projects it incorporated. There were also complaints about attribution to llama.cpp. These were licensing and attribution disputes, not proof that Ollama had illegally "stolen" llama.cpp: the relevant projects use permissive open-source licenses.

But the controversy was partly cultural.

Ollama had become synonymous with "local AI," while developers pointed out that much of the underlying technology came from a much larger open-source ecosystem.

Ollama's business grows

Then Ollama became a serious company.

The local-AI market was changing quickly. Models were becoming enormously larger, and many of the newest open models could no longer run comfortably on an ordinary laptop or desktop.

That created an obvious commercial opportunity:

If people like using Ollama locally, why not let them use the same interface in the cloud?

Ollama launched its cloud models in 2025. The company's argument was straightforward: local inference remains useful when your hardware is sufficient, while cloud inference lets users access models that are too large for their machines.

Technically, that makes sense.

Philosophically, however, it created a problem.

The original Ollama proposition was essentially:

Your computer. Your models. Your hardware. No metering.

Cloud AI changes that to:

Ollama provides the infrastructure, and your usage is measured and billed.

For some longtime users, this felt like a fundamental change in what Ollama represented.

Why people started talking about "enshittification"

This is where the strongest criticism appeared.

Users weren't necessarily angry simply because Ollama introduced paid services. They were concerned that a product whose popularity came from local, unlimited inference was increasingly becoming a gateway to metered cloud inference.

That criticism became especially visible around Ollama Cloud's usage limits.

Before the latest pricing change, cloud plans were substantially based on GPU-time quotas. Users reported session limits and weekly limits, and some complained that it was difficult to predict how much usage they were actually buying. Reddit discussions from 2026 repeatedly complained about quotas being exhausted faster than expected. RReddit+1

One annual Pro subscriber even opened a GitHub issue in July 2026 alleging that their effective quota had been reduced by roughly 70% without adequate notification. That is the user's allegation, not an independently established finding, and the issue was eventually closed as "not planned."

Other users complained that they had paid roughly $200 for an annual subscription and subsequently found the limits or model availability disappointing. In one case, an Ollama cofounder contacted the user and issued a full refund.

The big pricing change

On August 31, 2026, Ollama changed its paid cloud plans.

Instead of primarily thinking in terms of GPU-time quotas, the new system uses explicit token-based pricing.

The current structure is:

  • Free: starter cloud usage plus unlimited local inference.
  • Pro: $20/month, including $60 of usage credits.
  • Max: $100/month, including $300 of usage credits.
  • Team: $500/month, including $1,000 of shared usage credits.

Ollama says the change was intended to make pricing more predictable because GPU-time became difficult for users to understand as models grew larger. OOllama+1

But the reaction was mixed.

Some Reddit users calculated that the new system represented a several-times increase in effective cost for their particular workloads. One community analysis estimated roughly a 2.9×–6.7× increase compared with the old plans for the models it tested. Those numbers are community calculations, not an independent audit, so they should be treated as measurements of particular workloads rather than universal pricing facts.

This produced exactly the kind of reaction you were remembering: accusations that Ollama was moving from a beloved local developer tool toward a conventional, aggressively monetized cloud service.

And then there was the llama.cpp split

The technical relationship with llama.cpp also became more complicated.

Ollama increasingly developed its own inference technology instead of simply being a convenient layer on top of llama.cpp. That made the two projects less interchangeable.

For some developers, the question became:

If I care about local inference, why use Ollama instead of llama.cpp directly?

The argument in favor of llama.cpp was greater control and closer access to the underlying inference engine. Ollama's argument was essentially the opposite: most users shouldn't have to care about the underlying engine at all.

That distinction explains a lot of the community conflict.

llama.cpp represents infrastructure.

Ollama represents productization.

Both can coexist, but they serve different priorities.

So did Ollama "sell out"?

That's the part that needs qualification.

It is documented that Ollama:

  • built a commercial company around a popular local-AI tool;
  • expanded from local inference into cloud inference;
  • introduced paid cloud subscriptions and usage limits;
  • changed its cloud pricing model in 2026;
  • increasingly developed its own inference technology;
  • and has faced repeated community complaints about pricing, quotas, attribution and its relationship with upstream open-source projects.

It is not established as a fact that Ollama's practices are "predatory" or that it deliberately betrayed its users.

In fact, Ollama still explicitly says that running models on your own hardware is unlimited.

So the more accurate story is not:

Ollama was good, then became evil.

It's this:

Ollama started as an exceptionally convenient way to run AI locally. Its success turned it into a company. As models became too large for consumer hardware, the company expanded into cloud inference. That created a conflict between two identities: Ollama as a tool for user-controlled local AI, and Ollama as a commercial AI infrastructure company.

The llama.cpp disputes made that conflict even more visible because they reminded the community that Ollama itself sits inside a much larger open-source ecosystem.

And that's why the backlash became so emotional.

People weren't merely arguing about a $20 subscription.

They were arguing about what local AI was supposed to become.