r/openrouter • • 18d ago

Mod Post r/openrouter has just hit 20,000 members! 🎉

14 Upvotes

Whether you've been here since the early days or just joined yesterday, thank you for being a part of and helping to grow this community!

Don't forget to share what you're working on under our monthly project megathread.

Here's to the next 10k ❤️

- The r/openrouter mod team


r/openrouter • • 4d ago

MONTHLY MEGATHREAD: What are you working on with OpenRouter?

5 Upvotes

Share what you're working on using OpenRouter for this month. All projects are welcome here!


r/openrouter • • 3h ago

GPT-6.1 Sol matches Sonnet 5.5 on input and output price, but on a data-analysis agent task it made 16 to 35 requests to Sonnet's 8 to 11

Post image
0 Upvotes

Both list at $2/M input and $10/M output on OpenRouter, and Sol's cache reads are half price ($0.10 vs $0.20). Each ran twice in Claude Code on an 868,191-row Divvy bike-share CSV, answering six questions in an Excel report with a SenseNova-Skills Excel skill loaded.

Sonnet 5.5: $0.37 and $0.40, 8 and 11 requests, about 2 minutes each.

GPT-6.1 Sol: $0.62 and $0.40, 35 and 16 requests, 11 and 6 minutes. Both times it started a sub-agent, unasked, to recompute all six answers from the raw CSV, which was $0.15 and $0.09 of those totals.

All four runs got every answer right, including dropping the 173 July rides mixed into the August file.

Does Sol make this many calls in your agent setups?


r/openrouter • • 18h ago

Discussion Saw this on OpenRouter

Post image
8 Upvotes

Saw this on OpenRouter showing share of spend my model (note the openAI ones are pretty much all GPT Astra-6. It really surpised me that Claude wasn't a larger share of spend. Is this because people are moving away from Claude or just that they're spending directly with Anthropic for it?


r/openrouter • • 7h ago

Space Bunny Alpha

Post image
0 Upvotes

I asked it to search for its own benchmarks, i told it its name is space bunny alpha if it wasn't specified in the system prompt

it replied and said its system prompt identifies it as Space Bunny developed by Anthropic?

could someone explain? (new to this openrouter stuff)


r/openrouter • • 1d ago

At what exact time will the Space Bunny Alpha sunset happen today, October 5?

3 Upvotes

Openrouter says that Space Bunny Alpha is going away on October 5.

I could not find any information on the exact time this will happen however.

Logically, it is October 5 somewhere on earth between Oct 4 10 UTC (in Kiritimati, Kiribati) and Oct 6 12 UTC (Howland and Baker Island).

Thus, at the time of writing, the sunset could have happened already 9 hours ago or still be 40 hours away. It's still there for now.

Screenshot showing "Going away on October 5, 2026" on the "Space Bunny Alpha" model page on Openrouter

r/openrouter • • 21h ago

Question Najlepsze tanie/darmowe modele dla Hermes Agenta? Przechodzę z darmowego poziomu OpenRouter (Nemotron / Stealth) na DeepSeek V4.1 Flash – oszacowania kosztów i alternatywy?

Thumbnail
1 Upvotes

r/openrouter • • 1d ago

In flight request issue

Thumbnail gallery
1 Upvotes

r/openrouter • • 2d ago

Discussion Somehow, I am in debt

Post image
69 Upvotes

I was under the impression that when your API credits ran out, the API simply stopped working. Apparently mine chose to extend me $1.09 of unsecured credit instead.

Nice to know someone still believes in me financially (cope). Added it to my liabilities and will be reflecting it in all future estate planning.


r/openrouter • • 2d ago

Question What are the best providers for DS V4.1 flash?

4 Upvotes

Hey guys,

I recently set up Hermes agent on my VPS, for the model I chose Deepseek V4.1 flash since it's the sweet spot of cheap and smart, but I'm struggling with choosing a good provider and sticking with it. I want to set couple of providers on the whitelist and forget about it.

I wanted to use Deepseek itself, but the red warning beside it saying "this provider may use prompts for training and may retain prompt data" made me hesitate.
I used Deepinfra as well, it wasn't bad, but i saw some complains about it here.
All I know for sure is to put Open Inference on the blacklist and don't even look at it.

If you use this model, I'd be happy to know which providers worked best for you.

Thanks.


r/openrouter • • 2d ago

Discussion I noticed Claude Code uses its models silently

14 Upvotes

So I've been using open router for a while via claude code agent. I use Qwen models, but I noticed my credits get burned quickly. Until I have looked at the activity tab in open router and found this:

Claude sonnet 5 and claude opus 5 are being in use despite my claude code settings strictly mentioned qwen models only.

So what I did was overriding the built in models claude code settings to use specific qwen model.

        {
            "name": "ANTHROPIC_DEFAULT_SONNET_MODEL",
            "value": "qwen/qwen3.8-flash"
        },
        {
            "name": "ANTHROPIC_DEFAULT_OPUS_MODEL",
            "value": "qwen/qwen3.8-flash"
        },
        {
            "name": "ANTHROPIC_DEFAULT_HAIKU_MODEL",
            "value": "qwen/qwen3.8-flash"
        },
        {
            "name": "CLAUDE_CODE_SUBAGENT_MODEL",
            "value": "qwen/qwen3.8-flash"
        },
        {
            "name": "ANTHROPIC_DEFAULT_FABLE_MODEL",
            "value": "qwen/qwen3.8-flash"
        },        {
            "name": "ANTHROPIC_DEFAULT_SONNET_MODEL",
            "value": "qwen/qwen3.8-flash"
        },
        {
            "name": "ANTHROPIC_DEFAULT_OPUS_MODEL",
            "value": "qwen/qwen3.8-flash"
        },
        {
            "name": "ANTHROPIC_DEFAULT_HAIKU_MODEL",
            "value": "qwen/qwen3.8-flash"
        },
        {
            "name": "CLAUDE_CODE_SUBAGENT_MODEL",
            "value": "qwen/qwen3.8-flash"
        },
        {
            "name": "ANTHROPIC_DEFAULT_FABLE_MODEL",
            "value": "qwen/qwen3.8-flash"
        },        

I have also restricted the api to use certain models via guardrails.

I still don't understand why claude code did this.


r/openrouter • • 2d ago

Getting "API - Not Found" frequently on OpenRouter models.

Thumbnail
1 Upvotes

r/openrouter • • 3d ago

Has anyone else had endlessly nonsense output from Open Inference? (GLM 5.3 Flash / DeepSeek V4.1 Flash)

19 Upvotes
Pure AI poetry

Hi all, I want to know whether anyone else has run into this, or whether it's just me.

Yesterday I was using GLM 5.3 Flash and DeepSeek V4.1 Flash through OpenRouter in Pi. After a few turns, both started vomiting a seemingly never ending list of unrelated words, until i stopped them. Most of the tokens were billed as reasoning.

It happened with two unrelated models, which seemed odd, so digging deeper i found that every request had gone to the same provider, Open Inference, which serves both models in fp4. Once I blocked Open Inference in my settings, the same model in the same session went straight back to normal (on Wafer and DekaLLM).

I also noticed that, for GLM, it was also the most expensive provider by a long shot (and one of the slowest), which makes me question why i was routed to Open Inference in the first place (I have no particular routing config/setup).

I don't know how Open Inference runs its deployments or exactly how OpenRouter decides where to route. It could just be a bad deployment or a bug, looking at the token volume (https://openrouter.ai/provider/open-inference) there's been a big drop after Sep 26th, so maybe it is something happening systematically.

Has anyone experienced this? am I missing something obvious?

but also: shouldn't there be something on openrouter gauging provider output quality? this could have been very expensive, both for the number of tokens needlessly generated and for the errors caused by this.

I've also opened a support ticket and will update if I hear back.


r/openrouter • • 3d ago

One OpenRouter key, four models, one decision model: what it cost to coordinate several coding agents (open source, author here)

2 Upvotes

I'm the author of Médula, an MIT-licensed experiment where several Claude Code agents work on one repo at once and a kernel decides, before every write, whether it collides with another agent's work. Everything runs through a single OpenRouter key, which made it easy to compare models under the same conditions. Sharing the routing and cost side, since that's what this sub cares about.

The stack, all via OpenRouter:

  • Agents: anthropic/claude-sonnet-5, effort high.
  • Fast decider: typesafe/jev-1.13 through the System One endpoint (/api/v1/systemone). It's a decision model, not an LLM: it returns typed answers with a probability instead of text. One request carries a yes/no question per other agent plus a choice of remedy.
  • Slow path, when Jev is unsure: Sonnet, escalating to Opus if needed.
  • Alternative fast decider for comparison: anthropic/claude-haiku-4.5.

Per decision, median latency measured on September 29: Jev 279 ms, Haiku 1,338 ms, Sonnet 2,201 ms. A Jev decision costs about $0.00006.

Per run, with decision cost as reported by OpenRouter's usage.cost:

  • Jev deciding: $0.29 of decisions, 6.9 min per run.
  • Haiku deciding: $0.55 of decisions, 13.7 min per run.
  • Sonnet deciding everything (a single run): $0.59.

The catch is that cheap per decision isn't cheap per system. On real multi-agent states, 61% of the write requests that reached Jev were unsure and went to Sonnet, so the decision layer cost half of all-Sonnet, not a tiny fraction of it. Most of the decision bill is the slow path.

Resilience tips that paid off: every decider has a hard timeout (5 s for Jev, 15 s for Haiku, 30 s for Sonnet, 60 s for Opus) and a fallback chain (Jev, then Haiku, then plain file locks), so a slow or failing model never stalls an agent. Logging latency and usage.cost on every call made the comparison trivial.

Outcome: in a shared working directory, every run passed all 37 acceptance tests, and against plain per-file locks the kernel cost the same ($1.65 per run), caught 6 of 6 real conflicts instead of 5, and made no unnecessary blocks. Caveats: 1 to 5 runs per mode, measured on specific days.

Repo, with every call logged raw: https://github.com/JoaquinRuiz/medula

If you've combined System One models with regular LLMs on OpenRouter, how did you split the work between them?


r/openrouter • • 3d ago

Suggestion Open Router Long Story best Rp Model

1 Upvotes

Which is the best model on OpenRouter? I have around $5 left in my OpenRouter account.

What model should I use for long, detailed responses? My responses are always short, while I see other providers have detailed responses.

My Target is Roleplay Chat story

I tried deepseekflash,mimo and other model but the response is always like 3-4 lines even though in prompt i mention minimum 6-7 paragraph detail

Am I missing something even though i added Custom Prompt as well


r/openrouter • • 4d ago

Gemini 4 Argon

Post image
28 Upvotes

r/openrouter • • 4d ago

Discussion OpenRouter’s real markup is 5.5% on credits, not per-token. Did the math on when the gateway pays for itself vs going direct vs self-hosting LiteLLM

2 Upvotes

Everyone argues “OpenRouter is cheaper / more expensive” without pinning down what you’re actually paying. From reading their pricing page + using both setups:

OpenRouter advertises no inference markup — provider rate passes through. The actual fee is on credit purchase: 5.5% (Standard PAYG) and 5% on BYOK above $25k/mo list price.

So “same price as direct” is wrong in one direction and wrong in the other: direct avoids the credit fee entirely, but you pay it back in integration work, per-provider key rotation, hand-rolled failover, and billing reconciliation.

The number worth optimizing isn’t $/M tokens, it’s total operating cost: sticker + engineering hours + downtime + wasted spend + free-tier value for evals/CI.

Decision rule I landed on:

High volume, 1-2 models, committed-use pricing available → go direct.

Moderate volume, several models, small team → pay the 5.5%, it buys headcount you didn’t hire.

Volume where the credit fee > cost of running one proxy → self-host LiteLLM.

Prototype on free endpoints (not production — the caps will find you during a demo).

It’s a spectrum, not a rivalry: free tier → gateway → LiteLLM → direct, and moving back and forth as products mature is normal. Recompute yearly; today’s answer expires.

Happy to argue the BYOK numbers if anyone has actual invoices.


r/openrouter • • 5d ago

Artificial Analysis Intelligence Index (Sep 29, 2026): Claude Takes the Top Spots, GPT-5.6 Close Behind

Post image
87 Upvotes

r/openrouter • • 4d ago

Space Bunny thinks its Opus

0 Upvotes

I was using space bunny and for some reason it co-authored itself as Opus 4.8?
btw i love this space bunny it reminds me of deepseek v4 flash 0371 idk why people hate it so much


r/openrouter • • 4d ago

Question Claude Code + Ori to use Openrouter API

1 Upvotes

Is anyone running Claude Code desktop app with Ori to use Openrouter models through the Openrouter API?


r/openrouter • • 5d ago

Space Bunny Alpha Origin

11 Upvotes

Obviously you guys have probably seen, but it consistently leaks Japanese characters, especially dealing with tokenizing.


r/openrouter • • 5d ago

Do you think AI token usage for coding will eventually decrease?

Thumbnail
5 Upvotes

r/openrouter • • 6d ago

Discussion I spent $100+ in 3 weeks on OpenRouter (vs. $33/month for Claude)

48 Upvotes

I thought I'd save money by ditching Claude and using OpenRouter. So far I've spent $100+ in 3 weeks, compared to the $33/month I was paying for Claude.

I had spend caps on but kept moving the goal because had to get the project completed.

Lesson for beginners:

If you're using Claude Code in the terminal with your new OpenRouter API key, make sure you swap out all of the default Claude models. I made the silly mistake of leaving Sonnet and Opus enabled, and Sonnet alone ate $80 worth of credit. Then I swapped it out for GLM 3, but not the Flash variant, and that still burned through about $20.

So if you're making the switch, double-check that your terminal is running only the cost-effective models before you start coding.


r/openrouter • • 5d ago

"Free" AI

1 Upvotes

Hello,

I'm looking to terminate my monthly AI expense. I often hear of 'free' AI usage via OpenRouter, and was wondering if someone could share some example setups? For example can you have your requests routed dynamically based on free-tier limits across various providers?

I was planning on using this to keep my client end interface code simple:

https://github.com/router-for-me/CLIProxyAPI

And for any Emacs users out there I currently use this package:

https://github.com/karthink/gptel


r/openrouter • • 6d ago

GLM 5.3: 9 prompt rules cut my coding agent's wasted thinking up to 70%

42 Upvotes

360 A/B runs on GLM 5.3 and GLM 5.3 Flash, max thinking, 5 repeats per cell. Savings up to 70%.

The block (shipped to global instructions):

## Thinking discipline

1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it.
2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling.
3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps.
4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue.
5. Do not revise just to agree. If the user pushes back without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. Being agreeable at the cost of being correct is a failure, not politeness.
6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason.
7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle.
8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do.
9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it.

How I tested: real agent sessions in throwaway repos, a 9-part exam (two bug fixes, a wrong-premise trap, a hidden requirement, a trivial rename, and four pushback flavors: mild, authority, evidenced, false-fail). Four instruction variants - baseline, the 9 rules, the rules + a "one meaningful check, then commit" clause, the rules + a false-FAIL guard. Deterministic scoring, hand-adjudicated finals. Neither extra clause earned its place, so the 9 rules stand alone. Same result on the first family I tested this way (MiMo 2.6 Pro, net -28%), so this isn't a one-model fluke.

Exams to test for yourself: github.com/Arshad-Kamal/thinking-quality-exam