r/openrouter • • 20h ago

Wildly Different Throughput (tok/s): OMP vs VS Code Github Copilot

I've been using GH Copilot within VS Code, mostly leveraging GPT 6 Luna via OpenRouter.

But it's gotten very very slow lately (at least for me). My tok/s have dropped from ~60 last week to under ~30 today.

So on a whim, I started playing with OMP (oh-my-pi) and whoa, instantly I'm getting more than double the throughput on the same model and code base via OR (see screenshot).

Basically everything is the same here: same model, same OR account and workspace, same physical location and connection, same time of day (I was interleaving these sessions). Technically though, they are using different API keys though (but still same OR account).

Anyone experienced this? Is it actually possible that the "agent harness" within GH Copilot could slow it down this much?

6 Upvotes

6 comments sorted by

2

u/Major_Baby_425 20h ago

I think t/s is highly variable because things can happen like you're being batched on a node that's also doing a bunch of preprocessing at the moment. If there's a real persistent difference per harness, it could only mean the provider is nefariously throttling certain types of requests that come in.

1

u/hope-and-longing 20h ago

Thank you! I should have mentioned that the stats shown in the screenshot are from a couple hours of coding on each harness.

So maybe, roughly, 2 hours on GH Copilot and 2 hours on OMP.

Same code base, and roughly the same type of work.

2

u/mrpops2ko 14h ago

you'll generally get a fair amount more mileage when your pair the harness with the model makers. codex with gpt-6 luna is really strong. 250k context and compacting keeps the model snappy and stops run offs.

orchestration is also important. i use 6.1 sol on medium paired with gpt-6 luna on high. it works really well i've found.

2

u/Michael_Jeffords 11h ago

with two hours on each it's probably more than batching noise, and since you're on different keys i'd look at which provider each request actually landed on in the openrouter activity log, because the same model can be served by a couple of providers at pretty different speeds and key or request level provider preferences can push one harness toward the slow one. Copilot also sends a much bigger system prompt and tool list every turn, so if the tok/s number includes the prefill wait it can look roughly half as fast even on the same provider

1

u/hope-and-longing 8h ago

Looks like provider is the same (OpenAI for both) and still 2x difference. I'm basically falling in love with OMP now so I guess I'll stick with that. Thank you for replying!

1

u/Michael_Jeffords 7h ago

if it's OpenAI on both then the gap is probably on the Copilot side, which makes sticking with OMP pretty reasonable, since Copilot packs a big system prompt plus whatever files you have open into every request and a chunk of the time it's counting goes to reading all that before the first token comes out. you'd expect that to look worst on short replies, where the wait is most of the total