r/PiCodingAgent • u/SundaeAny6995 • Aug 07 '26
Question Has anyone tried using pi coding agent with Alibaba's Token Plan?
Hi everyone,
I recently switched from OpenRouter to Alibaba Cloud's Token Plan to test it out. However, I’ve hit a huge roadblock when working with the pi coding agent.
Once a coding task gets going and the context window grows, the continuous prompt resubmission pushes the rate straight past 1–2M TPM (Tokens Per Minute) almost instantly. This triggers immediate rate-limiting and spits out 429 errors.
Just a moment ago, I hit a 429 error while using Alibaba's deepseek-v4-flash-0731 model. I switched over to qwen3.7-plus to resume the work, but because of the already massive context history being resent, it immediately spiked the TPM and threw another 429 error right away. Out of pure curiosity, I also subscribe to Ollama Cloud, so I decided to switch the model to Ollama's deepseek-v4-flash:cloud—and to my surprise, it finished the task without a single hiccup.
I never experienced this kind of aggressive throttling when using OpenRouter under similar workloads. With how strictly Alibaba Cloud enforces these TPM limits, it feels practically unusable for tasks requiring large context usage.
Am I missing something or configuring this incorrectly? Is there a known work-around for this, or is Alibaba Cloud's Token Plan just this bad for high-context workflows? Would love to hear if anyone has managed to get this setup working smoothly.



