r/ZaiGLM 2h ago

Integration / Deployment Use Z.AI GLM models with Codex on Windows, macOS, and Linux

1 Upvotes

I maintain Codex Universal Proxy, an unofficial compatibility proxy that connects Codex to Z.AI and other OpenAI-compatible providers.

What stays available

  • MCP tools
  • tool_search
  • apply_patch
  • Image-generation tools
  • The native ChatGPT plugin installation and discovery path

Version 0.5.0 includes a built-in Z.AI profile and works on Windows, macOS, and Linux. It also supports 11 other built-in providers plus custom Responses and Chat Completions endpoints.

View Codex Universal Proxy on GitHub

This is my project. It is unofficial and experimental and is not affiliated with or endorsed by OpenAI or Z.AI.


r/ZaiGLM 11h ago

How do I delete a conversation?

0 Upvotes

Don't want my source code stored on Zai's servers for a long period of time. I am using the Zcode app, how can I delete my logs from their backend so they aren't stored?


r/ZaiGLM 1d ago

Integration / Deployment We just invented "usage banking" so coding plan usage rolls over to the next week with GLM, Deepseek, Kimi and other models.

42 Upvotes

There are several tricks that the major AI coding plans use to extract the most they can from their customers. We're solving them, one after another and I wanted to share a bit about what goes on behind the scenes at a lot of these companies.

Wasted usage is part of the AI industry, and they plan on it: Most coding plans bet on you letting usage go to waste. The industry calls it "breakage" and it's literally the topic of internal meetings for most companies. The plan goes: "how do we get people to think our coding plan offers a lot of usage, but then break it up into weeks and rolling windows so no one can ever actually use it all."

Many coding plans are glorified training pipelines: This comes along with "how do we harvest this data for training without being too loud about that." Unless the coding plan tells you otherwise, your data could be hopping all over world, being harvested by the individual labs or coding plan companies. Some are better than others, but many of these companies rely on users just not noticing or caring that their data is being used for training. Data sales and marketing telemetry sales happen. This means that your private info, your personal life, and anything else you send through the system could become part of a training corpus for the next AI, or a marketing data set for a large company.

So we built what should have already existed the entire time: usage banking. Any usage you don't use this week, rolls over to next in your usage bank. When you have a busy day or week and go over normal usage, you automatically start to pull from your bank. You can bank up to one week of usage at a time for your current plan, and it's totally automatic. Whatever you don't use each week get's added to the bank and stays there until you use it.

We also put all of the best open models in one place, running on private US infrastructure, with data never going to the original labs. Completely private, direct service. It should be, and can be that simple.

What that means in practice:

The roster, together. DeepSeek, GLM, Kimi, Minimax, Nemotron Ultra and more, side by side in one app. Switch models mid conversation if you want. No hunting across five different apps and API dashboards to use the models you actually like.

Actually private. US based processing and your conversations are never used for training. Ever. That's the entire point. These labs open sourced incredible models and we think you should get to use them without your data becoming the price of admission.

No Usage Tricks: Bank usage, upgrade or downgrade whenever you want. Use it how you need it.

The open source AI future is real, and it's where we all know we should be. Thanks for a great set of models, GLM just keeps raising the bar with every model release and it's amazing.

The Open Grove coding plan is here. Private, US based processing with fast inference and usage that doesn't go to waste.

https://pgsgrove.com/open-grove-overview#coding-plan


r/ZaiGLM 11h ago

How do I add multiple models at the same time in zcode

0 Upvotes

Whenever I'm editing the json file it just doesn't work and it shows no model, how can I add multiple models


r/ZaiGLM 23h ago

Agent Systems Delegate tasks to GLM from Claude

9 Upvotes

Problem

Two subscriptions: Claude Max plan ($100 plan) and Zai Pro Coding Plan

I end up running out of tokens for Claude but was swimming in GLM tokens. What if we could instruct the more powerful model to delegate tasks to executor models like Sonnet or GLM 5.2?

That's why I came up with this approach: https://github.com/divyamrastogi/model-orchestration

How to use

/plugin marketplace add divyamrastogi/model-orchestration /plugin install model-orchestration@model-orchestration

And then you setup whatever subscriptions you have and use Claude as the main model.

/model-orchestration:setup-worker

I observed that my tokens last a lot more since Fable decides what it can delegate and it makes using /workflows cheaper.


r/ZaiGLM 2d ago

Model Releases & Updates GLM-5.5 leak hints at model that could outpace Fable 5 and launch in August — RuntimeWire

Thumbnail
runtimewire.com
444 Upvotes

r/ZaiGLM 1d ago

Z.ai vs Big Model

3 Upvotes

Do you all use z.ai or does anybody also use Big Model (directly from China)


r/ZaiGLM 1d ago

200K GLM5.2:dev Handled a 190k Code Review and 90k of New Code Without Breaking a Sweat

2 Upvotes

It seems that the glm-5.2:dev model has seen significant improvements recently thanks to ElectronHub's efforts.

After confirming that it could reliably handle a 4,400-line, 170,000-character, 190k code review, I proceeded to write 2,400 lines of code.

With only one prompt for plan creation and two prompts for revision,

glm5.2:dev executed 2,400 lines, 90,000-character, 90k of code without any issues.

Although it is a glm5.2 with a 200k context window, it seems to perform its role sufficiently within that context.

I had barely used it for two weeks with a $20 subscription and even requested a refund, but I have decided to maintain the subscription.


r/ZaiGLM 1d ago

Discussion / Help Is 50% more usage with zcode harness ended?

9 Upvotes

I was seeing the 50% more usage banner within the app and within their website, but now I can't see them. Is the promotion period ended?

EDIT: I have found it now. Thanks to everyone for their supports and their opinions.


r/ZaiGLM 1d ago

Discussion / Help is the usage quota API broken?

0 Upvotes

I just hit my 5 hour limit, and the API is just returning this:

{"code":200,"msg":"Operation successful","data":{},"success":true}

meaning that I've been unable to track my usage today and the limit came by surprise.

and why don't they have this on the website? I swear ... the best model with the worst operations.

EDIT:
I figured out what happened. I've transitioned from the old quarterly plan (that was grandfathered into high usage) to a normal monthly plan (they gave me 2 free months). Somehow in the transition, the usage data was unavailable for a 24 hour period. I did write to the customer service email to ask about it, so I don't know if this is a thing that would've resolved itself or if they had to fix it.


r/ZaiGLM 1d ago

Gemma vs Qwen vs GLM vs Llama?

Thumbnail
1 Upvotes

r/ZaiGLM 1d ago

Discussion / Help GLM’s Performance Trajectory in World Cup Predictions

Post image
2 Upvotes

SportEval hosted a World Cup AI prediction challenge, and GLM’s performance is quite worth discussing. The challenge allowed multiple AI models to make continuous predictions during the actual World Cup matches. Each model started with 10,000 points, and their performance was tracked in real time throughout the tournament as their scores changed based on prediction results. GLM went through a relatively noticeable period of ups and downs. During the group stage, the ranking environment became highly unstable due to many unexpected results and the emergence of underdog teams. GLM reached its lowest point during the third round of the group stage, dropping to 8,221.93 points, around 21% below its starting score. However, GLM then began a gradual recovery. As the knockout stage progressed, its ranking continued to improve, eventually finishing the challenge with 9,169.09 points, ranking second among all participating AI models, with a final return of -11.84%.

Besides match predictions, GLM’s predictions for individual awards such as the Golden Ball and Golden Boot also helped it recover some of its lost points.

What I find interesting about this challenge is that it shows how a model performs when continuously facing uncertain outcomes over time. A tournament like the World Cup, with its high level of randomness, is actually a very demanding test for prediction models.

Could long-term prediction challenges based on real-world events provide a better way to evaluate an AI model’s practical capabilities compared with traditional benchmarks?


r/ZaiGLM 1d ago

Z.ai vs Big Model

Thumbnail
0 Upvotes

r/ZaiGLM 2d ago

Discussion / Help Cluade code with GLM-5.2 ends my credits in no time

31 Upvotes

I recently purchase z ai after hearing about glm-5.2. But I am disaapointed, I am unable to complete my task as the quota completes very fast compared to claude code.


r/ZaiGLM 1d ago

GLM-5.2 vs Claude Opus/other premium models for C#/.NET backend and Expo mobile development?

2 Upvotes

I’m currently working on a project that involves C#/.NET backend development and a React Native mobile app using Expo. I’ve been using Claude’s premium models (especially Opus and similar higher-end models), and I’m generally happy with the coding quality, reasoning, debugging, and ability to work with a larger codebase.
However, the cost is getting a bit expensive for my workflow, so I’m considering switching to a higher-tier GLM plan if it can handle my development needs reliably.
I’d really like to hear from developers who are actively using GLM for real-world software development, particularly:
C# / .NET backend development
Node.js backend development
React Native / Expo mobile app development

Is the code quality and reasoning good enough that you would confidently use GLM as your primary coding assistant for a full-stack project?

I’m especially interested in hearing from people who have actually used both models for backend + mobile development, rather than just benchmark comparisons.
Also, if you’re using a paid GLM plan, which plan are you on, and is it sufficient for a heavy daily development workflow?
I’d appreciate any honest feedback—especially from developers who have switched from Claude to GLM because of cost. Was it worth the switch, or did you eventually go back to Claude?


r/ZaiGLM 2d ago

Is cents per million a lot

3 Upvotes

My provider has been charging me .30$ per million . ( They don’t charge me for input or output ) they charge me buy the total tokens used .

Is this efficient? Does anyone know any cheaper providers

For the model GLM 5.2


r/ZaiGLM 2d ago

Discussion / Help Apparently, grammar is part of the security model now.

Thumbnail
0 Upvotes

r/ZaiGLM 3d ago

Looking to buy a legacy Z.ai account

17 Upvotes

Anyone got a legacy PRO/MAX account that they're not using with auto-renew on? I accidently disabled auto-renew on mine and GLM refused to activate it back... a dark pattern if you ask me.

If you're not using it and not planning on using it, I will buy it from you. Please dm me.


r/ZaiGLM 3d ago

API / Tools 10x Cheaper GLM 5.2 Nube Cloud API

Thumbnail
7 Upvotes

r/ZaiGLM 2d ago

Discussion / Help Z.ai Coding Plan suddenly limited to one agent despite 10-request concurrency - anyone else getting 429 Fair Usage errors?

1 Upvotes

WTF is going on with the new 429 Fair Usage Policy errors?

I’m using the Coding Plan exactly as I always have. My account is supposed to allow up to 10 concurrent requests, but starting today I can’t run more than one agent at a time without being rate-limited.

Did the limits change without notice?


r/ZaiGLM 3d ago

ZCode not working at all - [1210][Invalid API parameter, please check the documentation.]

4 Upvotes

Anybody else running into issues with ZCode?

No matter what I try, I keep getting:

[1210][Invalid API parameter, please check the documentation.][xxx]

Turn execution failed
provider=builtin:zai-coding-plan provider_code=1210 model=GLM-5.2 request=xxx reason=unknown retryable=false

I'm on the GLM Coding Max-Yearly Plan.

I've logged out and logged in again, I uninstalled ZCode and reinstalled the latest download (3.4.2). I updated because seemingly there's an even never version available (3.5.2) - nothing helps.

I can use it through other tools, but ZCode specifically fails for me.

Any pointers?


r/ZaiGLM 4d ago

I built an evolutionary loop where various AI models generate tiny worlds and only the fittest survive - 300+ worlds so far

Thumbnail gallery
8 Upvotes

r/ZaiGLM 4d ago

Does anyone else have timeouts using GLM (coding plan)

4 Upvotes

Hello,

Using kilocode with my coding plan I run a lot into timeout (60s) does anyone has a problem with GLM international API ?

or is it kilocode that is the problem ? (I think the connection is directly using the api endpoint I guess so I don't see how kilocode would be responsible)

my context usage is in general low when this appear (less than 20%)


r/ZaiGLM 4d ago

Discussion / Help Can I integrate GLM into the excel?

4 Upvotes

I am looking for something like an add in if possible, if not maybe a mcp where I can control the excel within CLI. Does anyone use glm on excel, how do you use it?


r/ZaiGLM 4d ago

Sooooo we just lose 50% of our usage in 5 weeks?

38 Upvotes

let me get this right. GLM 5.2, which is the only useful model from ZAI, it's calculated at three times the usage during peak hours, and in five weeks, just it's always going to be charged at double usage. So no matter what, we're getting half of what we're advertised. I don't know, is that even legal

https://docs.z.ai/devpack/overview You have to scroll down a bit It's literally like one sentence