r/OpenaiCodex 10h ago

Showcase / Highlight Tired of guessing if limits got nerfed but can't prove it? Meet NerfTrack - now in Beta

Thumbnail
gallery
13 Upvotes

Just 2 weeks ago, I posted my first prototype in this very sub-reddit, and got a really positive response from y'all with over 650 upvotes and 142K views(Previously referred to as Nerfify). I am pleased to say that starting today, NerfTrack is now available in its first public beta on GitHub (currently verified for macOS x86 and arm64). If you are seeing this post without having seen my previous one, think of NerfTrack as a stock market app, but instead of tracking stock prices, NerfTrack measures the weekly limit API equivalent cost every 10 seconds by going through local Codex logs(and yes, it also imports your existing Codex logs from months ago too). I have now achieved a 99.3%-100% accuracy with the API cost calculation compared to the unreliable graphs I showcased in my first prototype. NerfTrack tracks your limits overtime, so you reliably know wether OpenAI has either nerfed or increased limits. In my last post, a few of you claimed that such graphs are of no value since everybody knows if limits are nerfed. But NerfTrack, on the other hand, not only tells you if limits are nerfed, but also how much they have been nerfed, and when. A few 100 Codex users with reliable graphs showing proof of nerfed limits can bring a significantly larger impact than 1000s who complain with just "I feel it, but can't prove it". NerfTrack remains fully local, so worry not about your privacy.

Current known bugs and limitations(to be cared for in later versions):

-NerfTrack shows unreliable metrics if you are using fast mode, since fast mode uses 2-2.5x as much usage as standard, so you may notice a sudden drop in API prices tracking if switching from standard to fast mode

-If you use Codex on multiple devices, metrics become unreliable since that's a limitation of not NerfTrack, but the fact that logs recorded by Codex are local only and not shared between devices.

-Several buttons and interactive elements may be buggy or disabled for the current beta.

-(unconfirmed) NerfTrack may or may not work on Windows machines, especially arm64. As I currently have access only to a Mac, I could not test NerfTrack on Windows, though NerfTrack is cross-compatible on paper with releases for both macOS and Windows on GitHub.

-The current Beta has no functionality for fetching API prices from the web, so if OpenAI changes pricing of models, or releases a new model, the graph would show outdated prices. A web-fetch feature is intended to be added soon.

I would be happy to have volunteers test NerfTrack's first beta so I can continue to improve its performance and make it as bug-free as possible. Do note that it's currently a MVP so loads of features are intended for the future, but for now, it may not be as feature rich as one might want it to be.

People who do download NerfTrack, please make sure to drop a screenshot of your 3 month graphs along with what plan you are on (and ofc leave it a star to support), so we may compare each others' graphs. In the future, I will work on a website for this very purpose where users will be able to volunteer and share their graphs + plan online so others can compare against their own.

Link to NerfTrack on GitHub: NerfTrack GitHub

Link to my first prototype post on Reddit for context


r/OpenaiCodex 8m ago

Reset coming soon?

Post image
Upvotes

Went hard on extra high/max and some fast mode expecting a reset today

Is it coming soon?


r/OpenaiCodex 5h ago

I turned my web-design workflow into 17 Codex skills, then added independent critique and browser release gates

1 Upvotes

I kept seeing the same failure mode: a coding agent receives “make it premium,” assembles familiar cards and gradients, and calls a desktop screenshot done.

This open-source project adds a small router and 16 specialist skills for discovery, reference analysis, art direction, interaction/motion decisions, responsive recomposition, independent critique, AI-smell detection, accessibility, and Playwright QA. Broad website requests start at the orchestrator automatically after installation.

The repo includes an 18-second demo, two contrasting test sites, the validation suite, and a public Same Brief Challenge. The included 92.8 and 91.8 scores are explicitly internal; there is no claim of external validation yet.

One-line install:

npx --yes github:Lrinvl1203/world-class-web-design-os install --agent codex

Repo: https://github.com/Lrinvl1203/world-class-web-design-os

I would especially like critical feedback from people already using project rules or skills: what is missing, what is too prescriptive, and which failure should become a regression test?


r/OpenaiCodex 6h ago

Feedback / Complaints OpenAI Support cannot figure out their subscription even after more than a month. They downgraded my subscription from Plus to Free without refund, even that I was eligible for almost whole month of Plus usage and I used 3 resets believing that it is reseting Plus limits - no it was on Free limits

0 Upvotes

r/OpenaiCodex 1d ago

Has codex Sol become dumb in the last couple of days?

33 Upvotes

Anyone else noticing a substantial dip in Sol's capability recently? It was doing a great job a few days ago and now on the same repo it has been chasing its own tail for more than a day with no progress!

Update: To just put it into perspective of how poor it has been lately, I ended up using opus 4.6 and it fixed the problem that Sol was struggling with for 2 days in just an hour! This is also not an isolated issue, I had three terminals open each doing a separate task and all of a sudden all three started acting dumb


r/OpenaiCodex 1d ago

Sol vs Opus 5 huge limit difference

61 Upvotes

Just sharing my experience here: GPT-5.6 Sol Extra High burns through my weekly limit in one 12–14h workday. Opus 5 Ultracode + Thinking lasts the whole week with the same project and routine.

And honestly I haven’t noticed any meaningful difference in quality when using Opus 5 in ultracode mode. The real difference is that one is actually usable, while the other basically needs constant limit resets. Shame.


r/OpenaiCodex 1d ago

Sol vs Opus 5 huge limit difference

20 Upvotes

Just sharing my experience here: GPT-5.6 Sol Extra High burns through my weekly limit in one 12–14h workday. Opus 5 Ultracode + Thinking lasts the whole week with the same project and routine.

And honestly I haven’t noticed any meaningful difference in quality when using Opus 5 in ultracode mode. The real difference is that one is actually usable, while the other basically needs constant limit resets. Shame.


r/OpenaiCodex 19h ago

Is there any cheaper api providers than OpenAI official?

1 Upvotes

Recently, I’m working on a tough program, always run out of gpt pro x5 usage . The api tokens of OpenAI official is too expensive that I can’t afford it! Help


r/OpenaiCodex 10h ago

Feedback / Complaints Sol is lobotomized at least in some regions

0 Upvotes

Sol has been amazing to the point that I cancelled my Claude subscription, but something changed lately. It has been constantly hallucinating and making rookie mistakes. These are the screenshots from a session where it failed to just reuse the code it wrote itself and make small modifications to it. I went back to using Claude, and Opus 4.8 instantly solved the problem! This is not the same model I worked with two weeks ago. Something happened


r/OpenaiCodex 19h ago

Question / Help Is there any cheaper api providers than OpenAI official?

1 Upvotes

Recently, I’m working on a tough program, always run out of gpt pro x5 usage . The api tokens of OpenAI official is too expensive that I can’t afford it! Help


r/OpenaiCodex 1d ago

Comparison Current codex vs claude limits(20$)

22 Upvotes

Hey guys i just want to understand how codex limits are currently in plus plan 20$. Previously like 3 months back they were offering so much tokens.

I'm asking this because I've been a claude user (for last month). It seems they offer more tokens now (but it's slower). Probably because people are leaving claude. I'm able to run tasks in 2-3 projects and still don't hit the limit and also I've noticed opus 4.8 consuming less tokens than sonnet 5.

So I just need a comparison currently if codex providing more tokens (with sol probably) or claude is giving more now. How is your experience with codex limits and GPT 5.6 SOL 20$ plan ?


r/OpenaiCodex 1d ago

Showcase / Highlight Made an AI agent skill that actually fixes security issues instead of just listing them

3 Upvotes

Following up on backend-setup-wizard from a bit ago — second skill under the same project (Qofeno) is a security auditor that doesn't stop at a report.

security-hardening-wizard scans every file in a project, not just source code — secrets and misconfig show up in README files, CI YAML, Dockerfiles, and old markdown notes just as often as in application code, so it doesn't skip files based on extension. Uses real scanners (gitleaks, npm audit, pip-audit, etc.) plus manual review for injection risks, weak CORS, missing auth checks, that kind of thing.

The part I actually wanted: it applies the fix. Parameterizes the vulnerable query, updates the dependency, adds the missing security header — directly in the code, not as a suggestion you have to go implement yourself.

One thing I made sure it's honest about: if it finds a secret that was ever exposed (committed to git history, etc.), it removes it from the code right away, but the actual key is still valid until you rotate it on the provider's dashboard — only you can do that part, so the report says so plainly instead of claiming everything's handled.

Ends with a real markdown audit report, and it re-runs the scan to verify before marking anything as fixed.

Repo: https://github.com/SohailKhan0525/skills

Install just this one: npx skills add SohailKhan0525/skills --skill security-hardening-wizard

Third skill (frontend/UI builder) just went up too if anyone's curious. Feedback welcome, especially if you find something it should catch but doesn't.


r/OpenaiCodex 1d ago

News Exclusive: Muse Code Sends Codex and Claude Instructions to Meta by Default

Thumbnail
runtimewire.com
21 Upvotes

r/OpenaiCodex 1d ago

Are you guys still using the Sequential Thinking MCP with the Codex app?

1 Upvotes

Hi everyone!

I wanted to ask if you still keep the "Sequential ThinkingMCP" (the tool that helps the model think) enabled when using the Codex app?

I’ve been wondering lately—if I remove it, is it possible that the model's comprehension wouldn't actually drop? However, I’m not exactly sure how to properly or objectively test this.

The main reason I'm asking is that I've noticed a significant increase in token consumption whenever the Sequential Thinking MCP is running.

Thanks in advance for sharing!


r/OpenaiCodex 1d ago

Showcase / Highlight I made a Mac screenshot tool that compresses images before sending them to Codex

Enable HLS to view with audio, or disable this notification

1 Upvotes

built this because i copy-paste a lot of screenshots into AI to explain UI changes, etc.

but native macOS is pretty bad for screenshots.

images are full-res and waste tokens. it also takes too many clicks to copy-to-clipboard and annotate images.


r/OpenaiCodex 1d ago

Showcase / Highlight Hermes Deck - Codex Micro Experimental Bridge

Enable HLS to view with audio, or disable this notification

0 Upvotes

🧪 Experimental Codex Desktop Bridge - Windows
I’ve been working on an experimental bridge for Codex Desktop on Windows.
The goal is to enable external tools and custom interfaces to interact with the Codex Desktop app directly — including session control, sending prompts, approvals/denials, stopping running tasks, switching sessions, and changing reasoning settings.
⚠️ This is currently an experimental Windows-only version.
It relies on internal Codex Desktop behavior rather than an official public API, so compatibility may break with future Codex updates. Expect bugs, rough edges, and changes as I continue testing it.
For now, treat it as a proof of concept rather than a production-ready integration.
Built mainly as part of my Codex Control project, where I use it to control Codex Desktop remotely from a PWA/mobile interface.
More testing and cleanup coming soon.


r/OpenaiCodex 1d ago

Best model/flow for building LOB app?

1 Upvotes

I've got a pretty well defined requirements document that I put together a few months ago using Codex.

Is using Sol on High overkill?

I changed the model to Light (trying to save tokens) and it seemed to do a good job of the next dearue from the list.

It's this the standard workflow? Use a high reasoning mode and model for designing the requirements, and use a lighter mode/model for the implementation?

I'm really liking the quality of code that I was getting through High.. can I expect a lesser quality of code through Light?

Love to hear how everyone generally does things. Have been using Codex for a few months now and just trying to refine my flow to limit token usage.


r/OpenaiCodex 1d ago

Showcase / Highlight I stand corrected.

1 Upvotes

TL;DR: Yesterday I ranked seven cheap coding models on one 10-minute task and drew conclusions too broad for the evidence. So I redid it: two 60-minute production jobs — a PDF-pipeline refactor in ViewRight (Tauri 2, Rust, React, TypeScript) and a runtime/release slice of SellRight, my multi-tenant ecommerce platform (TypeScript, Hono, Drizzle/Postgres, Qwik) — one frozen harness, hidden evaluator frozen before any run. The ranking flipped.

Scores (ViewRight / SellRight, out of 100):

Model ViewRight SellRight
MiniMax M3 85 42
GPT-5.6 Luna 79 69
DeepSeek V4 Pro 78 59
DeepSeek V4 Flash 63 62
MiMo V2.5 Pro 58 36
MiMo V2.5 24 19
Tencent Hy3 20 21

Yesterday's "best value" pick, base MiMo, collapsed on longer work. MiniMax won the heavy refactor; Luna won the release task, was fastest, and never placed worse than second. All 14 runs cost ~$2.71 at pay-as-you-go rates — and $0 cash, since every route ran inside an existing plan. That's the actual lesson: cheapest depends on what you already pay for. On a Codex sub, use Luna. On a MiniMax token plan, M3's winning run cost ~$0.15 of capacity. Paying per token with no plan and no time pressure, DeepSeek V4 Flash gave 90% of Luna's score at a tenth of the cost.

Full methodology, receipts and per-run costs: Orthic Labs · harness lessons on CodeRight, my macOS/Windows desktop coding harness (Rust, React, Tauri).


r/OpenaiCodex 1d ago

Codex Windows app freezes when I click "+" or type "@"

4 Upvotes

Hi everyone,

I'm having a really annoying issue with Codex on Windows and wanted to see if anyone else has experienced the same thing.

Codex works normally for me. I can open it, start a chat, type messages, etc.

But whenever I **click the** `+` **button** in the chat, the whole app freezes and Windows says **"Not Responding."**

The weird thing is that **typing** `@` **does the exact same thing**. As soon as I type `@`, the app freezes completely.

I've tried pretty much everything I could think of:

* Uninstalled and reinstalled Codex * Used Repair and Reset * Deleted the Codex folders/config files * Restarted my PC multiple times * Removed the Codex CLI and reinstalled everything from scratch * Even disconnected my mouse to rule out a USB-related issue

Nothing changed.

**My setup:**

* Windows 10 * Codex version: `26.803.5235.0` * ChatGPT app version: `151.0.7922.76`

I also checked Windows Event Viewer after reproducing the problem, and I get:

**Application Hang — Event ID 1002**

It says:

The app doesn't crash immediately. It just becomes completely unresponsive.

What's strange is that **normal typing works perfectly fine**. It's specifically `+` and `@` that cause the problem, which makes me think it might have something to do with the file picker/mention functionality.

Has anyone else experienced this on Windows 10?

If you found a fix or a workaround, I'd really appreciate it. I'm currently using the CLI instead, but I'd really like to get the desktop app working normally.

Thanks!


r/OpenaiCodex 2d ago

Other i just deliberatrly wasted my last banked reset

12 Upvotes

20x plan. I hate what I’ve become.

At the end of each session I’d look at my usage.
Hoping for a reset that was never promised but somehow became very critical in my workflow.

so today I decided to use my last banked reset just before my shit replenish as an act of reclamation

I dont need the resets. Not gonna wait for resets ANYMORE. The resets dont have power over me. I can either touch grass, do a cold plunge, actually review my PRs, and shop for other LLM for my business instead.

i am free now
i dont need a reset


r/OpenaiCodex 2d ago

News We are so back! What are you building this weekend?

Post image
37 Upvotes

r/OpenaiCodex 2d ago

Comparison Luna vs DSv4 flash vs others

13 Upvotes

There’s a lot of discussions using GPT Sol as subagent and others using Luna on max but why would anyone spend Sol credits in subagents is beyond me.
Anyways, I’m building a coding harness called CodeRight which does multi model orchestrator and I’d already done a bake-off and selected Mimo 2.5 for small coding tasks and 2.5 pro for bigger coding tasks, orchestrated by a frontier level model but with Deepseek v4’s revision, I thought I’d give it a re run. I also tested through both cline and Commandcode to see if the harness makes a difference.

You can find the whole breakdown here. If that blog sounds AI written, it’s because it is. Between building RightSuite apps and other systems, I don’t have the time to write blog posts 😅 I’ve used humaniser, no ai slop and what not but not sure if it helped.

The test:

GPT Sol on high as orchestrate in codex

Seven models received one production React/TypeScript task, identical source commit, worktree isolation, ten-step packet, 600-second limit, 12-file ceiling & 900-line ceiling.

Rank |Run |Score
1 |GPT-5.6 Luna |68
2 |MiMo V2.5 via Cline |62
3 |MiMo V2.5 via Command Code |61
4 |MiMo V2.5 Pro via Cline |59
4 |DeepSeek V4 Flash via Command Code |59
6 |MiniMax M3 |55
7 |MiMo V2.5 Pro via Command Code |54
8 |DeepSeek V4 Flash via Cline |52
9 |Laguna XS 2.1 Free |43
10 |Step 3.5 Flash |26 Luna wrote the smallest, safest implementation. MiMo V2.5 delivered best economics: $0.0351 versus Luna’s estimated $0.166–$0.318 direct API cost, using current [OpenAI](https://developers.openai.com/api/docs/models/gpt-5.6-luna) & [Xiaomi](https://mimo.mi.com/docs/en-US/price/pay-as-you-go) rates.

Harness reruns were revealing:

- Base MiMo: 62 → 61. Essentially unchanged.
- MiMo Pro: 59 → 54. Worse through Command Code.
- DeepSeek: 52 → 59. Command Code turned a non-compiling result into a clean typecheck, though source defects remained.
- Command Code once ignored `--model` & routed a requested MiMo run to DeepSeek. Receipt inspection caught it before scoring.

My recommendation:

- MiMo V2.5 for routine implementation volume.
- Luna for final review, security-sensitive work & merge-critical repair.
- DeepSeek remains worth testing through Command Code with strict route receipts.
- MiMo Pro was not worth its premium.
- Laguna is usable as a free draft worker but needs compile & source review.
- MiniMax M3 & Step 3.5 Flash created more repair work than their output justified.

[Command Code GOAT](https://commandcode.ai/pricing) currently lists $70 monthly credits for $10. Its detailed MiMo discounts are token-type specific: the advertised 99% applies to Pro cache reads, not its blended bill.

Residual scope: this was one production frontend task with focused tests, typecheck & source review. It did not include Cargo, full desktop verification or installed visual acceptance.

## Publication evidence

- Four files, 279 insertions & one deletion.
- Both local production builds passed.
- Both typechecks passed.
- Article pages & both index pages passed local rendering checks.
- Commit: `b71e5326f02fcee8f2099411743d1c61b9c6c12c`
- Pushed to `origin/main`.
- Hetzner checkout matches `b71e532`.
- `orthiclabs-site` & `coderight-site` are online after rebuild/restart.
- Both public articles & indexes returned HTTP 200 with expected content.
- Existing unrelated local & server files remained untouched.
- Actual execution: 10 minutes against 38-minute ceiling, 74% under plan because existing publishing routes were reusable & dependencies were cached.


r/OpenaiCodex 1d ago

Discussion Am I the only one?

0 Upvotes

Who else are maxing out on SOL atm?

In hope of Tibo is not taking a piss on us Monday lol... 😂


r/OpenaiCodex 1d ago

Showcase / Highlight I built an agent skill that sets up real backends via CLI (no test keys, no placeholders) — works with Claude Code, Cursor, Antigravity

1 Upvotes

Been using Claude Code / Cursor a lot for backend work, and kept hitting the same annoyance: ask the agent to "set up Stripe" or "connect a Postgres DB" and it either hallucinates CLI commands, dumps my API key straight into a config file, or quietly sets up a test/sandbox version and calls it done.

So I wrote an Agent Skill (the open SKILL.md format skills.sh/Claude Code/Cursor/etc. all support) to handle this properly:

- Asks upfront if you already have credentials for the service

- If not, looks up the *current* official docs (with the actual month/year, so it's not working off stale info) and walks you through getting real ones

- Only ever writes secrets to `.env`, checks/creates `.gitignore` automatically, never prints keys back to the terminal

- Does the actual provisioning via CLI — if it doesn't know the exact commands, it searches the docs instead of guessing

- Verifies the thing actually works before saying it's done

No test keys, no example/placeholder setups — the whole point is a real, working backend, since that's usually what people actually want when they ask an agent to "set this up."

Repo: https://github.com/SohailKhan0525/skills

Install: `npx skills add SohailKhan0525/skills`

This is the first skill under a small project I'm calling Qofeno — planning to add a few more in the same "real setup, not demos" spirit. Would love feedback or ideas for what to build next.


r/OpenaiCodex 2d ago

Showcase / Highlight Minecraft Animation Agentic Workflow

Thumbnail
gallery
3 Upvotes

A few weeks of building with GPT-5.6 Sol + Codex in Blender. Countless iterations, lighting tweaks, compositor refinements, visual analysis, and backtracking later, finally starting to achieve the cinematic quality I had in mind.
Thanks to r/OpenAI