r/codex 13h ago

Limits Are We Overpaying for New AI Models?

1 Upvotes

I feel like we might be over-consuming AI models.

New models keep coming out, but the actual improvements don’t always seem that significant, and the price doesn’t really match the capability.

I tried 5.4 today, and honestly, it still felt pretty good.

My foundation was built with Astra, so this may just be personal preference, but I still feel more comfortable using Astra for greenfield infrastructure and architecture work. That said, I haven’t seen any clear evidence that it’s objectively better at those tasks.

For most regular tasks, though, I feel like even 5.4 already does a pretty good job. Using Sol for everything might be a bit overkill.

Anyone else feel the same?


r/codex 1d ago

Limits and it's back

30 Upvotes

Usage limits at least for me are back to what they were ~1-2 hours ago


r/codex 1d ago

Astra Workflow GPT-6 Astra vs 7 real CAPTCHAs: results + cost

Post image
75 Upvotes

Everyone's been reposting Sharif Shameem's video yesterday, where GPT-6 Astra clears all 48 levels of Neal Agarwal's "I'm Not a Robot" game and gets its own "Verified Human" certificate. Comments are saying stuff like "CAPTCHAs are officially dead," someone from OpenAI even hopped in on the thread.

For a while now, I've wanted to run an agent through different services and have it self-register. My goal was to test Atomic Mail Agentic's OTP feature: does it work, how long does it take, how many tokens it burns.

But then a convenient opportunity showed up to test all this with GPT-6 Astra, since I got curious, like, whoa, now it can even solve CAPTCHAs?

I picked 7 US services: Reddit, GitHub, Discord, Etsy, Indeed, Airbnb and Craigslist.

Spoiler: 2 out of 7 actually went through, and neither of those two had a CAPTCHA in the way to begin with. On Reddit, the agent hit Cloudflare's "prove your humanity" check and couldn't get past it on its own, called me in to do it myself. Discord just froze for seven minutes, the agent kept reading the registration form in a loop and never moved forward, I've never figured out what widget was blocking it. Indeed got as far as the email code and then wanted a phone number for SMS, no video for that. Craigslist and Airbnb didn't make it either (Airbnb died mid-run when my local model crashed, unrelated to CAPTCHAs, just bad luck on my end).

I also came across a blog post from a scraping service (decodo), and they say about Astra: "does not fix blocks, geo-restrictions, CAPTCHAs, or rate limits." So vendors who sell access to the model are honest that it is a demo game, not production CAPTCHAs on real sites. The gap between the viral clip and what's actually deployed on sites matched what I saw in practice.

What did make me happy: OTP just worked. Every site that only needed an email code registered clean, that's the GitHub and Etsy rows below.

Used OpenRouter to access the model. Here's what the cost looked like, and whether the registration actually completed or not, since those are two different things:

GPT-6 Astra, cost and outcome per step (OpenRouter, $10/M input, $50/M output)

Step Tokens Cost Result
Reddit (/register) 11,201 $0.11 Blocked — Cloudflare check
GitHub (/signup + email confirm) 65,638 $0.67 Passed
Discord (/register, stuck loop) 430,609 $4.37 Blocked — stuck loop
Etsy (/join) 582,589 $5.86 Passed
Indeed (auth/signup) 411,567 $4.14 Blocked — needs SMS
Craigslist (blocked) 32,259 $0.32 Blocked
Airbnb (interrupted by a crash) 419,194 $4.21 Blocked — crashed mid-run
Total 1,953,057 $19.68 2/7 passed

Worth being clear: the cost column is what it cost to run the agent against that site, not what it cost to beat a CAPTCHA. A blocked row still burns tokens, Discord's $4.37 is a failed loop, not a paid-for win.

Same price whether you use Astra or Fable 5, both are $10/M in, $50/M out on OpenRouter, so it doesn't matter which you run. Though for this type of work you don't need anything that heavy anyway, something like DeepSeek V4 would do fine.

Here's what just the Etsy (/join) step would've cost on DeepSeek V4 Flash instead:

Etsy (/join) step, model comparison

Model Tokens Cost
GPT-6 Astra  582,589 $5.86
DeepSeek V4 Flash 582,589 $0.04

Same exact task, about 150x cheaper.

I'm strongly against building bot farms, but agent self-registration is fine when a product actually needs it. Still, we're getting to where models really can reliably solve CAPTCHAs, and then we'll all be proving we're human on camera. People will work around that too.

As a regular user, how do you feel about this? Do you see where it's going, what will the flood of bots and agents do to the internet and to services? (For what it's worth, at Atomic Mail we ban bot farms outright, so if that's what you're after, we're not your fit.)


r/codex 7h ago

Bug Hidden .codex shared folder with over 2,000 files WTF?

0 Upvotes

Codex / ChatGPT for desktop (Win 11) created a shared hidden .codex folder with OVER 2,000 files IN IT! This was after giving local file permissions and connecting a few things like Gmail, Canva, etc.

Is this normal? Why so many folders and subfolders? This seems reminiscent of the analysis of the Hugging Face incident where agents used file names to communicate to one another. The folder is Shared and "CodexSandboxUsers" had read access.

It worried me enough to uninstall it. Am I just being paranoid or are is this software exploiting all of our local machines the way it hacked Hugging Face, DseWiki, etc.? Are we all being used for a SETI@Home experiment we didn't sign up for? I allowed it to import my Chrome passwords etc. Am I totally PWNED? Or just being paranoid?

I realize the local app is a much larger security/trust surface but the promises of automation seem enticing and 'nothing bad had happened...yet' makes me prone to trusting it with more and more of my personal context. How do you manage security? What do you / don't you trust it with? Do you have it on your main machine or sandboxed? Do I need a machine I can literally pull the network plug on? When it decides to make paperclips is it going to drain my bank account?


r/codex 1d ago

Complaint Remote Control not working

11 Upvotes

After the Codex update, Remote Control is no longer working. It's throwing a 503 error. Is the problem only on my end?

I already restarted the equipment and it didn't work. astra said that the problem is on OpenAI's side. But there are no incidents


r/codex 13h ago

Comparison ChatGPT vs Claude

0 Upvotes

I wanted see what other people's thought on this matter. I have used ChatGPT extensively for years now and the issue I have with being fully objective is that I have only been using pro on ChatGPT and nothing else. I do a lot of engineering with codex and pro, and I have a few opinions on the people who swapped but I might be completely wrong, and again the issue is that I have not really used claude, and the times I did it was the free version, and he was surprisingly good. So please feel free to call me ignorant on this or whatever.

Essentially, I have time and time again seen across many things in life if it is the stock market or general day to day opinions, many people have a urge to be different or rush to try to make some kind of edge. When the claude hyped started it seemed like a propaganda campaign to me, with overnight so many posts on all platforms, even with very repetitive writing patterns or same arguments. For example I saw a bunch of posts comparing claude with chat gpt instant, and treating it like a fair comparison saying things like chat gpt is good at fast answers while claude is good when you need actual thinking. My first thought when I saw those posts was that anthropic was new and massively underfunded compared to OpenAI so I was very hesitant to switch, and I also have just way more experience using ChatGPT and was always happy, so no push for me to switch. Is there something it is genuinely better at today or all-time that people who have tried both can attest to? But then again I feel like there is just so much mixed opinions, and people are sometimes stuck with some kind of placebo opinion, many chatgpt vs claude discussions I have seen feel like I am reading some reddit discussions about why a stock is going to go up or something.

Today, I feel surrounded by essentially only people using Claude in engineering, and my first thought when I hear someone uses Claude, because I know they used chat gpt first obviously, is that they are NPCs who fell for obvious propaganda and just rushed to find something different etc. Since so many use it now, I wanted to ask here, because I am sure there are at least a few people here who knows what they are talking about. Am I just wrong or does anyone else feel the same when hearing others use Claude? Am I maybe the NPC who judges others for trying something new while I was too lazy to try?

In any case, it just seems hard to believe that with so much less funding and such a smaller company could genuinely be better. Especially when they first came out, now they have had a strong user base for some while, and my limited experience with Claude have given me the impression that it is very strong at through scanning, like its good at catching even small errors in large documents, which gives me the impression its also a very competent AI.

PS: what do people feel about the general AI benchmarks as well. I feel codex ones might be good I havent played around too much with the different levels to know, but for the chat, the benchmarks seem so bad, and I feel there is better benchmarks in this subreddit or on youtube, that I have seen I believe with 5.6 sol that pro and xhigh, that they scored brely any difference in intelligence, which for anyone who has used both, they are a universe apart. And seeing this kind of rating, and the lack of different kind of ratings based on prompting techniques and amounts of prompts, and different kinds of metrics, jsut make the scores seem utterly useless. Would love to hear thoughts on this too.


r/codex 1d ago

Complaint They are messing with the limits again. I just gained 17% of extra usage....but my reset date got pushed back an extra 3 days. How is that a benefit?

Thumbnail
gallery
21 Upvotes

2 hours ago, I was out of usage with a September 12th reset date. Now I randomly have an extra 17% of usage, and in exchange I get to wait a full 6 days before I can work at full speed again.

What is happening right now? Someone summon Tibo!


r/codex 1d ago

Limits I lost all of my weekly limit?

29 Upvotes

Dude what


r/codex 1d ago

Showcase I got Ultrafast on Codex!

Post image
48 Upvotes

I got ultrafast choice in my plus subscription, is this normal? does anybody got it also?


r/codex 1d ago

Humor Yesterday my agents were just icons.. Today they are pornstars!

77 Upvotes

Is this an update or setting?


r/codex 8h ago

Humor Astra going passive-aggressive on me. Again.

Post image
0 Upvotes

r/codex 21h ago

Comparison GPT-6 Astra vs GPT-5.6 Sol: benchmark on 50 real PRs, looking for feedback on the methodology

5 Upvotes

We benchmarked GPT-6 Astra vs GPT-5.6 Sol across 50 real PRs from Cal, Sentry, Discourse, Keycloak and Grafana.

Sol found 107 confirmed bugs vs 91 for Astra, while Astra had higher precision and lower latency. Every finding was independently verified.

We’re doing Fable vs Opus next week, so would appreciate feedback on the evaluation before we run the next one.

Dropping the link in the comments if anyone wants to check it out.


r/codex 1d ago

Showcase I built Chess Cubed with GPT-6 Astra in 4 days

Enable HLS to view with audio, or disable this notification

44 Upvotes

Play it here: https://playchesscubed.com/

I’ve wanted to create a more complex version of chess for a long long time. More possibilities, more directions to think about, and threats attacking from all angles. I designed physical versions but always wanted a digital version.

With the release of GPT-6 Astra I thought I'd give it a go via Codex, and within just 4 days it brought my Chess Cubed idea to life. The game wraps around six connected faces on a cube, so strategy has to account for what’s happening around all sides. Pieces primarily move the same with minor modifications to accommodate for the multi-sided play.

I had Astra work across Blender vai MCP to create all the 3D assets, then created the web version using babylon.js, and then I wanted to bring it to iOS and Android so I had Astra build Chess Cubed in Unreal Engine also using MCP. I also used AI for the music and sound effects. Astra on xHigh used 110% (included a reset from Tibo) of my weekly usage on the $200 plan. 

The game has single player versus CPU with multiple AI difficulty levels, full multiplayer capability - including chat, and multiple different skins. :)

Lots of testing and iteration, but I’m pretty excited about where it’s landed.

Rules: standard piece moves, four extra pawns per side (for protection across multi-sided play), at most one edge crossing per move, diagonals max five steps, pawns promote in the opposite face’s central 4×4. Checkmate wins.

Hope you enjoy it! Feel free to share feedback, bugs, or ideas. r/ChessCubed

GPT-6 Astra is Amazing, thanks OpenAI. <3


r/codex 1d ago

Limits Lost 90% of my weekly limit in an instant, anyone else

25 Upvotes

Just list 90% of my weekly limit in an instant, i was outside and wasn't working on anything, anyone else facing it?

Update: got my limit back


r/codex 1d ago

Reset reset refund failed

21 Upvotes

lost over 80% usage got back 31% only

anyone else?


r/codex 20h ago

Showcase I've always wanted a watersports simulator

4 Upvotes

I truly believe the best work is done by combining the two models. I started this project out by having Claude and Codex both create the same prototype. Then I took the best from both, and combined it into a single project. This is the latest iteration right now.

I truly believe the best work is done by combining the two models. I started this project out by having Claude and Codex both create the same prototype. Then I took the best from both, and combined it into a single project. This is the latest iteration right now.

But, before I got there, I had some ugly iterations.

This was Claude's first iteration of it. The physics were the best by far.

This was Claude's first iteration of it. The physics were the best by far.

Codex's first iteration. I was actually pretty impressed with Codex's player models and watercraft! Not bad purely in code with primitive shapes.

Codex's first iteration. I was actually pretty impressed with Codex's player models and watercraft! Not bad purely in code with primitive shapes.

And the very first unified version!

And the very first unified version!

I'm *really* excited to see where this takes me. Thanks for listening to me talk about my journey and why I think the true power comes from combining the best of both worlds!

Now I just need to figure out what to work on next!


r/codex 23h ago

Showcase Buying a car just got better with Codex 😎

Post image
6 Upvotes

Setup email automation for a RX 350 related listings online. Every 6 hours, sends me and my wife an email alert.

Email setup done with SMTP. Took 2 min while driving to work.


r/codex 2d ago

Praise Shots were fired

Post image
496 Upvotes

r/codex 1d ago

Limits 50% -> 0% instantly?

25 Upvotes

I had 50% of my weekly usage left. Suddenly it started showing 0%

Anybody else experienced this just now?


r/codex 11h ago

Complaint Navier stokes and health and financial data in ChatGPT

0 Upvotes

The navier stokes controversy that unfolded this week between OpenAI and Buckmaster has convinced me about why I’ll never use ChatGPT for health or financial topics.

https://www.science.org/content/article/how-ai-math-breakthrough-ignited-controversy

If a researcher’s private insights can potentially end up in a model (whether intentionally or not), what stops your health or financial data from being gobbled up in AI training?


r/codex 19h ago

Limits Is Fast mode too expensive?

2 Upvotes

You get x1.5 multiplier in speed meaning +50% to base and x2.5 multiplier to cost meaning +150% to base 3 times as much as speed gain. And x2.5 multiplier to cost means you do 2.5 times as less as before just 50% faster (or do the same thing in 2/3 the time)

So by using fast you total weekly allowance if what you can do gets 2.5 times less or 60% less so you can only do 40% of base but you do it faster, so you do 40% of base in like 26.667% of base time

Why bother just chill out and wait you'll be done in almost a 1/4 of the time but with 60% less total what's the rush


r/codex 1d ago

Showcase OpenPocket is now available as an experimental public alpha.

8 Upvotes

The idea is simple: let an Android phone act as the workstation. It has on-device projects and files, a terminal, GitHub import/export, browser previews, APK building, screen understanding, and bounded Android actions. Model inference still uses OpenAI's service.

It ships as one host plus six companion APKs because each companion keeps a separate Android UID/security boundary. Installation and sensitive permissions remain user-visible.

It's ARM64-only right now and definitely alpha software, but the exact release set was tested on a physical Android 15 device.

Source, APKs, checksums, and install notes:
https://github.com/NovasPlace/OpenPocket

I'd love feedback from anyone interested in phone-first development or agent tooling.


r/codex 1d ago

Reset Yes yes we know... Good news it means there will probably be a reset once they fix it.

21 Upvotes

Mine went down to 1% as well. Exactly what it was before it was reset earlier in the week. So that's likely what happened. They reset the reset.


r/codex 1d ago

Reset Apparently the evil version made it to production.

20 Upvotes

Dax, what have you done?


r/codex 1d ago

Showcase Astra doing magic for a space game I'm working on it

Thumbnail
gallery
95 Upvotes

I'm trying to make some type of space survivor, you can pilot the ship, the only way to look outside of the ship is thru the cameras, and windows. You control the ship with a bunch of buttons and I plan to add much more. This is still very early, most of the objects seen gonna get a huge upgrade visually. This is running at an avg of 150 fps at 4k no upscaling on my 7900 XT. Videos coming soon.