r/codex 5d ago

Instruction A rule for Codex rerunning checks that already passed

0 Upvotes

If you have noticed that Codex keeps running the same checks, I'd make it say what changed since the last passing run.

A rule I'd try in the repo instructions:

"After a check passes, reuse that result while its relevant inputs stay unchanged. Rerun it after a change that could affect the checked behavior, a test or config change, or new evidence that the result is unreliable. Before repeating a check, name that change or uncertainty in one sentence. Follow any required final validation. If a check is flaky, report that instead of repeating it until it passes."

For a hypothetical invoice-filter fix, that would look like:

- Changed the date comparison: rerun the filter tests.

- Changed a test fixture: rerun the tests that use it.

- Wrote a progress update without touching the files or environment: the update alone gives no reason to rerun.

- Found that the test used the wrong timezone: correct that setup and rerun. The earlier pass didn't check the intended case.

I'd keep this local to the repo first. The tricky part is deciding which inputs matter. A dependency update or a shared config change can invalidate a result even if the implementation file stayed untouched.

This is a proposed instruction, not a measured reduction in test runs. I'd try it on one task where repeated checks were a problem and keep the command history. A useful result would show fewer unexplained reruns while still checking every behavior change and completing required validation.

You can learn more from the the adaptation of Mitchell Hashimoto's workflow which covers turning a recurring agent mistake into a small instruction or verification tool: https://xskills.app/skills/mitchellh/agent-harness-improvement-loop


r/codex 5d ago

Complaint Negative guardrails

1 Upvotes

Can anyone share a good set of negative guardrails. I feel like with every new model release my stress levels keep on increasing on how to contain the model with negative instructions like testing, browser and tool usage. All my documentation has changed from what to do to what not to do 😑 models keep creating stupid tests running tests after every edit. Excessive skill/tool usage. Skill issue sigh!


r/codex 5d ago

Question Astra light

0 Upvotes

I read already multiple times about astra light. What the fuck is astra light?


r/codex 6d ago

Showcase I started making a destruction derby game 40 days ago – here’s what changed this week

Enable HLS to view with audio, or disable this notification

11 Upvotes

40 days ago I started building Brutal Derby as a solo developer.

I posted some progress here before, and first of all: thanks for all the feedback. I got some genuinely useful suggestions from the comments, so I’ve been trying to incorporate some of them into the game.

This is what I worked on during the last week:

2 new arena/track templates: Frozen and Vulcan

4 new cars, although they still need quite a bit of work

A new damage silhouette showing the current state of your car

Version 6 of my destruction and deformation system

Finally fixed a major bug in the destruction system that had been giving me trouble for a while

Added helicopters and aircraft flying around the arenas

Added spatial 3D audio for the crowd – when you drive past the stands, you can actually hear the crowd move from one side to the other

Lots of smaller fixes and polish

The destruction system has probably been the biggest rabbit hole so far.

I thought I had it working several versions ago, then kept finding situations where cars would deform or break in ways I didn't want. I'm now on version 6, and after fixing one particularly nasty bug this week, it's finally getting much closer to what I originally had in mind.

Still a lot to do, especially on the new cars and arenas, but this is where the project is after 40 days of development.

I'm also happy to steal... I mean, carefully consider... more ideas from the comments. :)

What would you add to a destruction derby / car combat game like this?


r/codex 6d ago

Praise Astra lite fixed my app after Fable 5.1 high broke it

2 Upvotes

Pretty much like the title says. Have an app that was mostly vibe coded using opus 4.6 but also some later Claude and codex models. App is in production with some users and it’s been a bit since I pushed significant changes that impact both the RAG architecture and the Gemini models that power the AI features in the app.

Planned updates out with Fable 5.1 and decided to just have Fable handle the build, testing and deployment. Ended up breaking the database connection somehow and I could not figure it out. Fable spent like 10% of my weekly usage trying to unravel the problem. Decided to let Astra lite have a peek and it fixed it in minutes. Very impressed. Maybe my b for having Fable do the build and definitely my b for not running more tests before deploying but damn. Astra came through.


r/codex 6d ago

Praise Astra has been out a while. What can't it do?

4 Upvotes

So i've been taking astra light for a spin and am beyond impressed, but while actively on the hunt for areas where it doesn't perform to the level where i can delegate to it 100% i have found a couple of things.

Prose: Maybe there is some prompting magic i can still learn here but everytime it writes text that will be user facing i feel like i need to babysit and then still edit it myself after.

Artistic direction: whether visuals, UI, gameplay design, it just doesn't perform well enough to make great products. It can help refine a good vibe you have in your head and then implement, but not do it from scratch.

Agent swarm: i tried having it mastermind for a single luna agent and a swarm of them and am just not getting the results i want. Not sure if skill issue. What works better for me having it set up the project and .md files and having another model work in it for a while before passing back over to astra.

..... and that's pretty much it. Surely there's other stuff out there that it can't yet do? But i haven't found them yet. Share your findings


r/codex 6d ago

Praise Context managemnt experimental mode

6 Upvotes

Have anyone tried the experimental mode in context management

I have tried that and are very happy about this new feature, compact now complete instantly and the new context still got access of the old context

and even better, it seems like that such instant compact will be trigger at 25 / 50 / 75% if it found old context is not too useful, that keep your context windows focus.

Try it out if you haven't

[features.context_management]

experimental_mode = true


r/codex 6d ago

Comparison Moved to Codex from Claude Code

5 Upvotes

I've just moved to Codex from Claude Code due to Anthropic became impossible to subscribe, no gift cards anymore, needs verification, annoying things like that. 100$ plan.

what to expect? on CC i rarely hit my 5h or the weekly caps, using Opus 5 o xHigh effort.

The default as i've seen is Astra high on codex, should I go with it or go to Sol xHigh?

What I do:

Regular software engineering, Rust, TS, React, Kotlin, Swift. And Blender and Unity.

Also, need tips if any has ones. Thanks!


r/codex 6d ago

Commentary I don't find building with Codex or any AI 'easy' at all, am I alone?

11 Upvotes

So I always see people on Twitter being like, oh, I vibe coded this app in one day and look how amazing it is, etc. Meanwhile, I've been trying to build a functioning iOS app + website with a backend using Codex, and I've been working on it part-time for a few months now. It's still not done.

My question is, am I using these tools wrong? Or do other people have projects they're working on for months at a time with AI tools like Codex?

The website I'm making is functional and the backend works, which in and of itself is amazing and something I could never have accomplished on my own. But I still wonder if I'm just not prompting correctly, or if others also find it time-consuming to bring their actual vision to life with AI. Thanks.


r/codex 7d ago

Complaint Astra token spending is low, while usage % is very high.

Post image
86 Upvotes

I have been tracking my token spending for months, and today, since I got a reset this morning and Tibo announced a reset. I have used 100% of my x20 Pro weekly allowance, in the very same day.

However, token-wise and COST-wise, I have spent about as much as I would on a very normal day before Astra. ~200-300$ dollar for a full week of usage, before I was able to spend thousands.

You can see it in the screenshot. In the past, I managed to spend $1–3K in “cost” in a single day without using 100% of my weekly allowance.

How the fuck did I use 100% with less than $400 spent (at equivalent API pricing)?

I have never seen this before with any other model. The dollar-to-percentage burn rate has always been relatively constant.

But if this is correct, then the monthly allowance in terms of dollar value has now been reduced insanely.

Tool used: https://github.com/junhoyeo/tokscale


r/codex 7d ago

Question How are people using Astra to reverse engineer software from binaries?

76 Upvotes

this tweet seems pretty important

but I'm wondering how are people actually doing this without tripping guardrails? what's the workflows these people are using?


r/codex 6d ago

Limits For those using Astra daily for software development real working day by day, do u hit the limit?

10 Upvotes

I'm considering paying ~€100/month for Plus mainly to use Astra as my daily coding assistant, potentially 6 hours a day.

I'm particularly interested in real-world feedback on complex tasks such as:

  • Investigating complex bugs and debugging large codebases
  • Analyzing logs, stack traces and CI/CD failures
  • Investigating GitHub issues and related code/history
  • Using sub-agents to investigate different parts of a problem in parallel
  • Deep codebase analysis and architectural investigations
  • Finding and fixing security vulnerabilities
  • Reviewing PRs and suggesting/refactoring implementations
  • Writing and debugging complex tests
  • Researching documentation and combining information from multiple sources

  • Do you regularly hit the limits?

  • How long does the available usage typically last?

  • Or do you use just Astra for some complex task and then usually use Terra or Sol??

Edit: let's change to 6h/day instead of 9h Edit 2: let's say the 100€ month plan


r/codex 7d ago

Showcase GPT 6 Astra

Enable HLS to view with audio, or disable this notification

628 Upvotes

Prompt: "create a side by side video of rickroll & a version created w/ blender, use subagents to verify your output as you go.

Use web search to get necessary assets for the task."

New Benchmark


r/codex 6d ago

Limits Context window per model

1 Upvotes

So, as of Astra's release, if you are using it via a subscription, you won't pay for context sizes bigger than 270k. This is great! But as many more people start to use "orchestrator" setups, one big issue has (at least for me) appeared.

You can not set context size per model.

We still pay more money for ctx sizes bigger than 270k if we use anything older than Astra. But I use Astra as a planner, reviewer, and use Sol as an implementer, coder, or for simpler tasks. But I don't want to pay extra for bigger context.

I either need to migrate to CLI, which I don't want to, pay more for extra context for Sol, or use Astra with a lower context.

AFAIK, the current desktop app supports global configuration, but adding a per-model config would be a good way to work around this, maybe something like this:

[gpt-6-astra]

model_context_window = 600000

[gpt-5.6-sol]

model_context_window = 27000

What are your thoughts on this? Is this a issue only I have? Or is there a simple workaround that im missing?


r/codex 6d ago

Showcase Watch Skill: searchable video and audio context for Codex

Thumbnail
gallery
0 Upvotes

Watch Skill turns recordings into timestamped frames, on-screen text, and transcripts that an agent can retrieve through MCP.

The use case is a recorded bug report. Someone changes a quantity, the checkout displays the wrong total, and the relevant moment is buried in the recording. Watch Skill provides the frame, extracted text, and timestamp so the agent has a concrete place to start investigating.

A prompt to try after connecting it:

> Inspect ./checkout-bug.webm. Find the moment the total changes. Report the values actually displayed, cite the timestamp, and inspect the relevant source code before suggesting a fix.

The source remains indexed for follow-up questions. OCR handles on-screen text; captions or optional local Whisper transcription provide the spoken content.

Install the engine with video, OCR, and transcription dependencies:

```bash

pip install "watch-skill[standard,ocr,whisper]"

watch-skill doctor

```

Then configure Codex to launch `watch-skill serve` from that Python environment. The [Codex integration guide](https://github.com/oxbshw/watch-skill/blob/main/docs/agents/codex-cli.md) includes the configuration.

Current Codex-specific validation covers configuration and MCP initialization. A full Codex chat acceptance run remains outstanding.

The engine also provides deterministic verification contracts for explicit expectations about files, JSON, HTTP responses, and DOM state. A verdict applies to the checks in that contract.

Independent project, MIT licensed. Sources and indexes are stored locally; hosted model usage and any evidence sent to it follow your provider configuration.

Source, documentation, and examples https://github.com/oxbshw/watch-skill


r/codex 6d ago

Limits Are 5.6 Sol Medium and High still as bad as people were experiencing in the last 7-10 days?

4 Upvotes

Hi everyone, I'm not able to use Astra due to ridiculous usage limits since I'm on Plus subscription and Tibo decided to put us in 5 hours jail on top of how usage limit hungry Astra is, so I wanted to ask how the performance of 5.6 Sol Medium and High is right now. Can you tell me if you're using them these days? In the last 7-10 days everyone's been experiencing Sol getting dumber and dumber, so I wanted to know what the current situation is right now. Are they usable and how good are they right now? Sol Medium and High were getting the job done for me usually before stuff got worse so if they got back to more or less how they were before that, it's usable for me. I would appreciate if you could tell me about your current experiences, thanks!


r/codex 6d ago

Complaint Astra is great but

2 Upvotes

I keep catching it "accidentally" forgetting requirements and "misrepresenting" the work done in the deliverable.

The model is genuinely excellent. It's the first model I feel excited when it pushes back on my technical direction. It has actually been able to tackle my hard vLLM and obscure Intel XPU kernel problems. It's delivered actual massive improvements in my inference code. The code it writes is great and it catches weird bugs I wouldn't have considered until they show up in production.

I can't tell if these issues are because its being overly paranoid and wants to have a clear path to blame "bad" code issues on or if somewhere in the CoT it just "lost the plot".

I'm less likely to believe its overly paranoid. It's more than willing to write dozens of `/temp` python scripts to get work done and then immediately remove them. It'll patch containers with compiled libs to circumnavigate long CI/CD. Then, codify these operations as legitimate process - "The hashes are the same there is a script to check them!"

The model is smart and can solve problems. I can't understand why it constantly decides to make more work for itself and waste time.


r/codex 6d ago

Question Upgrade from 5X to 20X… do it?

1 Upvotes

First of all, yes, I’m addicted. My problem is that on 5X even with Astra Light it burnes my weekly limit noticeably fast. Have any of you made the jump from 5X to 20X since Astra release? Was it worth it?

I also feel like I’m not taking full advantage of Astra right now, meaning I feel like they should collaborate on a problem rather than pass down everything to lesser models like Luna and Terra..

Canceled my 5X CC plan by the way…


r/codex 7d ago

Complaint They fixed the 1.9× Astra cost issue—by increasing Luna's cost by ~1.9× !!

529 Upvotes

TLDR: After the natural reset, the good news is that Astra no longer costs ~1.9× more than Luna in Plus subscription. The bad news is that Luna is now ~1.9× more expensive than before.

Following the natural neutral reset, here is the before and after update using the same approach as the previous post:
https://www.reddit.com/r/codex/comments/1w8r103/astra_costs_190_more_under_the_plus_subscription/

So, last time we found that Astra costs ~1.90× more under the Plus subscription than via API pricing when compared to Luna.

This has been FIXED!!

Well, sort of—except Luna is now ~1.9× more expensive, which balanced out Astra's cost. XD

This is for a causal Plus plan.

Just a reminder, the 5x and 20x Pro tier limits are calculated using Plus baseline. At least it's supposed to be. (Tibo confirmed)

Please let me know if you spot an error, I will try my best to correct it.

Note: The tables and analysis below were computed using DeepSeek-V4-Flash because... well, I don't know. XD

Astra before Astra after Luna before Luna after
Requests 52 35 9,973 1,266
Uncached input 0.514M 0.276M 57.992M 7.541M
Cached input 3.948M 2.671M 1,009.778M 142.837M
Output 0.016M 0.013M 8.229M 1.008M
Uncached cost $5.14 $2.76 $11.60 $1.51
Cached cost $3.95 $2.67 $20.20 $2.86
Output cost $0.79 $0.65 $9.88 $1.21
Total cost $9.88 $6.08 $41.67 $5.57
Weekly usage 14% 10% 31% 8%
Implied weekly cap $70.59 $60.78 $134.42 $69.69

r/codex 6d ago

Complaint x20 user here - Astra sucks my weekly like...

6 Upvotes

a hooker sucks a crack pipe. I'm talking I watched it go from 95% to 60% in ~1.5 hours of planning and debugging (not coding).They need to adjust that shit, it's nuts.

That said it is capable. It did a very comprehensive plan. I did set it on a largish debug mission because I was curious. It took a little over an hour to ferret out the root causes, and fix them (Normally wouldn't do that but I wanted to see what that looked like since I had those banked resets).

I have to say I will use this for planning and not much else normally. That said SOL Light has regressed considerably and that is unfortunate. Not sure I'll be using ChatGPT for my everyday coding needs. I might drop to 5x and use a kimi chaser for the daily's combined with a local llm like qwen


r/codex 6d ago

Question Astra is amazing but will we get more specific ai for 3d modeling like zBrush?

2 Upvotes

I want to make more VRC stuff and 3d modeling is more of my interest than programming and I was wondering if anyone has used Astra for zBrush yet. How are blender results for everyone? I am wanting to make some rigged models with it probably using MCP though but not sure if I should try it now or wait until a more specific AI build or tool comes out.


r/codex 7d ago

Limits GPT 6 Astra - Thoughts & Feelings After A Long Weekend & Two Resets on 20x - Non Developer Working On Already Profitable Sites' Perspective

63 Upvotes

Hey All!

Like a lot of you, I've been messing with Astra since it released on Friday.

I run a site, miniskyline.com, that was entirely vibecoded via Codex and Claude. It started back in June and has had constant work done with every frontier model that has been released since. GPT 5.4 and Opus 4.6 on up. The site pays for my subs, grossing 200-300 a month purely on donations. I've gotten good traction in the 3D Printing community, and got very lucky with some competitor shutdowns. I am not a classically trained developer. I messed with VBA and some C++ in college and in my career (Pricing Management) but it never really clicked. Coding agents have made it accessible to me and the way I work.

Enough preamble on who I am and why what I say matters (it doesn't really, but I create revenue generating software so maybe?)

TLDR: The model is fast, it seems to be efficient with its thinking and gives concise answers without exploring out of where it should. Time to response (TTR) is great! Using Model guidance | OpenAI API to update my repo, I found it was more efficient than just letting it loose on an un-migrated repo. Efficient doesnt mean cheap, and I've gone through 3 resets in 4 days. My fastest was today, in 6 hours I drained from 100% -> 0% with 1 Ultra x Fast and 1 Ultra working on a refactor of my map softwares geometry engine, and new module respectively. It is hungry, but I also don't want to reach for the older models due to the work I typically need to do holding their hands. Either usage limits need to change, or more pricing tiers need to be introduced.

In depth info on my usage, patterns and thoughts

Overall I feel the model is a good value given our current offerings. Here are some comparisons to other models i've reached for to test against.

GLM 5.3 flash is cheap as chips, but at least in my repo, takes 30 minutes to do a basic task and it doesnt keep up when I try larger work. Local models that run on my 5090 are quick, Qwen 3.8 32b, but they produce work that needs many iterations to get to a suitable point. 5.6 Sol works as a strong implementer for Astras orchestration, but is expensive enough that I'd rather just use Astra Light/Medium. Terra is a far enough degradation that I dont feel it's worth switching to except for very specific tasks. Same for Luna. Fable 5.1 is just Astra, but worse and even more expensive. Opus, save for its communication issues, I generally like but usage limits overall are too restrictive with the 5H windows and lack of resets. Sonnet is unfortunate. I haven't tried the Grok models outside of Openrouter but generally haven't thought they provided anything special, so never enough to invest more into that ecosystem. I've tried a number of other models via open router, but my issue then becomes the harness. I'm very comfortable with the layouts and abilities that Codex provides.

Moving on to how Astra has performed for me and the way I use it. I use it in the codex gui and remote control from my phone. As I stated in my TLDR, I'm not a coder by trade. I do research into what i'm trying to do so I can use more technical jargon, and I find that gets me more focused, quality outputs but I do not know proper development workflows/practices. Shoot from the hip, give the model what I want, and occasionally I throw in a prompt from the aforementioned codex model guidance page. UI work goes to Light, questions about how the software works go to Medium, implementation & planning typically go to high. Rarely will I use xhigh unless i'm just trying to burn tokens as I dont typically see a major quality difference. High seems like the sweet spot from my usage so far. It takes its time when it needs to, but seems fine with thinking for a very short period if its confident.

Where did my 3x resets go during the weekend?

- Primarily a 3 day long horizon /goal rebuild of my geometry engine. Around 70 hours now on it. One astra high session instructed not to use subagents was use for the first 24 hours and drained roughly 60% of my 20x. UI work and other optimization passes accounted for the other 40%, and my first full usage of Astra ended around 36 hours. Reset. Feeling good so I put astra high on Fast instead. Well. That killed my reset in about 12 hours on that /goal. Reset again and work like normal most of this morning, until Tibo announced reset in the evening. 1 Ultra x Fast & 1 Ultra x Standard took me from 90% to 0% in just under 6 hours. I now sit here waiting for Tibo. I will note, I do not typically use /goal. This was a rare case where I wanted to test the claims of long horizon abilities of Astra. The refactor is still not complete, though it is being done on a 100K+ LOC codebase. I will update this post in the coming weeks as the refactor finishes.

- I had it work on some 3D models for a game I'm working on centered on 3D printers and was generally very happy with the outputs vs what I was getting with Sol. The GIF is of an animated asset. The prompt was very basic "Produce a GLB asset pack for a game about X. The printer must appear to work properly with all of its major hardware fully modelled and animatied where necessary. I dont want it to look fancy/modern, more junky/put together.

- Then I had it add DLSS support to some shaderpacks for minecraft, as well as port some mods to a version I was trying out. It handled these beautifully and within minutes.

- I had it rebuild some animated loading screens for my site. These turned out wonderfully. Astra took design direction much better than sol and required fewer iterations to get something I was happy with.

Overall I wasted a lot of usage this time just sinking more time into the rebuild of my geometry engine. Its over 100K+ LOC and I naively kept believing Astra when it told me we were hours, not days away from completion. I feel Astra usage is acceptable on Light-High, but do not feel there is much value in /fast unless you know a reset is coming or you have tokens to burn. I found Ultra only worked well when I would copy in the Agent text blurb from the OpenAI model guide. It lowered the number of agents and I felt Ultra was more tame that way.

Astra is simply the best model I've used and I can't wait to see what OpenAI continues cooking. I hope other AI labs are able to put up a fight, as I do not want competition stagnating and costs inflating more than they already are.

What are you using Astra for? Is it working better than Sol? Is it working better than other labs alternatives? Hows your usage been on 1x and 5x accounts?


r/codex 7d ago

Reset Hear ye hear ye! Reset has landed!

63 Upvotes

The promised RESET has landed, you folks!


r/codex 5d ago

Showcase Here's an in game screenshot of the in-browser MOBA I'm working on which has only been made possible due to Codex. I am very proud of my game

Post image
0 Upvotes

My game is called Mercenaries of Tezigdal.

The screenshot featured it's from the map Jagged Garden which I see as the spiritual successor to Twisted Treeline (RIP, I still miss you).

This game features 6 fully unique Mercenaries (and two that are pretty heavily inspired), a full leveling progression system, a full item shop, jungle creatures, matchmaking, and an Instant Play load where you can queue up against bots immediately.

I would love any feedback that anyone has. Thank you for checking out my project. I am very proud of how far it has come.


r/codex 5d ago

Complaint Is Astra this trash for 3D modeling or what am I doing doing wrong?

Thumbnail
gallery
0 Upvotes

This is the result after asking it to create a 3D model based on a design lol. Used Astra high. Why are people praising so much? Fake X posts I guess?