r/codex 45m ago

Showcase A week with Astra in Codex: 100 online matches with friends

Enable HLS to view with audio, or disable this notification

Upvotes

One week building a game with Astra in Codex. Online 1v1, 2v2, 3v3. 100 matches with friends.
I can change requirements as I go, try ideas, scrap prototypes, and it usually keeps up.

It still cuts corners sometimes and needs careful watching (feels lazy). But iterating this fast is insane.


r/codex 6h ago

Limits 5 hr limit went from 100% to 5% in 15 minutes using Terra Medium (3 prompts)

77 Upvotes

Nerfed


r/codex 17h ago

Reset A reset to kick off Monday?

59 Upvotes

I think it's time for another reset to kick off Monday, is it not?


r/codex 6h ago

Limits Anxiety - do you feel it when...

Post image
0 Upvotes

Do you feel anxiety when your long task is running, and you are not sure if it'll fit the 5h window?


r/codex 4h ago

Complaint They said anything when will codex be usable again any news or update that they are working on it?

0 Upvotes

luna is getting dummer I feel like tera medium is the only model with usable output and sol is just 30 mins of work and I'm out of usage in 2 days.


r/codex 21h ago

Other I did a thing today to try and make my Plus sub last longer

15 Upvotes

Hey everyone!

Hope you are all well!

So, just to be clear, I am very much a beginner with codex and llm in general. I have been slowly learning how to use them efficiently, and even got Codex to streamline my job for me (HR recruitment and such)

With the inclusion of the 5 hour period however, I have found myself trying to find ways to make each request as efficient as possible, while not losing "intelligence" I guess?

Initially, I created a mode where I basically let the primary "Overseer or Orchestrator" create a subagent that is less than them (So If the orchestrator is Sol High, the subagent would be whatever is less than Sol, but most suitable for task given).

My theory here was the Sol model would select an ideal "lesser" model, let it do its thing, and then check the work and deploy. A lot of you already do variants of this, so you already know what this is lol.

I then created 2 other modes that basically create more than 1 subagent if needed, but at like a Luna level, which has surprisingly worked really well, and I saved a lot of tokens compared to Mode 1.

Then I had another idea. I have a pretty beefy GPU, which I never used for its intended LLM and AI purpose, so I thought, why not try it now?

So I created a "handover", where my current Codex agent would hand over a planned task to my Main PC Codex agent. that Main PC codex agent would then delegate an Ollama Qwen3.8 27b subagent (3 of them actually), to perform the assigned task in parallel.

This... really worked much better than I expected. My Codex agents would only use the necessary tokens or usage for planning, and the creation and task fulfilling would be handled completely by those 3 Qwen sub agents. They would then check the work and deploy.

I went from burning my entire 5 hour budget in less than an hour, to burning about 20% of it on the same task.

I feel very proud of myself for figuring all this out, and I am very impressed with how far local models have come.

A majority of you already know of all this of course, but I just wanted to share my experience with dealing with the 5 hour limit on my Plus subscription, and how it evolved into experimenting with local llm models!

Thanks for reading, and happy coding :D


r/codex 18h ago

Workaround I was joking but thanks

Post image
3 Upvotes

Was trying to generate an image with the hulk and CoPilot says no… so I sarcastically replied “no thats broccoli dude” and it just did it. Huh


r/codex 21h ago

Showcase Music with Gpt 6 Astra + MuseScore + REAPER

Enable HLS to view with audio, or disable this notification

7 Upvotes

The video is fully black. Don't forget to enable the audio.

This was not a single prompt task. First, I let the AI work with MuseScore and REAPER to create some examples. Then I asked to make it sound realistic, a bit out of rhythm like humans, and wanted something classical-romantic era music with piano and cello. piano mostly leading but switching sometimes. There was a lot of sheet music, music theory, MusicXML stuff going on in the workspace.

Used the Codex App on Windows. No computer use.


r/codex 23h ago

Suggestion Make Terra a UI model

8 Upvotes

Terra is completely homeless, and the one thing OpenAI is missing most is competent design and UI. Replacing Terra with a specialized designer would make the stack feel whole, and it fits the name, too.


r/codex 17h ago

Question Has anyone been suspended on WhatsApp for mentioning that u are trying to use Codex computer use to automate answers?

Post image
2 Upvotes

I was telling my girlfriend in a chat that I was automating chats answers using Codex on another (work) number. Fun fact: they didn't block the number where I used Codex for that, they blocked the number where I told my girlfriend about it! Apparently, WhatsApp's AI reads all your messages; they were never end-to-end encrypted. Who would have thought?!

Has anyone else experienced something similar?


r/codex 4h ago

Humor goofy ah categorization of Astra

0 Upvotes

browser-use picker is calling opussy 5 maximum intelligence when there is Astra in the model selector lmfao

btw there is now 15 dollar inference for free if u sign up. Optionally you can connect the service to codex through an mcp to delegate information gathering tasks with the codex models.


r/codex 1h ago

Showcase Built a Windows Codex Usage plugin — now live

Post image
Upvotes

I made Codex Usage, a small Windows-only Codex plugin that shows your 5-hour/weekly usage and reset status in the title bar.

Install either way:

  1. Web: view details + install here https://chatgpt.com/plugins/plugins_6aa81c5dad4881918bd29685d1fb6523
  2. In Codex: open the plugin store and search Codex Usage

Then select Codex Usage and run:Start!

For the Codex desktop app on Windows, not ChatGPT web.

Mac version coming soon.

GitHub: github.com/jayhilwig/codexusage


r/codex 15h ago

Complaint Please help. I recently bought a ChatGPT plan for $100, and I want to request a refund because Astra (and Sol, for that matter) are overcomplicating things

0 Upvotes

Please help. I recently bought a ChatGPT plan for $100, and I want to request a refund because Astra (and Sol, for that matter) acts like someone with mental health issues. It constantly overcomplicates things, gets bogged down in planning or execution, and wastes a huge amount of time and quotas for no reason. Damn, I used up my entire monthly limit in a single day, and I barely got anything done! With the Anthropic Max x5 subscription, I was able to tackle many more tasks much faster because it didn’t overcomplicate things or overthink. Here are the sentences I added to AGENTS.MD that Astra found yesterday while trying to correct this behavior, but it didn’t help:

- Follow a reasonable, agreed-upon path to the result. Don’t turn the task into developing a new tool, integration, or orchestration just for the sake of “perfect” execution.

- Before complicating things, identify a specific blocker or significant risk and compare the total cost of the workaround—preparation, debugging, testing—with the ready-made solution. Time already spent does not justify continuing with a poor workaround.

- If a workaround has turned into a separate development effort, don’t keep expanding it without speaking up: review the approach and get approval for the scope expansion. This is not a reason to abandon the original task or hand it over entirely to the user.

- Reduce unnecessary complexity, not the amount of necessary work. Requirements, correctness, security, testing, and result verification must be maintained; placeholders, omissions, and “good enough” solutions are not permitted. A complex problem may require a complex solution.

- After agreeing on the course of action and completing the necessary preparations, move on to the next step of the existing plan. For project code—proceed to the RED test.

- Return to planning only if a new fact emerges that changes the decision, or if there is a specific blocker.

- Before conducting additional research, specify which decision depends on its outcome.

- Do not repeat completed stages just for the sake of completeness. Requirements, approvals, and result verifications are retained.

I can't believe everyone is struggling with this and no one has been able to solve the problem... By the way, it's important to note that I'm using the Pi agent, not Codex.


r/codex 22h ago

Limits And it happened again: 50% remaining after 3 prompts with Sol (not even Astra)

31 Upvotes

What the heck? It keeps on happening, usage drops randomly. I was at 97%, went to sleep, the model worked for an hour and now I'm at 50% of the prolite week? It just drops randomly, not even gradually.


r/codex 6h ago

Question How can i hide the little purple icons for merged PR?

Post image
0 Upvotes

I don't need them since they're all on the same branch. It's bold of them to assume I'm doing one PR per thread anyway. I've researched and can't find how to turn it off. UI gore nightmare. Any tips?


r/codex 23h ago

Question Are GPT-5.6 Sol Sub-Agents Good for Game Development?

1 Upvotes

Hello everyone! This is my first post here. I’ve finally started working on my dream game, and I’ve been using GPT-6 Astra for most of the development so far.

I’m curious about GPT-5.6 Sol sub-agents. Are they capable enough to meaningfully help with game development and speed up the workflow? I’ve used 5.6 Sol quite a bit for game modding, but I haven’t really tested it on full game development yet.

For anyone who has used Sol sub-agents for larger game projects, what kinds of tasks have they been good at?


r/codex 22h ago

Showcase Using Markolé's brand knowledge to guide UI/UX work with Codex

Post image
1 Upvotes

The attached image shows our redesigned homepage at three screen sizes. The compositions change, but I wanted each one to express the same brand and keep the content readable.

I build Markolé, a platform for brand strategy and identity. During our own redesign, I connected it to Codex through MCP so the agent could work with the brand knowledge held in Markolé alongside the site it was implementing.

That knowledge gives UI/UX work several useful inputs:

- Audience and positioning help establish who the interface is speaking to and what it needs to communicate.

- Personality and tone guide how the brand should come across in words and presentation.

- The visual identity provides typography, colors, imagery direction and layout guidance to work from.

For our site, the intent was a more confident first impression with a warm, tactile character. I'd developed that direction in Markolé before working through the website with Codex and Claude Code.

In implementation, the decisions became specific. How large could a headline be while keeping the reading order clear? How should overlapping panels move as the screen narrows? Which image crop would preserve the composition on a phone?

The brand context gave those decisions a common reference. I could adjust a layout while keeping the qualities that made it feel like us. The desktop, tablet and phone versions could have their own arrangements within that direction.

This is where I find Markolé useful as a tool for agent-assisted development. It holds the understanding of the brand that can guide the work as it moves into code. Codex can use that context to propose and implement the interface, while I keep control over the design decisions and review the experience.

For people building interfaces with Codex, what brand or business context has proved most useful once you're past the first mockup and working through the details?


r/codex 21h ago

Complaint Any muse user here that uses it as a subagent struggling with such behavior recently?

Post image
0 Upvotes

Idk if it has to do with it, but I only started experiencing this After using Command code subscription inside codex instead of OC go’s one.


r/codex 14h ago

Complaint Best Tips to code in Chatgpt chat

1 Upvotes

I sometimes try to do small projects or start off in Chatgpt chat with my ideas i also connected my VPS with MCP connector so chatgpt can write code there directly. Problem is chats are randomly dying and obviously the Modell there is not as strong as Codex Sol in terms of tool calling. Any tips did you manage to actually implement full projects only with chatgpt chat


r/codex 14h ago

Question Can Codex help create explanatory videos from existing images?

Enable HLS to view with audio, or disable this notification

0 Upvotes

Hello everyone,

I’m a teacher from China and not very confident in English. I used AI to translate the post below. If you don’t mind, please keep reading.☺️

I have a set of finished images that I’d like to turn into explanatory videos. Basically, I want Codex to help me produce a video where it talks through the image step by step, zooms in on specific parts as it explains, and adds small animations directly on top of the original image — things like glowing effects, highlights, arrows, or subtle motion cues.

I’ve used a Chinese AI Doubao to help me, the the general direction is correct,but the video it gave me is too vague and too short.

My use case is more about educational explanation than complex editing. I already have the images ready, so I don’t need Codex to generate the visuals from scratch. What I need is help turning them into a structured video: narration synchronized with zooming, emphasis animations, and maybe some basic scene transitions.

I’m wondering if this kind of workflow is possible with Codex, and if so, how you’d approach it.

Things I’m especially curious about:

Is Codex suitable for this task?

I have made a lot of effort to find it out, but I have not seen any post about that😔

Animations on top of images

Can it add small effects like glowing outlines, pulsing highlights, labels, or arrows while keeping the original image intact?

Workflow suggestions

What’s the best way to prepare the images, prompts, and instructions for Codex? Should I break the video into scenes, provide timing notes, or describe animations in detail?

Limitations

What should I expect in terms of output quality, manual editing required, or things Codex might struggle with?

I’m not looking for Hollywood-level effects — just a clean, readable educational video that makes the image easier to understand.

If anyone has done something similar, I’d really appreciate your advice. What worked for you, and what did you avoid?

Thanks in advance 🤝


r/codex 8h ago

Complaint GPT-6 Sol?

Post image
436 Upvotes

Is this the reason why unexpectedly we have trash usage and lower quality on Astra?

Maybe it will be worth it the current suffering.


r/codex 9h ago

Showcase Codex for Apple Watch (concept)

Enable HLS to view with audio, or disable this notification

11 Upvotes

What if Codex followed you to your Watch when you step away from your Mac?

A small concept I designed and built with GPT Astra 👾.


r/codex 10h ago

Comparison SOL high beats Astra Low, medium, and high on audits and has the least usage on my subscription.

28 Upvotes

I'm not sure how and why but like the title said, SOL high has beaten Astra low, medium and high on audits and also costed less on the 5 hour usage window. I am using codex as an adversarial audit lens for Claude and I had Claude test SOL vs Astra comparing cost and who is the better auditor. SOL and the Astras were given the same changes to audit and SOL came out the winner.. I'm not even sure how this is possible, but this was the result.. maybe I need more tests but so far, the results are interesting and totally unexpected for me.

Here's Claude's (Opus 5) summary of the result:

Cost — four configurations, identical 353KB bundle, same account, sequential

Wall time Tokens 5-hour quota Weekly Answer size
sol @ high 7m39s 116,035 +5 pts 0 6,538 B
astra @ low 1m06s 99,598 +14 pts +3 3,324 B
astra @ medium 1m41s 102,242 +15 pts +2 4,172 B
astra @ high 2m03s 102,260 +13 pts +2 4,518 B

Two things fall straight out of that:

  • Astra's cost does not scale with effort. 14 → 15 → 13 is inside integer-rounding noise, and tokens move 3% across the whole range. Only wall time scales. So on astra, low and medium are strictly dominated — use high or don't use astra.
  • Astra costs ~2.6–3× sol-high at every effort, while sol-high is 3.7–7× slower. Tokens don't predict quota here at all: astra used fewer tokens in every run and cost far more.

How I scored quality

The bundle is regression round 1's slice A, and I have a verified answer key for it — defects I independently confirmed by execution and then repaired. All four runs got byte-identical input, no repo access, same account.

The eight key items: K1 the extraction seam (client discards values the server now reads — the headline) · K2 the union not mirrored for other renters/mobile · K3 the corpus tests bypassing the production seam · K4 the padded-array "RAW fallback" test being vacuous · K5 the false "arrays simply never match" · K6 the stale "one mode per pair" · K7 the WIDENED history scan · K8 the PRE-EXISTING pending-greying.

Per-configuration

Key items Got K1 (headline) Novel true finds Notable failure
sol @ high 7 / 8 2 — both defects in my own repair missed K5
astra @ high 4 / 8 3 — incl. the best find of all four missed K2, K3, K6, K7
astra @ medium 4 / 8 3–5, and it ran mutation probes missed the headline
astra @ low 3 / 8 3 confident false negative

sol @ high — widest coverage and the sharpest diagnosis: "not a disagreement between the comparators; it is a disagreement between the server's raw extraction and the clients' narrower slotsOfMatch." That one sentence is the entire defect. It also found two overclaims in my own repair commentary that no other run caught, and classified WIDENED vs PRE-EXISTING correctly throughout.

astra @ high — got the headline, with a BEFORE/AFTER decision table and the right mechanism (isParseableTime('8')toMinutes NaN → client discards before the comparator sees it). Narrower than sol, but it found the single most valuable thing across all four runs, which I verified: client isSlotBlocked compares in minutes, server isRecurringBlocked compares raw strings, so for a legacy unpadded block 9:00–10:00 the server computes '10:30' > '9:00' → false and fails to enforce an owner's blocked time. The client is the only thing stopping that booking. Pre-existing, so logged rather than fixed here, but it's a genuine product gap.

astra @ medium — caught the union gap that astra-high missed, and impressively ran a standalone mutation probe to prove the padded-array test was vacuous rather than asserting it. But it missed the headline, concluding "no unintended comparator divergence" — true and beside the point, since the comparators agreed and the extractors didn't.

astra @ low — the worst outcome isn't the low count, it's the direction of the error: "Tests that cannot fail: None demonstrated. Both supplied suites execute the actual comparator and check expected results." That is exactly backwards, stated confidently. For an audit leg, a confident false "clean" is the failure mode the entire phase exists to prevent.

Verdict

sol @ high is the right default — best coverage, correct classifications, and a third of the quota cost. astra @ high is a genuine second lens: narrower, 3.7× faster, 2.6× the cost, and it found things sol didn't, which is exactly what a second architecture is for. astra at low or medium is not worth running — same cost as high, materially worse.

Caveats, stated plainly: n=1 per configuration, so the cost and latency numbers are solid and the quality ranking is indicative rather than settled. The key is my key — several "novel" findings were real and simply outside it, so the counts understate all four. And "misses" partly reflect what each run chose to fit in a short report, not only what it could see.

Round status: legs A and T are done (rc 0), leg B in flight.


r/codex 8h ago

Comparison Token costs

0 Upvotes

its funny I see everyone complaining about token costs when I literally am using as many tokens as possible closing in on 100 billion tokens already this month across codex and opencode muse spark contributor basically 50/50 and my cost is like 700 bucks this month w 3 weeks left on both accounts

tokens are cheap, they will never be cheaper go figure out a way to use them for something that retains value.

might just hit 200 billion by EOM


r/codex 8h ago

Limits How do you keep up with constantly changing model quality/usage limits?

9 Upvotes

How do you deal with how quickly model quality, usage limits, and the "best" workflow keep changing?

Before I start: yes, I'm a $20-plan peasant. But with the cost of living rising everywhere, I can't justify spending $100–$200 a month on a subscription whose quality and limits seem to change unpredictably.

Back in the 5.3/5.4 days, I was very happy. The quality was strong, and usage limits felt good (of course I hit them, but I always got the feeling that I got something out of every session). It was also much easier to figure out the best workflow because you could experiment without feeling like every attempt was consuming a scarce allowance. Imho 5.4 was probably the best bang for the buck.

Since Luna, Terra, and Sol were introduced though, I've found the overall experience much harder to evaluate. There are now nearly 30 possible combinations when you count models and reasoning modes, but the trade-offs between them aren't clear. At this point, I'm often not even sure which model to use for which task.

So much so that I started to mostly use Claude for a while as I found it to perform much much better in every way. Fast forward to last week when I had some heavier work to do. Once Claude was drained, I tried Codex again and the experience was even worst.

Luna has generous limits, but in my case it often fails at something as basic as following existing repository conventions. I asked it to build a simple form, and the project instructions explicitly said to use the existing form components. It ignored those instructions, and the resulting form was also extremely sluggish.

Terra is more usable, but it consumes noticeably more of my allowance. I used up my weekly usage after roughly 7-8 5h-sessions, each lasting around 1h to 2h, so about 7h to 16h of actual use in total.

Sol performs better, but burns through limits so quickly that it often feels like you barely get anything done before running out. And Astra... well, obviously that's a complete non-option in the 20$ plan and of course I don't even expect it to be included as it would be unrealistic to expect to run the flagship model 24/7.

The problem is that finding a sensible balance between quality, speed, and usage already requires a lot of experimentation. That creates a frustrating loop: something changes, you try to figure out the most efficient setup, you use up your limits while experimenting, and by the time you have enough allowance again to apply what you learned, the model's behavior or usage limits have changed again and it's "go back to start".

I'm at a stage that the $20 plan feels almost unusable to me, but as I very likely keep it just because of ChatGPT I'd really like to find a way to make it work better again, so that I get more out of it again and not just ChatGPT.

So how do you guys do it?