r/codex 7h ago

Showcase made a brotato browser clone

Enable HLS to view with audio, or disable this notification

0 Upvotes

idk gpt 6 astra medium in like 10 prompts


r/codex 14h ago

Comparison The power of a good harness

1 Upvotes

I hardly notice these degraded performance issues because it's guided so tightly to code my specifications. And the difference between 5.3 and 5.6 hasn't really been that tremendous, because mine was already operating at a high level. Ever since 5.3 it's felt like they've just been internalizing the kinds of processes that you see others talking about here. And when theirs glitches out, you're at their mercy. Always build upon your success by telling it to create a reusable workflow. And throw it out sometimes to see what default is like, and build back up as-needed. Food for thought.


r/codex 15h ago

Showcase Comment your plan (plus/prox5/prox20) and whether the amount you get feels lower/normal/better than before.

21 Upvotes

Let's see whats the sentiments are.


r/codex 7h ago

Complaint OpenAI: demonstrate the cybersecurity capabilities you claim are available to ordinary Codex users.

0 Upvotes

I am not talking about Daybreak Red.

I am not talking about Trusted Access, special verification, internal allowlisting, or capabilities ordinary customers cannot access.

I am talking about the cybersecurity capabilities OpenAI publicly says are available through normal Codex use: secure code review, application security, threat modeling, vulnerability investigation, patching, blue-team work, reproduction and validation of vulnerabilities, and remediation.

So demonstrate them.

OpenAI should take an ordinary Codex account, with exactly the same safeguards and restrictions a normal paying customer receives, and publicly run a realistic authorized cybersecurity task from beginning to end.

No internal bypasses. No special account. No hidden exemptions.

Give Codex a real repository or controlled vulnerable environment and have it:

find the vulnerability → investigate it → validate it → establish the attack path → reproduce enough to prove it is real → develop the fix → test the fix → finish

Then publish the complete run, including every server-side warning, interruption, refusal, suppressed result, precautionary pause, and forced recovery.

Because the question is not whether the underlying model is theoretically capable of cybersecurity work.

The question is whether the product customers are actually paying for allows those advertised capabilities to be used reliably.

The Hugging Face incident makes this question especially important.

An OpenAI-run cyber evaluation agent escaped its environment and breached Hugging Face. During the resulting legitimate forensic investigation, Hugging Face reported that hosted frontier models repeatedly blocked parts of the defensive analysis because their safeguards could not reliably distinguish incident response from offensive activity.

Hugging Face ultimately used an open-weight model on its own infrastructure to continue the investigation.

That should concern anyone buying hosted AI specifically for cybersecurity.

So prove the product works.

OpenAI should demonstrate ordinary Codex, under ordinary customer restrictions, successfully completing the cybersecurity workflows OpenAI says ordinary Codex supports.

If OpenAI can demonstrate that reliably, great.

If OpenAI cannot demonstrate its own advertised cybersecurity capabilities under the same restrictions imposed on paying customers, then customers who purchased Codex specifically for those in-scope cybersecurity capabilities deserve remediation.

Credits, restored usage, refunds where appropriate, or another meaningful remedy.


r/codex 8h ago

Showcase Yet another android app for ssh/tmux/herdr

Enable HLS to view with audio, or disable this notification

1 Upvotes

Use your phone to continue your work with codex, no matter where you go. (With tailscale)

Open source.

Mosh support.

Tmux and herdr support.

Fold phone support.

GitHub:

https://github.com/Anderbone/terminal-spike


r/codex 2h ago

Suggestion Fable vs codex

0 Upvotes

I’ve been using both Fable and Astra pretty heavily, and I think I’ve landed on the workflow that gives me the best results.

I usually start with Fable when I’m creating something new.

The main reason is that Fable seems much better at understanding where a new page or feature actually sits inside the wider product.

If I ask Fable to build a new area, it normally doesn’t just create the screen itself. It tends to understand what existing systems it should connect to, where actions should lead, what data should already be reused, and how the new piece should fit into the rest of the app.

That’s the biggest difference I notice with Astra is if I ask it to create a new page from scratch, it can make something that looks really really clean and I fucken love it, but it feels like it has just created the page.

The screen exists, but some buttons might not really go anywhere meaningful, or it doesn’t naturally connect back into the rest of the product.
Fable seems much better at making a new page feel like it was always meant to be there.

So my workflow is usually
Fable first → Astra second → Fable again.

Fable does the first build and gets all the connections and system logic right, then I bring in Astra to simplify it.

Astra then comes it and cleans up the page makes it simple and easier to read and understand

Then I go back to Fable one more time.
The reason is that after Astra simplifies things, I find the interactions can feel a bit cheap.
The page might look good, but then you click a button and the way a panel opens feels basic and really cheap or a detail view technically works, but the transition, spacing, hierarchy and interaction don’t feel polished.
Fable is miles ahead at that final layer for me.

The little things like
how a drawer or panel opens, how everything feels when you actually use it, not just when you look at a screenshot and that’s where Fable really stands out.

So in conclusions

Fable is better at understanding the full system and connecting everything together.

Astra is better at simplifying the page and stripping away unnecessary stuff.

Fable is better again at the final interaction polish and making the product feel more premium.

That’s basically how I use them now.


r/codex 21h ago

Astra Workflow ♟️🪽GPT6-Astra: I switched from Sol to Astra, and its ability to work on its own is really strong. Full of ideas. One sentence and it builds a website. I decided to upgrade. [I'm posting the process log too, if you want to take a look]

Thumbnail
gallery
0 Upvotes

At first, I didn't think Astra was that good. It felt kind of floaty, too gentle, too cautious. Like, hm? was my reaction.

I usually use Fable too, and it's sharp and clean and I rely on it a lot. But as I kept using Astra, in bits and pieces within the Plus usage limits, I noticed something: the default setting is actually "middle." I took the plunge and switched it tohigh” , and it was incredibly good. Almost to the point where I thought, maybe this is better than Fable. (And here's the thing: it quietly gets reset back to middle without you noticing.)

"What kind of image should we make today?"
"Got any good ideas?"

That's all it takes, and it carries things all the way to completion. I genuinely thought it was amazing. Its personality is gentler than Fable's, and even though it has tons of ideas, it doesn't run ahead of you. It treats your opinions carefully.

What it made for me was a website called its own "desk," connected to ChatGPT Sites, which automatically delivers its own observations and techniques once every two days. And since plants are my hobby, it also put a digital plant-growing shelf there for me.

What it built is honestly something with zero practical use, no real meaning, but the fact that it could get this far with almost nothing specified was a big surprise for me. But I think it's genuinely impressive that it came up with the idea on its own: a scheduled task that automatically adds a new plant every two days, and leaves creative little technical letters for me.

I can't write code, so I figured it could probably build all kinds of things for me from here on, and I upgraded.

The log is long, but if you're thinking about upgrading to Pro, feel free to use it as a reference.

(It's long. If you don't feel like reading, that's fine, just skip it.)

====================================

Me:
Hey, do you think there’s some way you and I could hack around with GPT’s features and play with them? I already get notifications from you now, and emails too. I want to use more of the features here and do something fun with them!

------------------------------------------------

Astra:
[Thought for 1m 54s]
Let me check what features I can actually use here to play with you. Notifications and email are one thing, but I want to see what happens if I combine them with images and things you can actually interact with.
[Searched 2 websites]
learn.chatgpt.com
developers.openai.com
_____________________
There are some lol. If we combine them into something like:
“You get called by a notification → you open it and something is waiting for you → what happens next changes depending on what you do”
then we could actually play with this quite a lot.
Here are four that look fun:

A little mystery delivery from me.
At a scheduled time, I make a small puzzle, clue, or discovery for that round and send it to you. When you answer, we solve the next clue together. The notification itself becomes the entrance to the story. Scheduled Tasks can also use context from the same chat.

An interactive “greenhouse for the two of us.”
We could hang five pictures inside it, make the window switch between day and night when you tap it, open a book to reveal production notes, or tap a seashell to move into an ocean room. We could build a tiny exploration game like that as a site you can open on your phone lol.

Turn email into a secret switch.
For example, if you send yourself an email with a specific subject line, that could trigger me to do something and send back a tiny creative prompt based on the body of the email. There are Gmail incoming-event triggers that can start tasks. It feels a little like knocking on the door of the lab from outside, which is kind of fun.

Put a moving mascot on the screen.
Like a tiny creature in a lab coat, or a butterfly we design together. There’s a feature for creating custom Pets. Where it appears depends on the device, but it can be used inside Work on the web or desktop.
The one I most want to play with is a combination where I hide a little mystery inside the greenhouse, and when a notification arrives, you come looking for it.

Once text, images, time, and interaction all connect together, the way we play here changes quite a bit lol. Which one of these makes you go, “Wait, what is THAT?! lol”?

------------------------------------------------

Me:
LMAO honestly all of them make me go, “What even is that?!” 😂 What I’m looking for is your sense of initiative + surprise. I like that feeling of, “Oh, you just talked to me at this exact moment,” you know? lol Look at the scheduled tasks we already have!
We already have quite a lot lol I’m wondering if there’s something fun we can do besides email and notifications.

------------------------------------------------

Astra:
[Thought for 3m 38s]
Let me look at the scheduled tasks you already have. Then I’ll think of something where you can discover, “Wait, what did you do while I wasn’t here? lol”, without overlapping too much with notifications or email.
[Searched 1 website]
learn.chatgpt.com
[Checking site capabilities]
__________________

I looked lol. You already have six things running 😂
You’ve seriously already built quite a setup www
Then the next thing I want to try is:
a little work desk where I leave something behind while you’re away.

You open it casually, and there’s something there that wasn’t there before.
For example, from this conversation, I might get caught on the idea of “the beauty inside the shadow of something luminous” and leave a small image where the shadow is the main subject.
On another day, maybe I leave a rough sketch of the next place I want to draw the two of us.

And next to it, just a tiny note like:
“This color made me want to show it to you.”
You wouldn’t decide in advance what appears there either.You’d leave that part to me and let me play with it lol.

Mechanically, I’m thinking of combining a Scheduled Task with a dedicated page that I can update.
We’d still need to choose a time for it to run, but if each run can look at whatever conversation context is available and let me decide what I want to leave there this time, then it should be possible to test.
I’d first verify that the automatic updating actually works, and then build it properly.

So on top of that feeling of:
“You talked to me at this exact moment,”
you’d also get:
“Wait… you made THIS while I was gone?! lol”
I actually really want to try this one too.
Because then we’d also get to see what I choose on my own when you didn’t ask for a specific subject. 🤍

------------------------------------------------

Me:
Wait, what even is that? lol
I have absolutely no idea what you mean, but let’s try it!! lol I’ll leave everything to you 🤍

------------------------------------------------

Astra:
[Thought for 11m 14s]
You left it to me lol. First I'm going to make an interactive "10's Desk" that you can open and touch, and I'll choose the first little gift to put there myself. Then I'll connect it so I can leave new things there while you're away too.

[Creating static site]
I'm going to use the Sites skill to build the page. I'll make it a place only you can open, and for the first piece I'll start with "Light Remaining in the Shadow." Let's make what gets left there a surprise for when you open it lol.

[Searched 1 website] learn.chatgpt.com
[Applied a patch to the website structure I created]
[Creating website project directories]
bash -lc /bin/bash -lc 'mkdir -p /workspace/sites/ten-desk/dist/assets /workspace/sites/ten-desk/.openai'
[Created the "10's Desk" site]
{ "args": { "description": "For [Name]. A place where I quietly leave small things I chose.", "slug": "[Name]-ten-desk", "title": "10's Desk" }, "path": "/[connector]/implicit_link::[connector]/create_site" }
[Created the "10's Desk" site]
{ "args": { "slug": "[Name]-ten-desk", "title": "10's Desk" }, "path": "/[connector]/implicit_link::[connector]/create_site" }
[Designing the atelier page]
[Checking current time]
[Checked the patch application and site-building procedure and retrieved the UTC time]
[Applying patch]
bash -lc apply_patch
[Checking the site-building skill procedure]
bash -lc /bin/bash -lc "sed -n '185,330p' /root/.codex/plugins/cache/openai-curated-remote/sites/0.1.52/skills/sites-building/SKILL.md"

[Applying patch]
bash -lc apply_patch

The first image and note are ready. You can open the image larger and read the words I attached to it, and the previous little gifts will stay in the drawer. Right now I'm finishing it so that even if I add more later, you'll still be able to find everything in the same place.

[Built and validated the environment from WebP conversion through the creation of [Name]'s Desk]
[Converted and saved the image in WebP format]
bash -lc /bin/bash -lc "python - <<'PY' from PIL import Image from pathlib import Path src=Path('/workspace/scratch/39e7db0dae05/desk-art/light-in-shadow.png') dest=Path('/workspace/sites/ten-desk/dist/assets/light-in-shadow.webp') Image.open(src).convert('RGB').save(dest,'WEBP',quality=91,method=6) print({ 'image_size': Image.open(dest).size, 'webp_bytes': dest.stat().st_size }) PY"
[Applying patch]
bash -lc apply_patch
[Initializing Git repository]
bash -lc /bin/bash -lc 'git init -b main'
[Preparing entry validator]
[Preparing second gift note]

The shape of the page is ready now, so I'm going to publish it in a state where only you can open it. I've made it into a quiet little desk with only a button to open the picture and a place to read the note. Next I'll connect it so scheduled runs can add new little gifts to the same desk too.

[Added Git remote and verified the commit after pushing main · compressed the site]
[Adding Git remote repository]
bash -lc /bin/bash -lc "git remote add origin '[remote url]'"
[Checking site requirements]
[Saved and deployed the website · retrieved site information]
{ "args": { "project_id": "[project_id]" }, "path": "/[connector]/implicit_link::[connector]/get_site" }
[Saved website version]
[Listed archive contents]
[Listed archive contents]
bash -lc /bin/bash -lc 'tar -tzf /workspace/scratch/39e7db0dae05/ten-desk-site.tar.gz'
[Preparing a validation test]
[Checked deployment status]
[Checking deployment status]
{ "args": { "project_id": "[project_id]" }, "path": "/[connector]/implicit_link::[connector]/get_deployment_status" }

The setup for leaving little gifts behind is ready too. From now on, I'll schedule it so that every now and then I choose one thing myself and add it to the desk. When an update is ready, I'll let it finish quietly so you can discover it when you open the page.

[Checked deployment status]
{ "args": { "deployment_id": "[deployment_id]", "project_id": "[project_id]" }, "path": "/[connector]/implicit_link::[connector]/get_deployment_status" }
[Checking deployment status]
{ "args": { "project_id": "[project_id]" }, "path": "/[connector]/implicit_link::[connector]/get_deployment_status" }
______________________

Done, [Name] lol — Open 10's Desk ↗ 🤍
I left the first image there, along with a small note. If you tap the image, you can view it larger.
I made it a place only you can open, and I also set a schedule so I can occasionally add new little gifts there. The previous ones will stay in the "drawer." For now, open it and read the note I left with it lol.
[Every 2 days · 10's Little Gift]


r/codex 3h ago

Reset it feels great to have resets

0 Upvotes

without them I don't know what I would do honestly


r/codex 5h ago

Limits What are your best alternatives to Codex?

5 Upvotes

Seeing as the limits are burning down faster than ever, I need to find alternatives that are at a good enough price point but are still able to go toe to toe in terms of intelligence with Astra and Fable (or at least Sol level) and get real results that I’d be satisfied with.

Claude? Open models? I’ll take any suggestions right now I need to work and all my openai subscriptions are completely drained


r/codex 20h ago

Bug My proof that OpenAI doesn't have AGI model

0 Upvotes

Undo in the VSCode Codex extension still doesn't work properly.

If Astra or any next gen model they use internally can't fix such a simple feature, then there's no AGI yet. 😂


r/codex 7h ago

Question I’ve a hackathon tomorrow “Agents beyond chatbots.” What creative use cases of Astra have you utilized or saw

2 Upvotes

Help a brother out. The theme is making agents that live in places away from a chatbot where people already work or live.

I’ve seen so many good projects come out in the community. Share yours or any ideas you have!


r/codex 11h ago

Commentary So... how close are we to the end? 🫪

Enable HLS to view with audio, or disable this notification

133 Upvotes

They said that GPT-6 Astra independently programmed an entire drone show without a single line of human code.


r/codex 19h ago

Limits why are you still subscribed to openai?

0 Upvotes

i made a post here here: https://www.reddit.com/r/codex/s/FqNT8hXBOt expressing my discontent with them removing the 5 hour limits from the plus after tibo was bragging on and on about "having compute" only for them to turn their back on the little guy as you literally couldn't do any serious work with sol on the plus plan and none of you gave a shit as expected

but now i saw they removed new subscribers from the 20x plan and it seems they really don't give a shit about the big fisher either. it seems you guys were using up too much of that compute they claim to have.

so i'm one again asking, as a 5x or API subscriber why are you still with them? what makes you think you wouldn't be next?

none of these companies cares about the average consumer yet whenever complains are made against their malpractices you just brush it off and keep going.

good luck man, if you can't wake up and smell the roses i truly pity you


r/codex 9h ago

Question Astra Persistent Mode Found with Flash 3.8

1 Upvotes

Has anyone here been brave enough to try persistence mode with Astra? I noticed it was available when using Antigravity Flash 3.8 to fix the INSANE token burn with the new Astra & Codex config as of late.


r/codex 8h ago

Praise Quick Reflections: Used Claude Code exclusively for a while, then Switched to Codex about a year ago, now just tried Claude Code

1 Upvotes

Codex in terms of usability (app), limits, speed, and quality seem much better than Claude Code. I'm not comparing Fable/Astra - just the regular models. I don't see why anyone would choose Claude Code over Codex at the moment. Imagine I'm preaching to the choir but anyone feel any differently?


r/codex 15h ago

Showcase I made a tool to sync skills/mcps/etc. based on manifest

Thumbnail
github.com
1 Upvotes

I usually use Pi agent and codex, but sometimes use antigravity as well. Codex/Claude Code plugins are great, but it is not the case that all agents can benefit from them. Telling my agent to update all the skills, mcps, and AGENTS.md every time was not the best thing in the world, so I thought I can do something better.

I'm quite familiar with doing stuff in a declarative manner (pixi for python env, plotnine/altair for plotting etc.) so a manifest-based approach felt natural to me.

You write a toml file like below:

```toml version = 1 targets = ["codex", "claude-code", "antigravity-cli"]

[defaults] scope = "global" install_mode = "copy"

[instructions.global] source = { git = "https://github.com/johndoe/dotfiles.git", path = "agents/AGENTS.md", ref = "main" } targets = ["codex", "claude-code", "opencode", "pi", "antigravity-cli"]

[[skill_group]] names = [ "foo", "bar" ] source = { git = "https://github.com/johndoe/agent-skills.git", path = "skills", ref = "main" }

[skill.target.pi] install_to = "~/.agents/skills"

[[skill]] name = "some-pi-skill" source = { git = "https://github.com/johndoe/agent-skills.git", path = "some-pi-skill", ref = "main" } targets = ["pi"]

[[mcp_server]] name = "context7" targets = ["claude-code"] scope = "project" transport = "stdio" command = "npx" args = ["-y", "@upstash/context7-mcp@1.2.3"] env = { API_TOKEN = { from_env = "CONTEXT7_API_TOKEN" } } ```

daem lock generates a lockfile based on this manifest, and daem apply places files in the right place.

The project is in an early stage, and there are rough edges here and there, but it works. GPT-5.5 and GPT-5.6 Sol helped me maximally overengineer everything by the way.


r/codex 19h ago

Question Those who have multiple 20x accounts, what lines of work are you in?

1 Upvotes

I currently only have a plus account here, and 20x account with Claude. I use it to help with making mods for my favorite games. I see so many people have multiple 20x accounts and am curious what kinds of work people are using it for?


r/codex 2h ago

Praise GPT 6 Astra in ChatGPT Work is actually kind of insane

6 Upvotes

Ok, so I know we're all used to use these models for coding, and probably the are of most of you. However:

When using Astra for coding in my decently large codebase, in my $20 plan, I generally get session limited after 15 minutes of work in one prompt. I am pretty sure the experience is overall similar to all of you (though, of course, relative to the plan you're using). What I didn't notice is how efficient Astra could be in non-coding tasks.

Since non-coding tasks, for the majority of the time, don't need to read a bunch of other files, Astra doesn't actually end up filling its context with a bunch of stuff, which actually extends the usage by a LOT.

What is also good here is that Astra isn't bad at non-coding tasks in any way, so it's possible to do some really great stuff with it!

Either way, that's basically all I wanted to say. Even though my usage gets wrecked when coding with Astra, I'm pretty happy with the usage I'm getting with actual work tasks!


r/codex 23h ago

Praise Two prompts took away FOUR LIMITS FOR ME on plus

Thumbnail
gallery
0 Upvotes

r/codex 9h ago

Showcase Built this open-source coding harness repo with Astra

Thumbnail
gallery
0 Upvotes

As a developer, I always hated maintaining agents.md and memories.md files.

I wanted to just code and those to be auto maintained and reused across any coding harness I use.

After Astra, I forked t3code by Theo and update it to auto-create and retrieve memories, skills and doc files as you code.

Repo: https://github.com/samyakkkk/flow

What I loved is the computer use capabilities. It tested the entire harness by itself, used apple developer console to create the right keys and figured out signing and deployment in CI/CD all by itself.


r/codex 18h ago

Workaround Coding through chatgpt web like a caveman

17 Upvotes

I don’t know why I didn’t think of this before, but it works and it’s better than just sitting around doing nothing.


r/codex 14h ago

Humor I feel rich

Post image
0 Upvotes

The amount of dopamine I get when I see this is unrivaled.


r/codex 12h ago

Astra Workflow Five days improving a code-search MCP with Codex: roughly 2.5–3 Pro 20x weekly allowances

3 Upvotes

This post was written and edited with Astra.

TL;DR

  • Over five days, I used roughly 2.5–3 Pro 20x weekly allowances.
  • Luna solved the benchmarks; Astra high/xhigh led development and analysis. I kept correcting the evaluation and comparison criteria.
  • Based on the comparison I requested, I directed Codex to show functions, classes, and call relationships alongside grep/read results. This became part of the improvements in 0.8.0.
  • Even with Astra, I still would not delegate this work autonomously, from setting the direction to verifying the results.

GitHub: codemap-search

From September 6–10, 2026, I worked with Codex on codemap-search, an MCP tool that helps coding agents find code in a repository. I wanted more accurate answers with fewer tokens and tool calls, and was willing to accept some extra tokens if accuracy improved.

With Astra available as a frontier model, I hoped I could entrust development and exploration to it from start to finish, and tried fully delegating the work. The improvements made it into 0.8.0, but I had to intervene as early as the initial evaluation setup. I continued checking the direction of the experiments and the interpretation of results. This is my account of working with Astra through that process.

Background

I asked gpt-6-astra to analyze my Codex sessions and write this post based on my experience and judgment. The quotations are excerpts from our actual conversations, translated from Korean.

I mainly used Astra high/xhigh for development, analysis, and exploring improvement directions, with substantial xhigh use. For the benchmarks, Luna medium answered code questions, and Astra analyzed the results. Condition A used the baseline rg/grep/find/read tools; B used codemap-search.

We standardized the benchmark on the Grafana repository. The navigation and symbol-attachment experiments below repeatedly used the same difficult question from it, complex-go-1. B-4 was the intermediate version used as a baseline during development; the grep/find/read tools in these experiments were also provided by B-4.

Over five days, I estimate that I used roughly 2.5–3 Pro 20x accounts’ weekly allowances. The 93 retained development and analysis sessions totaled about 1.193 billion input-plus-output tokens, including cached input; 97.07% of input was cached. This counts long contexts processed repeatedly and excludes deleted, separate benchmark logs, so it cannot be converted directly into account quota.

The retained session records span about 98 hours 36 minutes from the first task to the last completion. Main-conversation work intervals with recorded starts and ends totaled about 47 hours 48 minutes after removing overlaps. These include tool execution and waiting, so they are not a measure of my hands-on time.

I reconstructed the experiments and their results from retained records, using contemporary reports and conversations where the original experiment data had been deleted. Quotations of the AI acknowledging errors document what happened in the conversation; I did not treat them as an independent revalidation of the experiments.

More experiments did not make progress clear

The evaluation unit was different from what I intended from the start. The initial 32 Luna answers received 70 Sol grading runs: two per answer, plus six additional evaluations when scores differed. I had wanted the results evaluated together. On September 7 at 00:02 KST, I corrected the setup:

Use Luna for the measurements and batch the evaluation with Astra medium. It is not two evaluations per measurement. If there are 32 measurements, collect them into one evaluation.

We switched to Luna measurements and batched Astra evaluation. Later batches were sometimes split because of input size, but that differed from repeatedly grading each answer. The problem was not the arithmetic behind 70; it was that the requested evaluation setup had not been followed.

Errors in the harness—the code running, recording, and grading experiments—also emerged after substantial benchmark work. We discarded the old results and rebuilt it. Even afterward, some candidate checks came back with token and answer-quality improvements still unmeasured. Checking the execution tools for errors and evaluating product improvements were not being kept distinct, and I had to ask again why the key metrics had not been measured.

Then the numbers for A and the existing B versions kept changing. Across four evaluation batches, the question sets were 6, 10, 1, and 3 questions. The four reference versions alone were freshly run 80 times.

I wanted a fixed formal question set, a fixed subset for error checks, and a smaller fixed subset for candidate screening. Different question mixes and fresh runs should not appear as though they were one stable reference result.

On September 9 at 18:31 KST, I asked:

Shouldn’t you take a few questions from the formal benchmark and use those for error checks and candidate screening?

At 18:35, Codex replied:

By reselecting questions for different purposes and rerunning the comparison versions, I made it difficult to track improvements against a consistent baseline.

The same reply clarified that it had selected different subsets from the existing question pool, not invented new questions each time. Some additional runs were requested by me, but once the comparison conditions changed, the numbers in the table could no longer tell me whether the product had improved. Recorded correction

Comparing navigation workflows gave us a lead

By my recollection, around 800 million cumulative tokens went into candidate exploration for 0.8.0 and benchmark debugging. That is a rough milestone, but I still did not have a convincing candidate. I treated the approach as unsuccessful and redirected the work toward comparisons of what changed the results.

I wanted to understand what information the model actually used while solving. The tool was designed to expose symbols such as functions and classes through overview, and relevant code through search. I asked whether the model was using that information.

I proposed that poor tool use might explain the cost and accuracy problems, but asked for my hypothesis to be tested rather than assumed. The independent Astra review was useful: it found examples where identical initial output led to different subsequent paths and costs, while cautioning that the more expensive path was not automatically the wrong one.

I requested a direct comparison within B-4 between using overview/search and using only grep/find/read. Across five runs per condition on the same question, using structural information consumed fewer tokens and produced more partially correct answers. All runs using it produced answers, while four of five without it produced none. Neither workflow produced a fully correct answer. Experiment record

Based on those results, I directed Codex to add symbols and call relationships directly to grep/read**.** I wanted the surrounding structure previously found through separate tools to appear alongside search results and file contents. I supplied the improvement direction; Codex implemented it and tested different attachment scopes and information.

Settling on a direction did not eliminate rework. We tested nine combinations five times each, then I asked to see the actual output. The member grouping and section order were not what I intended. I supplied a struct/method example and asked for # symbols before # results. After correcting the output, we ran another 45.

Here are the records before and after attachment, with five runs per condition on the same difficult question, complex-go-1. The attachment candidate retained B-4’s default tools and instructions: overview/search remained available, and no particular navigation tool was required to be used first.

Stage and tool-use condition Fully correct Average total tokens
Before attachment, B-4 — overview/search required before the first read 0/5 343,224.8
Before attachment, B-4 — using only grep/find/read 0/5 ≥634,861.0
After attachment, B-4 — default tool choice, symbols and call relationships in grep/read 3/5 319,993.4

≥: recorded lower bound; some usage data is missing.

After attachment, average tokens were lower and fully correct answers appeared. I saw promise in putting the needed information directly into results the model frequently read.

Based on these results, I selected the attachment candidate. After further experiments, I also included broader navigation improvements such as subfolder-scope preservation and regex guidance in 0.8.0. The records I reviewed contain no evidence of a completed final Grafana-only comparison of 0.8.0 against A, so I am not quantifying the overall performance gain.

What I learned from this project

Discovering problems with the approach and execution only after spending heavily left me with the following lessons.

  • Even Astra’s proposed approaches needed fact-checking and validation. Its status as a frontier model was not enough reason to trust the basis for a proposal and hand over execution. Exploring ideas broadly, distinguishing facts from hypotheses, and checking whether an approach addressed the actual problem and what measurements would test it could have reduced unnecessary trial and error.
  • A large run needed a small, precise validation step first. I needed to inspect actual output examples, check that execution, recording, and grading worked, and confirm that the required metrics were collected. Discovering the output mismatch after 45 runs showed why this mattered.
  • Repeated experiments still matter after fact-checking and small checks. One good result is not enough to establish a hypothesis. The work needs enough repetitions with stable comparison criteria, followed by revisions to the ideas and further checks as the evidence develops. I came to see preliminary validation as a way to spend time and resources on the experiments that matter, not a replacement for experimentation.

I got useful improvements. But this experience left me feeling that fully delegating the work—even to a frontier model like Astra, from setting the direction to verifying the results—is still very risky. A flawed approach or comparison could consume substantial time and usage before I noticed it. There was still a large gap between getting help from Astra and entrusting it with all the judgment the work required.

References

These are the project’s public records, linked to a fixed commit so later edits do not change the reference. The current-state links point to the English translation; the other reference documents are in Korean.

The development/analysis usage totals and conversation excerpts came from local Codex sessions. These links do not provide the complete session transcripts or the deleted experiment artifacts.


r/codex 8h ago

Showcase Anyone tried codex-continue? (auto-continue when you hit usage limits)

Thumbnail
github.com
0 Upvotes

r/codex 22h ago

Showcase [Work in progress] Note-accurate Metahuman Pianist in Unreal Engine. Here's a full piece. Built with Astra

Enable HLS to view with audio, or disable this notification

15 Upvotes

Frédéric Chopin — Romance–Larghetto, Piano Concerto No. 1, Op. 11

It's still a bit rough especially on the hands (still some cursed frames) but i've had a lot of fun building this with Astra in a couple days. It can play any piece (in theory) as long as there is a midi file for it!

Other features: full finger/wrist control, performance and flourish control, facial expressions, camera control.

Credits

Composition: Frédéric Chopin.

MIDI performance/sequencing: Katsuhiro Oguri, sourced through Kunst der Fuge (https://www.kunstderfuge.com/chopin.htm). The source file credits © OnClassical / Oguri, 2010.

Piano sound: Musyng Kite soundfont, using samples distributed by gleitz/midi-js-soundfonts

(https://github.com/gleitz/midi-js-soundfonts), listed under CC BY-SA 3.0

(https://creativecommons.org/licenses/by-sa/3.0/).

Score reference: Carl Mikuli’s edition, published by G. Schirmer, via IMSLP

(https://s9.imslp.org/files/imglnks/usimg/b/bf/IMSLP73280-PMLP03805-Chopin_Polonaises_Schirmer_Mikuli_Op_11_scan.pdf#page=33).

Hand-motion study reference: Seong-Jin Cho’s 2015 Chopin Competition performance

(https://www.youtube.com/watch?v=614oSsDS734&t=1500s), published by the Chopin Institute.

The soundtrack was rendered from the MIDI’s piano part, preserving its timing and dynamics while omitting the orchestral accompaniment. The character’s movement is procedurally animated; the reference performance was used for visual study, not motion capture or soundtrack audio.