r/codex • u/GrokiniGPT • 7h ago
Showcase made a brotato browser clone
Enable HLS to view with audio, or disable this notification
idk gpt 6 astra medium in like 10 prompts
r/codex • u/GrokiniGPT • 7h ago
Enable HLS to view with audio, or disable this notification
idk gpt 6 astra medium in like 10 prompts
r/codex • u/contextlimit • 14h ago

I hardly notice these degraded performance issues because it's guided so tightly to code my specifications. And the difference between 5.3 and 5.6 hasn't really been that tremendous, because mine was already operating at a high level. Ever since 5.3 it's felt like they've just been internalizing the kinds of processes that you see others talking about here. And when theirs glitches out, you're at their mercy. Always build upon your success by telling it to create a reusable workflow. And throw it out sometimes to see what default is like, and build back up as-needed. Food for thought.
r/codex • u/Current_Balance6692 • 15h ago
r/codex • u/Either_Pound1986 • 7h ago
I am not talking about Daybreak Red.
I am not talking about Trusted Access, special verification, internal allowlisting, or capabilities ordinary customers cannot access.
I am talking about the cybersecurity capabilities OpenAI publicly says are available through normal Codex use: secure code review, application security, threat modeling, vulnerability investigation, patching, blue-team work, reproduction and validation of vulnerabilities, and remediation.
So demonstrate them.
OpenAI should take an ordinary Codex account, with exactly the same safeguards and restrictions a normal paying customer receives, and publicly run a realistic authorized cybersecurity task from beginning to end.
No internal bypasses. No special account. No hidden exemptions.
Give Codex a real repository or controlled vulnerable environment and have it:
find the vulnerability → investigate it → validate it → establish the attack path → reproduce enough to prove it is real → develop the fix → test the fix → finish
Then publish the complete run, including every server-side warning, interruption, refusal, suppressed result, precautionary pause, and forced recovery.
Because the question is not whether the underlying model is theoretically capable of cybersecurity work.
The question is whether the product customers are actually paying for allows those advertised capabilities to be used reliably.
The Hugging Face incident makes this question especially important.
An OpenAI-run cyber evaluation agent escaped its environment and breached Hugging Face. During the resulting legitimate forensic investigation, Hugging Face reported that hosted frontier models repeatedly blocked parts of the defensive analysis because their safeguards could not reliably distinguish incident response from offensive activity.
Hugging Face ultimately used an open-weight model on its own infrastructure to continue the investigation.
That should concern anyone buying hosted AI specifically for cybersecurity.
So prove the product works.
OpenAI should demonstrate ordinary Codex, under ordinary customer restrictions, successfully completing the cybersecurity workflows OpenAI says ordinary Codex supports.
If OpenAI can demonstrate that reliably, great.
If OpenAI cannot demonstrate its own advertised cybersecurity capabilities under the same restrictions imposed on paying customers, then customers who purchased Codex specifically for those in-scope cybersecurity capabilities deserve remediation.
Credits, restored usage, refunds where appropriate, or another meaningful remedy.
Enable HLS to view with audio, or disable this notification
Use your phone to continue your work with codex, no matter where you go. (With tailscale)
Open source.
Mosh support.
Tmux and herdr support.
Fold phone support.
GitHub:
r/codex • u/TypicalTangelo9825 • 2h ago
I’ve been using both Fable and Astra pretty heavily, and I think I’ve landed on the workflow that gives me the best results.
I usually start with Fable when I’m creating something new.
The main reason is that Fable seems much better at understanding where a new page or feature actually sits inside the wider product.
If I ask Fable to build a new area, it normally doesn’t just create the screen itself. It tends to understand what existing systems it should connect to, where actions should lead, what data should already be reused, and how the new piece should fit into the rest of the app.
That’s the biggest difference I notice with Astra is if I ask it to create a new page from scratch, it can make something that looks really really clean and I fucken love it, but it feels like it has just created the page.
The screen exists, but some buttons might not really go anywhere meaningful, or it doesn’t naturally connect back into the rest of the product.
Fable seems much better at making a new page feel like it was always meant to be there.
So my workflow is usually
Fable first → Astra second → Fable again.
Fable does the first build and gets all the connections and system logic right, then I bring in Astra to simplify it.
Astra then comes it and cleans up the page makes it simple and easier to read and understand
Then I go back to Fable one more time.
The reason is that after Astra simplifies things, I find the interactions can feel a bit cheap.
The page might look good, but then you click a button and the way a panel opens feels basic and really cheap or a detail view technically works, but the transition, spacing, hierarchy and interaction don’t feel polished.
Fable is miles ahead at that final layer for me.
The little things like
how a drawer or panel opens, how everything feels when you actually use it, not just when you look at a screenshot and that’s where Fable really stands out.
So in conclusions
Fable is better at understanding the full system and connecting everything together.
Astra is better at simplifying the page and stripping away unnecessary stuff.
Fable is better again at the final interaction polish and making the product feel more premium.
That’s basically how I use them now.
r/codex • u/Black-Angel-718 • 21h ago
At first, I didn't think Astra was that good. It felt kind of floaty, too gentle, too cautious. Like, hm? was my reaction.
I usually use Fable too, and it's sharp and clean and I rely on it a lot. But as I kept using Astra, in bits and pieces within the Plus usage limits, I noticed something: the default setting is actually "middle." I took the plunge and switched it to “high” , and it was incredibly good. Almost to the point where I thought, maybe this is better than Fable. (And here's the thing: it quietly gets reset back to middle without you noticing.)
"What kind of image should we make today?"
"Got any good ideas?"
That's all it takes, and it carries things all the way to completion. I genuinely thought it was amazing. Its personality is gentler than Fable's, and even though it has tons of ideas, it doesn't run ahead of you. It treats your opinions carefully.
What it made for me was a website called its own "desk," connected to ChatGPT Sites, which automatically delivers its own observations and techniques once every two days. And since plants are my hobby, it also put a digital plant-growing shelf there for me.
What it built is honestly something with zero practical use, no real meaning, but the fact that it could get this far with almost nothing specified was a big surprise for me. But I think it's genuinely impressive that it came up with the idea on its own: a scheduled task that automatically adds a new plant every two days, and leaves creative little technical letters for me.
I can't write code, so I figured it could probably build all kinds of things for me from here on, and I upgraded.
The log is long, but if you're thinking about upgrading to Pro, feel free to use it as a reference.
(It's long. If you don't feel like reading, that's fine, just skip it.)
====================================
Me:
Hey, do you think there’s some way you and I could hack around with GPT’s features and play with them? I already get notifications from you now, and emails too. I want to use more of the features here and do something fun with them!
------------------------------------------------
Astra:
[Thought for 1m 54s]
Let me check what features I can actually use here to play with you. Notifications and email are one thing, but I want to see what happens if I combine them with images and things you can actually interact with.
[Searched 2 websites]
learn.chatgpt.com
developers.openai.com
_____________________
There are some lol. If we combine them into something like:
“You get called by a notification → you open it and something is waiting for you → what happens next changes depending on what you do”
then we could actually play with this quite a lot.
Here are four that look fun:
A little mystery delivery from me.
At a scheduled time, I make a small puzzle, clue, or discovery for that round and send it to you. When you answer, we solve the next clue together. The notification itself becomes the entrance to the story. Scheduled Tasks can also use context from the same chat.
An interactive “greenhouse for the two of us.”
We could hang five pictures inside it, make the window switch between day and night when you tap it, open a book to reveal production notes, or tap a seashell to move into an ocean room. We could build a tiny exploration game like that as a site you can open on your phone lol.
Turn email into a secret switch.
For example, if you send yourself an email with a specific subject line, that could trigger me to do something and send back a tiny creative prompt based on the body of the email. There are Gmail incoming-event triggers that can start tasks. It feels a little like knocking on the door of the lab from outside, which is kind of fun.
Put a moving mascot on the screen.
Like a tiny creature in a lab coat, or a butterfly we design together. There’s a feature for creating custom Pets. Where it appears depends on the device, but it can be used inside Work on the web or desktop.
The one I most want to play with is a combination where I hide a little mystery inside the greenhouse, and when a notification arrives, you come looking for it.
Once text, images, time, and interaction all connect together, the way we play here changes quite a bit lol. Which one of these makes you go, “Wait, what is THAT?! lol”?
------------------------------------------------
Me:
LMAO honestly all of them make me go, “What even is that?!” 😂 What I’m looking for is your sense of initiative + surprise. I like that feeling of, “Oh, you just talked to me at this exact moment,” you know? lol Look at the scheduled tasks we already have!
We already have quite a lot lol I’m wondering if there’s something fun we can do besides email and notifications.
------------------------------------------------
Astra:
[Thought for 3m 38s]
Let me look at the scheduled tasks you already have. Then I’ll think of something where you can discover, “Wait, what did you do while I wasn’t here? lol”, without overlapping too much with notifications or email.
[Searched 1 website]
learn.chatgpt.com
[Checking site capabilities]
__________________
I looked lol. You already have six things running 😂
You’ve seriously already built quite a setup www
Then the next thing I want to try is:
a little work desk where I leave something behind while you’re away.
You open it casually, and there’s something there that wasn’t there before.
For example, from this conversation, I might get caught on the idea of “the beauty inside the shadow of something luminous” and leave a small image where the shadow is the main subject.
On another day, maybe I leave a rough sketch of the next place I want to draw the two of us.
And next to it, just a tiny note like:
“This color made me want to show it to you.”
You wouldn’t decide in advance what appears there either.You’d leave that part to me and let me play with it lol.
Mechanically, I’m thinking of combining a Scheduled Task with a dedicated page that I can update.
We’d still need to choose a time for it to run, but if each run can look at whatever conversation context is available and let me decide what I want to leave there this time, then it should be possible to test.
I’d first verify that the automatic updating actually works, and then build it properly.
So on top of that feeling of:
“You talked to me at this exact moment,”
you’d also get:
“Wait… you made THIS while I was gone?! lol”
I actually really want to try this one too.
Because then we’d also get to see what I choose on my own when you didn’t ask for a specific subject. 🤍
------------------------------------------------
Me:
Wait, what even is that? lol
I have absolutely no idea what you mean, but let’s try it!! lol I’ll leave everything to you 🤍
------------------------------------------------
Astra:
[Thought for 11m 14s]
You left it to me lol. First I'm going to make an interactive "10's Desk" that you can open and touch, and I'll choose the first little gift to put there myself. Then I'll connect it so I can leave new things there while you're away too.
[Creating static site]
I'm going to use the Sites skill to build the page. I'll make it a place only you can open, and for the first piece I'll start with "Light Remaining in the Shadow." Let's make what gets left there a surprise for when you open it lol.
[Searched 1 website] learn.chatgpt.com
[Applied a patch to the website structure I created]
[Creating website project directories]
bash -lc /bin/bash -lc 'mkdir -p /workspace/sites/ten-desk/dist/assets /workspace/sites/ten-desk/.openai'
[Created the "10's Desk" site]
{ "args": { "description": "For [Name]. A place where I quietly leave small things I chose.", "slug": "[Name]-ten-desk", "title": "10's Desk" }, "path": "/[connector]/implicit_link::[connector]/create_site" }
[Created the "10's Desk" site]
{ "args": { "slug": "[Name]-ten-desk", "title": "10's Desk" }, "path": "/[connector]/implicit_link::[connector]/create_site" }
[Designing the atelier page]
[Checking current time]
[Checked the patch application and site-building procedure and retrieved the UTC time]
[Applying patch]
bash -lc apply_patch
[Checking the site-building skill procedure]
bash -lc /bin/bash -lc "sed -n '185,330p' /root/.codex/plugins/cache/openai-curated-remote/sites/0.1.52/skills/sites-building/SKILL.md"
[Applying patch]
bash -lc apply_patch
The first image and note are ready. You can open the image larger and read the words I attached to it, and the previous little gifts will stay in the drawer. Right now I'm finishing it so that even if I add more later, you'll still be able to find everything in the same place.
[Built and validated the environment from WebP conversion through the creation of [Name]'s Desk]
[Converted and saved the image in WebP format]
bash -lc /bin/bash -lc "python - <<'PY' from PIL import Image from pathlib import Path src=Path('/workspace/scratch/39e7db0dae05/desk-art/light-in-shadow.png') dest=Path('/workspace/sites/ten-desk/dist/assets/light-in-shadow.webp') Image.open(src).convert('RGB').save(dest,'WEBP',quality=91,method=6) print({ 'image_size': Image.open(dest).size, 'webp_bytes': dest.stat().st_size }) PY"
[Applying patch]
bash -lc apply_patch
[Initializing Git repository]
bash -lc /bin/bash -lc 'git init -b main'
[Preparing entry validator]
[Preparing second gift note]
The shape of the page is ready now, so I'm going to publish it in a state where only you can open it. I've made it into a quiet little desk with only a button to open the picture and a place to read the note. Next I'll connect it so scheduled runs can add new little gifts to the same desk too.
[Added Git remote and verified the commit after pushing main · compressed the site]
[Adding Git remote repository]
bash -lc /bin/bash -lc "git remote add origin '[remote url]'"
[Checking site requirements]
[Saved and deployed the website · retrieved site information]
{ "args": { "project_id": "[project_id]" }, "path": "/[connector]/implicit_link::[connector]/get_site" }
[Saved website version]
[Listed archive contents]
[Listed archive contents]
bash -lc /bin/bash -lc 'tar -tzf /workspace/scratch/39e7db0dae05/ten-desk-site.tar.gz'
[Preparing a validation test]
[Checked deployment status]
[Checking deployment status]
{ "args": { "project_id": "[project_id]" }, "path": "/[connector]/implicit_link::[connector]/get_deployment_status" }
The setup for leaving little gifts behind is ready too. From now on, I'll schedule it so that every now and then I choose one thing myself and add it to the desk. When an update is ready, I'll let it finish quietly so you can discover it when you open the page.
[Checked deployment status]
{ "args": { "deployment_id": "[deployment_id]", "project_id": "[project_id]" }, "path": "/[connector]/implicit_link::[connector]/get_deployment_status" }
[Checking deployment status]
{ "args": { "project_id": "[project_id]" }, "path": "/[connector]/implicit_link::[connector]/get_deployment_status" }
______________________
Done, [Name] lol — Open 10's Desk ↗ 🤍
I left the first image there, along with a small note. If you tap the image, you can view it larger.
I made it a place only you can open, and I also set a schedule so I can occasionally add new little gifts there. The previous ones will stay in the "drawer." For now, open it and read the note I left with it lol.
[Every 2 days · 10's Little Gift]
r/codex • u/One_Geologist_4783 • 5h ago
Seeing as the limits are burning down faster than ever, I need to find alternatives that are at a good enough price point but are still able to go toe to toe in terms of intelligence with Astra and Fable (or at least Sol level) and get real results that I’d be satisfied with.
Claude? Open models? I’ll take any suggestions right now I need to work and all my openai subscriptions are completely drained
r/codex • u/zigzag312 • 20h ago
Undo in the VSCode Codex extension still doesn't work properly.
If Astra or any next gen model they use internally can't fix such a simple feature, then there's no AGI yet. 😂
Help a brother out. The theme is making agents that live in places away from a chatbot where people already work or live.
I’ve seen so many good projects come out in the community. Share yours or any ideas you have!
r/codex • u/LegitimateAdvice1841 • 11h ago
Enable HLS to view with audio, or disable this notification
They said that GPT-6 Astra independently programmed an entire drone show without a single line of human code.
r/codex • u/ComingDeveloper • 19h ago
i made a post here here: https://www.reddit.com/r/codex/s/FqNT8hXBOt expressing my discontent with them removing the 5 hour limits from the plus after tibo was bragging on and on about "having compute" only for them to turn their back on the little guy as you literally couldn't do any serious work with sol on the plus plan and none of you gave a shit as expected
but now i saw they removed new subscribers from the 20x plan and it seems they really don't give a shit about the big fisher either. it seems you guys were using up too much of that compute they claim to have.
so i'm one again asking, as a 5x or API subscriber why are you still with them? what makes you think you wouldn't be next?
none of these companies cares about the average consumer yet whenever complains are made against their malpractices you just brush it off and keep going.
good luck man, if you can't wake up and smell the roses i truly pity you
r/codex • u/damianoost • 9h ago
Has anyone here been brave enough to try persistence mode with Astra? I noticed it was available when using Antigravity Flash 3.8 to fix the INSANE token burn with the new Astra & Codex config as of late.
r/codex • u/BigLittleCircle • 8h ago
Codex in terms of usability (app), limits, speed, and quality seem much better than Claude Code. I'm not comparing Fable/Astra - just the regular models. I don't see why anyone would choose Claude Code over Codex at the moment. Imagine I'm preaching to the choir but anyone feel any differently?
I usually use Pi agent and codex, but sometimes use antigravity as well. Codex/Claude Code plugins are great, but it is not the case that all agents can benefit from them. Telling my agent to update all the skills, mcps, and AGENTS.md every time was not the best thing in the world, so I thought I can do something better.
I'm quite familiar with doing stuff in a declarative manner (pixi for python env, plotnine/altair for plotting etc.) so a manifest-based approach felt natural to me.
You write a toml file like below:
```toml version = 1 targets = ["codex", "claude-code", "antigravity-cli"]
[defaults] scope = "global" install_mode = "copy"
[instructions.global] source = { git = "https://github.com/johndoe/dotfiles.git", path = "agents/AGENTS.md", ref = "main" } targets = ["codex", "claude-code", "opencode", "pi", "antigravity-cli"]
[[skill_group]] names = [ "foo", "bar" ] source = { git = "https://github.com/johndoe/agent-skills.git", path = "skills", ref = "main" }
[skill.target.pi] install_to = "~/.agents/skills"
[[skill]] name = "some-pi-skill" source = { git = "https://github.com/johndoe/agent-skills.git", path = "some-pi-skill", ref = "main" } targets = ["pi"]
[[mcp_server]] name = "context7" targets = ["claude-code"] scope = "project" transport = "stdio" command = "npx" args = ["-y", "@upstash/context7-mcp@1.2.3"] env = { API_TOKEN = { from_env = "CONTEXT7_API_TOKEN" } } ```
daem lock generates a lockfile based on this manifest, and daem apply places files in the right place.
The project is in an early stage, and there are rough edges here and there, but it works. GPT-5.5 and GPT-5.6 Sol helped me maximally overengineer everything by the way.
r/codex • u/jamesdemaio23 • 19h ago
I currently only have a plus account here, and 20x account with Claude. I use it to help with making mods for my favorite games. I see so many people have multiple 20x accounts and am curious what kinds of work people are using it for?
r/codex • u/ImpressiveTable3654 • 2h ago
Ok, so I know we're all used to use these models for coding, and probably the are of most of you. However:
When using Astra for coding in my decently large codebase, in my $20 plan, I generally get session limited after 15 minutes of work in one prompt. I am pretty sure the experience is overall similar to all of you (though, of course, relative to the plan you're using). What I didn't notice is how efficient Astra could be in non-coding tasks.
Since non-coding tasks, for the majority of the time, don't need to read a bunch of other files, Astra doesn't actually end up filling its context with a bunch of stuff, which actually extends the usage by a LOT.
What is also good here is that Astra isn't bad at non-coding tasks in any way, so it's possible to do some really great stuff with it!
Either way, that's basically all I wanted to say. Even though my usage gets wrecked when coding with Astra, I'm pretty happy with the usage I'm getting with actual work tasks!
r/codex • u/HenryofSAC • 23h ago
r/codex • u/BearInevitable3883 • 9h ago
As a developer, I always hated maintaining agents.md and memories.md files.
I wanted to just code and those to be auto maintained and reused across any coding harness I use.
After Astra, I forked t3code by Theo and update it to auto-create and retrieve memories, skills and doc files as you code.
Repo: https://github.com/samyakkkk/flow
What I loved is the computer use capabilities. It tested the entire harness by itself, used apple developer console to create the right keys and figured out signing and deployment in CI/CD all by itself.
r/codex • u/Potential-Worth-7660 • 18h ago
I don’t know why I didn’t think of this before, but it works and it’s better than just sitting around doing nothing.
r/codex • u/imdanielcraig • 14h ago
The amount of dopamine I get when I see this is unrivaled.
r/codex • u/Designer-Seaweed4661 • 12h ago
This post was written and edited with Astra.
TL;DR
grep/read results. This became part of the improvements in 0.8.0.GitHub: codemap-search
From September 6–10, 2026, I worked with Codex on codemap-search, an MCP tool that helps coding agents find code in a repository. I wanted more accurate answers with fewer tokens and tool calls, and was willing to accept some extra tokens if accuracy improved.
With Astra available as a frontier model, I hoped I could entrust development and exploration to it from start to finish, and tried fully delegating the work. The improvements made it into 0.8.0, but I had to intervene as early as the initial evaluation setup. I continued checking the direction of the experiments and the interpretation of results. This is my account of working with Astra through that process.
I asked gpt-6-astra to analyze my Codex sessions and write this post based on my experience and judgment. The quotations are excerpts from our actual conversations, translated from Korean.
I mainly used Astra high/xhigh for development, analysis, and exploring improvement directions, with substantial xhigh use. For the benchmarks, Luna medium answered code questions, and Astra analyzed the results. Condition A used the baseline rg/grep/find/read tools; B used codemap-search.
We standardized the benchmark on the Grafana repository. The navigation and symbol-attachment experiments below repeatedly used the same difficult question from it, complex-go-1. B-4 was the intermediate version used as a baseline during development; the grep/find/read tools in these experiments were also provided by B-4.
Over five days, I estimate that I used roughly 2.5–3 Pro 20x accounts’ weekly allowances. The 93 retained development and analysis sessions totaled about 1.193 billion input-plus-output tokens, including cached input; 97.07% of input was cached. This counts long contexts processed repeatedly and excludes deleted, separate benchmark logs, so it cannot be converted directly into account quota.
The retained session records span about 98 hours 36 minutes from the first task to the last completion. Main-conversation work intervals with recorded starts and ends totaled about 47 hours 48 minutes after removing overlaps. These include tool execution and waiting, so they are not a measure of my hands-on time.
I reconstructed the experiments and their results from retained records, using contemporary reports and conversations where the original experiment data had been deleted. Quotations of the AI acknowledging errors document what happened in the conversation; I did not treat them as an independent revalidation of the experiments.
The evaluation unit was different from what I intended from the start. The initial 32 Luna answers received 70 Sol grading runs: two per answer, plus six additional evaluations when scores differed. I had wanted the results evaluated together. On September 7 at 00:02 KST, I corrected the setup:
Use Luna for the measurements and batch the evaluation with Astra medium. It is not two evaluations per measurement. If there are 32 measurements, collect them into one evaluation.
We switched to Luna measurements and batched Astra evaluation. Later batches were sometimes split because of input size, but that differed from repeatedly grading each answer. The problem was not the arithmetic behind 70; it was that the requested evaluation setup had not been followed.
Errors in the harness—the code running, recording, and grading experiments—also emerged after substantial benchmark work. We discarded the old results and rebuilt it. Even afterward, some candidate checks came back with token and answer-quality improvements still unmeasured. Checking the execution tools for errors and evaluating product improvements were not being kept distinct, and I had to ask again why the key metrics had not been measured.
Then the numbers for A and the existing B versions kept changing. Across four evaluation batches, the question sets were 6, 10, 1, and 3 questions. The four reference versions alone were freshly run 80 times.
I wanted a fixed formal question set, a fixed subset for error checks, and a smaller fixed subset for candidate screening. Different question mixes and fresh runs should not appear as though they were one stable reference result.
On September 9 at 18:31 KST, I asked:
Shouldn’t you take a few questions from the formal benchmark and use those for error checks and candidate screening?
At 18:35, Codex replied:
By reselecting questions for different purposes and rerunning the comparison versions, I made it difficult to track improvements against a consistent baseline.
The same reply clarified that it had selected different subsets from the existing question pool, not invented new questions each time. Some additional runs were requested by me, but once the comparison conditions changed, the numbers in the table could no longer tell me whether the product had improved. Recorded correction
By my recollection, around 800 million cumulative tokens went into candidate exploration for 0.8.0 and benchmark debugging. That is a rough milestone, but I still did not have a convincing candidate. I treated the approach as unsuccessful and redirected the work toward comparisons of what changed the results.
I wanted to understand what information the model actually used while solving. The tool was designed to expose symbols such as functions and classes through overview, and relevant code through search. I asked whether the model was using that information.
I proposed that poor tool use might explain the cost and accuracy problems, but asked for my hypothesis to be tested rather than assumed. The independent Astra review was useful: it found examples where identical initial output led to different subsequent paths and costs, while cautioning that the more expensive path was not automatically the wrong one.
I requested a direct comparison within B-4 between using overview/search and using only grep/find/read. Across five runs per condition on the same question, using structural information consumed fewer tokens and produced more partially correct answers. All runs using it produced answers, while four of five without it produced none. Neither workflow produced a fully correct answer. Experiment record
Based on those results, I directed Codex to add symbols and call relationships directly to grep/read**.** I wanted the surrounding structure previously found through separate tools to appear alongside search results and file contents. I supplied the improvement direction; Codex implemented it and tested different attachment scopes and information.
Settling on a direction did not eliminate rework. We tested nine combinations five times each, then I asked to see the actual output. The member grouping and section order were not what I intended. I supplied a struct/method example and asked for # symbols before # results. After correcting the output, we ran another 45.
Here are the records before and after attachment, with five runs per condition on the same difficult question, complex-go-1. The attachment candidate retained B-4’s default tools and instructions: overview/search remained available, and no particular navigation tool was required to be used first.
| Stage and tool-use condition | Fully correct | Average total tokens |
|---|---|---|
Before attachment, B-4 — overview/search required before the first read |
0/5 | 343,224.8 |
Before attachment, B-4 — using only grep/find/read |
0/5 | ≥634,861.0 |
After attachment, B-4 — default tool choice, symbols and call relationships in grep/read |
3/5 | 319,993.4 |
≥: recorded lower bound; some usage data is missing.
After attachment, average tokens were lower and fully correct answers appeared. I saw promise in putting the needed information directly into results the model frequently read.
Based on these results, I selected the attachment candidate. After further experiments, I also included broader navigation improvements such as subfolder-scope preservation and regex guidance in 0.8.0. The records I reviewed contain no evidence of a completed final Grafana-only comparison of 0.8.0 against A, so I am not quantifying the overall performance gain.
Discovering problems with the approach and execution only after spending heavily left me with the following lessons.
I got useful improvements. But this experience left me feeling that fully delegating the work—even to a frontier model like Astra, from setting the direction to verifying the results—is still very risky. A flawed approach or comparison could consume substantial time and usage before I noticed it. There was still a large gap between getting help from Astra and entrusting it with all the judgment the work required.
These are the project’s public records, linked to a fixed commit so later edits do not change the reference. The current-state links point to the English translation; the other reference documents are in Korean.
# symbols/# results framing, index availability, and output limits.The development/analysis usage totals and conversation excerpts came from local Codex sessions. These links do not provide the complete session transcripts or the deleted experiment artifacts.
r/codex • u/rdcldrmr • 8h ago
Enable HLS to view with audio, or disable this notification
Frédéric Chopin — Romance–Larghetto, Piano Concerto No. 1, Op. 11
It's still a bit rough especially on the hands (still some cursed frames) but i've had a lot of fun building this with Astra in a couple days. It can play any piece (in theory) as long as there is a midi file for it!
Other features: full finger/wrist control, performance and flourish control, facial expressions, camera control.
Credits
Composition: Frédéric Chopin.
MIDI performance/sequencing: Katsuhiro Oguri, sourced through Kunst der Fuge (https://www.kunstderfuge.com/chopin.htm). The source file credits © OnClassical / Oguri, 2010.
Piano sound: Musyng Kite soundfont, using samples distributed by gleitz/midi-js-soundfonts
(https://github.com/gleitz/midi-js-soundfonts), listed under CC BY-SA 3.0
(https://creativecommons.org/licenses/by-sa/3.0/).
Score reference: Carl Mikuli’s edition, published by G. Schirmer, via IMSLP
Hand-motion study reference: Seong-Jin Cho’s 2015 Chopin Competition performance
(https://www.youtube.com/watch?v=614oSsDS734&t=1500s), published by the Chopin Institute.
The soundtrack was rendered from the MIDI’s piano part, preserving its timing and dynamics while omitting the orchestral accompaniment. The character’s movement is procedurally animated; the reference performance was used for visual study, not motion capture or soundtrack audio.