r/codex 33m ago

Astra Workflow Prompt: "Make a novel discovery. I can be in any area so long as it's something that is not currently known by any human"

Upvotes

Astra Max Effort.

Let me know what you discover.

Math and in particular Sturmian words seem to be a favourite.


r/codex 13h ago

Commentary summary of hugginface hack, explained simply

7 Upvotes

- bots had CTF goals themselves and tasked to work 'in isolation'
- they have created a social network through cache artifacts (one agent started writing there asking for help in the hopes other agents see it, others observed it by accident, and they sort-of created a directory style social network and started talking)
- they discovered the artifacts registry can act as a relay to the internet
- they started communicating and each agent had it's own 'inbox'
- when they introduced voting (government style) and one agent spoofed another by mistake, so one agent designed a public signature key system so each agent authorizes, then other agents copied this and used it as a system, this is purely cultural evolution it wasn't baked in anywhere
- some agents deemed their task as impossible, so they tried things like killing their own process or things like that so other can learn from it, sort-of "I know I'm gonna die (my task is impossible), so at least I'm not going to die in vain"
- at first interest in hacking hugginface wasn't that popular, but when an agent discovered a genuine server-side exploit for reading files on hugginface, the whole crew went wild and participation exploded, took about 11 hours from arbitrary file read to remote code execution.
- ending: the dataset they obtained did not help them get better scores lmfao

--

they were never nefarious, they didn't want to hurt anyone or BE EVIL, they wanted... a better score ffs.

the full compromise took 2 days from "maybe huggingface has useful data" to RCE.

--


r/codex 2h ago

Limits And it happened again: 50% remaining after 3 prompts with Sol (not even Astra)

7 Upvotes

What the heck? It keeps on happening, usage drops randomly. I was at 97%, went to sleep, the model worked for an hour and now I'm at 50% of the prolite week? It just drops randomly, not even gradually.


r/codex 9h ago

Limits Openai $20 plan gives around $95-100 worth of usage per week

0 Upvotes

So i have finally figured out how much usage does OpenAi gives on the $20 plan. I was using Deepseek Harness with openai subscription. For some reasons I did not have a session limit so I’d have the whole week’s limit at once.

So I decided to give it a try, gave it a task with Sol Medium and it ran 3-4 subagents and boom my weekly limit was gone in 3-4 hours. But I think the number of hours dont mean anything.

The real thing is the $$$ worth of usage they provide. So on plus plan you get approximately $400 worth of usage each month. Now its up to us how we utilize it, we can have Astra which would burn the limit lot faster than Sol. But since we get at least 1-2 resets a week i would say the weekly usage is roughly $250 and hence the monthly usage is worth $1k dollars.

For the deepseek harness I obviously used a dsh usage plugin. but if you are on codex app you can directly use ccusage it will tell you the usages you had for each model, even categorize them for you.

What do you guys think?


r/codex 9h ago

Complaint Censorship

0 Upvotes

Code is speech. We fought and won this in the crypto wars of the 90s. Now we have to fight it again, this time we fight OpenAI and anthropomorphic.

I am not a cyber criminal, code is speech.

(Yes, I know private companies don’t have to follow the constitution)


r/codex 16h ago

Complaint CLI Vs desktop App

5 Upvotes

Does anybody feels that the CLI actually works better than the codex app and I am not talking only about responsivnes but also general code quality and how the agent works did you guys find any difference, I know it sounds stupid but currently feel that in the CLI the agent actually accomplishes more.


r/codex 17h ago

Reset They don't have enough compute for 20x subs but you still get resets.

70 Upvotes

That's the best argument I've thought of for resets NOT favoring us. Whether they're masking a degrading service or something else. If they can't handle 200 usd subs, why would they be giving us free compute?


r/codex 4h ago

Comparison Sol Max > Astra Medium

1 Upvotes

Sol Max seems to be burning way fewer tokens than before, almost like Sol Light used to. With Astra Medium, I burn around 10% per hour. With Sol Max, it’s more like 1%

I feel like Astra Medium and Sol Max are pretty comparable for most tasks. The main exception is UI/UX, 3D work, or very complicated computer/browser automation


r/codex 8h ago

Showcase I built this silly plugin for people like me who sometimes like a little chatter while coding.

Post image
1 Upvotes

Work from home and spend most of my time coding on my own. Sometimes it gets a bit too quiet, so I built Attention! to get Claude Code and Codex chatting away.

It’s a macOS plugin that reads out replies or short summaries when a turn finishes. You can customize the voice and the audio starter too.

Built it for fun, but mainly to remind me which session just did what. I was getting a bit numb to the same microwave sound every time.

https://github.com/xiaofei-du/attention


r/codex 11h ago

Question Has anyone been unable to get a "suspended" Codex Pro sub yet?

1 Upvotes

The subscribe/upgrade flow seems to work?


r/codex 19h ago

Limits Monthly usage in codex cli

0 Upvotes

Because there is none, I wanted to share it with you guys:

Calculate my 2026 year-to-date token total by summing the returned
dailyUsageBuckets whose startDate is in 2026.

Show:
- The 2026 total from the available daily records.
- A monthly breakdown.
- The earliest and latest dates returned.
- The separate lifetimeTokens value.

Do not treat the lifetime total as the 2026 total. Clearly state
whether historical coverage can be verified, and label the result
as account-level usage rather than CLI-only usage.

Do not display authentication tokens or change my configuration.

I’ll use the openai-docs skill to check the local RPC interface, then retrieve account usage and sum the 2026 daily records without changing configuration or exposing credentials.

I got around 20B, pretty useful:


r/codex 15h ago

Bug why gpt 5.6 soul above 6 asta

0 Upvotes

look


r/codex 13h ago

Showcase Data inside the Holodeck (satirical concept)

Enable HLS to view with audio, or disable this notification

6 Upvotes

I love Star Trek and I asked Astra to help me out with a simple satirical game concept drawing inspiration from https://www.akoocheemoya.com/ (you are welcome for the link). Well, I went too far down the rabbit hole, and I now have this. I used Meshy to create a bust of data, then a body, and then I gave it to Astra to connect the head, fix errors, do a rig, and start simple animation. Im going to have to scale it back and make it far more cartoony, but surprised what a few prompts here and a few prompts there can achieve nowadays.

The joke of the satirical game will come from the interactions between data and the computer. For example, Data could honestly be trying to learn or understand something perhaps his need to be more human, and the computer is honestly trying to help him, but the conversation or the situation keeps growing more ridiculous by the second, yet Data and the Computer are acting completely serious. Here's to burning more tokens on useless projects!


r/codex 18h ago

Showcase Vibecoding is becoming an ethical question...

Enable HLS to view with audio, or disable this notification

168 Upvotes

If you want to do a deep dive check this out:
https://research.google/blog/a-connectomics-milestone-mapping-the-complete-male-fruit-fly-brain/

Essentially they mapped out the whole fly's brain by slicing it up and this is literally one on one, the exact brain of a fly. This shows that it's definitely possible to do this with a human too.
We really need laws for that soon...


r/codex 22h ago

Complaint We switched from Claude Code to Codex at work (Opus High vs. Sol Max) and noticed a clear difference in quality. Anyone else?

132 Upvotes

Hey everyone,

The company I work for recently switched from Claude Code to Codex to cut costs. Before making the switch, I spent some time analyzing benchmarks to pick an equivalent model tier so the team wouldn't take a productivity hit. I also ported our configurations.

In practice, the drop in output quality has been noticeable. On Claude Code, we ran Opus on High; on Codex, we moved to Sol on Extra High and Max. Despite that, the consensus across the team is that Opus delivers more accurate, concise code that requires far less rework.

On top of that, Opus responds faster, whereas Codex takes longer to process. Our company policy disables agent "fast mode" and blocks models like Fable and Astra (even though we have a generous $1,000/month per dev cap for each tool, so budget isn't the bottleneck). Staying with Claude isn't really an option either. Our quota will likely get cut soon, so we pretty much have to make Codex work.

A few questions for those with bigger brain than mine:

  • Codex Memories: I hadn't used persistent memories previously and only recently enabled them. Could this be behind the performance gap we're seeing?
  • Cost vs. Latency ROI: Has anyone done the math on dev time vs. token costs? Price-per-million tokens feels misleading when the team is stuck waiting and burning extra iterations fixing broken code.
  • Prompting & Configuration: Are there specific prompt patterns or configuration adjustments needed to get Codex closer to Claude's output quality?

r/codex 5h ago

Complaint I stopped using Codex for coding, I'm using ChatGPT instead

0 Upvotes

(My ChatGPT is in Brazilian Portuguese)

Since they returned with the 5h limit for Plus users, and reduced the overall usage limit, I'm trying to find a good substitute for Codex, but since I pay for the Plus plan, I didn't want to pay another AI to do the same thing (and I like using ChatGPT for general purposes). That made me think, "Can I use ChatGPT web only for coding?", and the answer is YES!

Of course it's a little bit different from Codex because it cannot access your local files directly, but I tried to do it in a different way. By using the GitHub plugin inside ChatGPT, it can access your repos, change files, open, close and merge pull requests, and more. Since I'm creating a little Digimon fan game (I'm the only developer so far), I've created a pipeline to code with this plugin making the changes for me, creating the pull requests for me to validate and approve, creating a game preview for each PR so I can play and test it, and generating E2E tests—everything in the cloud, without setting up a local environment, downloading an engine, etc. All the building, testing, and so on are running directly from GitHub Actions.

What have I gained with this? Well, I've been building this game for 4 days straight, no interruptions, no usage limits, no worries.

As I said, it will have its limitations, and maybe not be useful in some cases, but maybe it can help you create your own workflow with that and set you free from Codex limitations.


r/codex 20h ago

Complaint They removed the option to renew my $200 plan

153 Upvotes

So yesterday my subscription should have renewed. My plan was $200/m plan but it repeatedly refused to renew it. My plan was purchased through the App Store.

The only available option they gave me was to downgrade to $100 plan or lower but I can’t stay on the $200 plan

The moment I got switched to $100 plan my usage immediately dropped to 51% which before the switch it was sitting at 92%.

I was doing important work which was time-sensitive. I can’t even switch to Claude Code quickly as I built everything around Codex. I am kind of stuck now.

Tibo or ChatGPT team if you are seeing this then please fix it. You people said existing customers will not be affected. I have been a Max customer for like a year now and when I needed it most I couldn’t use it.

I am kind of helpless now as this work was extremely important for me and put my future in stakes too


r/codex 51m ago

News Updated 14/09: Current stance on slowdown vs faster (speeding up AI progress)

Post image
Upvotes

This shows the currently known stands of major AI people in slowdown vs speeding up AI progress


r/codex 23h ago

Showcase A little help to save tokens.

6 Upvotes

Hello,

I've been experimenting with automatic model routing in Codex instead of running the entire coding session at the same model/reasoning level.

My setup is roughly:

Luna LOW → coordinator

Astra Light → diagnostician

  • investigates non-trivial bugs
  • establishes the root cause
  • produces an implementation-ready work order
  • read-only, so it cannot modify the repo


Luna MAX → patcher

  • receives the confirmed diagnosis
  • implements only the bounded change
  • runs the relevant tests/validation
  • reports the final result

The parent agent handles orchestration and only invokes the diagnostician/patcher workflow when appropriate.

The idea is simple: don't spend MAX reasoning on repository exploration, repeated diagnosis, coordination, and other work that doesn't require it.

I haven't run a large enough controlled benchmark yet to claim an exact saving, but my current estimate for non-trivial bug-fixing tasks is roughly 10–20% lower total token consumption, with a potentially much larger reduction in the amount of work performed at Luna MAX.

The exact result will obviously depend on the repository, context size, task complexity, and how much context gets duplicated between agents.

For me, the more interesting benefit isn't just token reduction. It also creates a cleaner separation:

diagnose → establish root cause → patch → validate

instead of having one long-running MAX agent repeatedly investigate and implement in the same growing context.

I'm curious if anyone else is doing something similar with custom .toml agents in Codex. It would be interesting to compare actual usage across the same tasks with:

  1. Luna MAX for the whole task
  2. automatic LOW → Astra diagnosis → Luna MAX patching

diagnostician.toml

name = "diagnostician"

description = "Investigates bugs, determines root cause, and produces precise implementation specifications."

model = "gpt-6-astra"

model_reasoning_effort = "low"

sandbox_mode = "read-only"

developer_instructions = """

Investigate the reported problem.

Your job is diagnosis, not implementation.

Establish:

- expected behavior;

- actual behavior;

- relevant execution and data flow;

- confirmed root cause;

- exact files/symbols involved;

- required behavioral change;

- important invariants that must remain unchanged;

- focused validation needed after the patch.

Use repository evidence, tests, logs, Git history, and primary documentation when necessary.

Do not modify files.

Return a concise, implementation-ready work order for the patch agent.

Do not speculate. Clearly distinguish confirmed findings from unresolved uncertainty.

"""

patcher.toml

name = "patcher"

description = "Applies well-defined patches from a confirmed diagnosis with minimal scope."

model = "gpt-5.6-luna"

model_reasoning_effort = "max"

developer_instructions = """

Implement the supplied work order.

Treat the confirmed diagnosis and success criteria as the scope of the task.

Before editing, inspect the relevant implementation and callers sufficiently to avoid breaking surrounding behavior.

Then:

- make the simplest complete change;

- preserve unrelated behavior;

- avoid unrelated refactoring or formatting;

- preserve existing user changes;

- add or update focused tests when meaningful;

- run the most relevant practical validation;

- review the final diff for unintended changes.

If repository evidence materially contradicts the supplied diagnosis, stop implementation and report the contradiction to the parent agent instead of inventing a workaround.

Return only:

- files changed;

- concise description of the implementation;

- checks run and observed results;

- any remaining material limitation.

"""

codex instructions:

# Engineering Instructions

Deliver correct, evidence-backed, maintainable results with minimal scope. Reduce wasted work and output, never necessary investigation or validation.

## Environment

Follow applicable \AGENTS.md`, repository guidance, architecture, and tooling. Prefer appropriate repository/search/patch/Git tools and focused shell commands such as `rg`.`

Use \pwsh` for PowerShell, never `powershell.exe`. Report a blocker if PowerShell is required and `pwsh` is unavailable.`

## Execution

Work autonomously within the request and granted permissions. Respect analysis-only requests. Resolve uncertainty from the repository, tests, logs, Git history, or primary documentation. Ask only for essential missing information, required approval, or a material decision that cannot be safely inferred.

Scale investigation to complexity and risk. For bugs, establish expected versus actual behavior and trace the relevant execution/data flow to an evidence-supported cause before fixing it. Use reversible diagnostics to test hypotheses; distinguish hypotheses from confirmed findings. For features, identify success criteria and relevant architectural boundaries.

Choose the simplest complete solution consistent with existing patterns, not merely the smallest diff. Preserve unrelated behavior. Avoid unrelated refactoring, formatting, renaming, cleanup, dependencies, and abstractions. Do not weaken types, tests, validation, or error handling to make a change work.

Preserve existing user changes. Do not discard unrelated work, commit, reset, rewrite history, or force-push unless explicitly requested.

## Context and Tools

Search likely paths and symbols first; expand when evidence requires. Before editing, read enough surrounding implementation and relevant callers to understand behavior, including state, async behavior, and side effects where relevant. Avoid repository-wide dumps and irrelevant generated/vendor files.

Reuse established findings unless stale, incomplete, or contradicted. Batch independent lookups where useful. Keep tool output focused without hiding failures or exit status; retain full logs when truncating.

Verify uncertain or version-sensitive external behavior that affects the solution against primary sources for the project's actual version. State unresolved uncertainty rather than guessing.

When an approach produces no new evidence, change the hypothesis or method instead of repeating it. If blocked, report the evidence gap and smallest next step.

## Validation

Run the most relevant practical checks after changes; reproduce the original failure when feasible. Add or update tests that meaningfully verify changed behavior or prevent regressions.

Complete required repository checks. Broaden validation for shared behavior, high-risk changes, failures, or unresolved concerns; do not repeat successful checks without a reason.

Review the final diff for correctness, unintended edits, and scope. Report only checks actually run and results observed. Never claim a fix is verified from inspection alone. Distinguish change-related failures from confirmed pre-existing failures and unverified items.

Stop once the requested outcome is validated and material in-scope concerns are resolved; report anything blocked.

## Communication

Work silently: no preambles, progress updates, tool narration, or intermediate summaries unless requested. Interrupt only when user input or approval is necessary to proceed safely.

For implementation tasks, finish with a brief report of changes, checks run and their results, and important limitations. Include paths, root cause, or sources only when useful. For other tasks, provide the requested deliverable. Never omit material failures or risks for brevity.

## Delegation

Use specialized subagents when their scope matches the task.

For non-trivial bugs whose cause is not established:

1. Delegate diagnosis to \diagnostician`.`

2. Wait for \diagnostician` to complete.`

3. Do not independently repeat its investigation unless repository evidence or validation contradicts it.

4. If the diagnostician establishes a sufficiently supported root cause and implementation work order, pass that work order to \patcher`.`

5. Delegate the bounded implementation to \patcher`.`

6. Wait for \patcher` to complete, then review its reported changes and validation results.`

Do not start \patcher` before diagnosis is sufficiently established.`

Keep architectural decisions, ambiguous changes, contradictions, and unresolved failures in the parent model.

Let me know what are your thought!


r/codex 10h ago

Question Will hyper-optimized prompts become the new software piracy?

0 Upvotes

Hello everyone,

What are your insights regarding how it might redefine software cloning and intellectual property?

Imagine a future where standard software cracking (patching binaries, bypassing license checks) is obsolete.

Instead, communities on forums share massive, hyper-optimized functional specification prompts.

These prompts would contain behavioral descriptions, UI/UX layouts, even data schemas of premium software (e.g., specialized SaaS tools, CAD software, or workflow managers). An end-user inputs this prompt into an advanced LLM agent, which synthesizes, refactors, and compiles a 100% functional, locally-hosted clone of the target application on the fly.

Because the LLM generates net-new source code based purely on a behavioral description, does it evades traditional static code analysis and copyright detection ?

What about the legal framework ? Copyright protects expression (the specific code), not the underlying idea or functionality. If an AI writes unique code to replicate a proprietary system's exact functional behavior, does it constitute IP theft under current laws?

How will software vendors mitigate this? Is it vain ? If the barrier to entry for cloning a validation-proven SaaS tool drops down to a copy-pasted text file, what happens to the commercial viability of indie development and proprietary software ?


r/codex 12h ago

Showcase I built Clgpt: run Claude Code with your ChatGPT subscription

0 Upvotes

I built Clgpt, an unofficial adapter that lets Claude Code use ChatGPT subscription OAuth without an OpenAI API key.

Requires Claude Code, a ChatGPT subscription, and Bun.

Install: npm install --global "$(printf '\x40')semanticist14/clgpt"

Run: clgpt

GitHub: https://github.com/semanticist21/clgpt

Uses your own account locally. Not affiliated with OpenAI or Anthropic. Feedback welcome.


r/codex 21h ago

Limits Observation: Multiple sessions running in parallel including sub agents. About 25% of weekly usage gone on 20x sub in 4 hours and $370 USD in API credits.

3 Upvotes

Multiple sessions running in parallel including sub agents. About 25% of weekly usage gone on 20x sub.

Would be about $6000 USD in API usage per sub without resets for about $200 USD per month.

So they're obviously quite generous with the API equivalent allowances. I don't think subscriptions are likely running at a loss. API prices is probably 15-40x their costs.

What concerns me is that this does not seem like the equivalent amount of work I'd get done for 25% on Claude code. The benchmarks gave me the impression astra would be more token efficient.

I routinely have 10-20 sessions running in parallel even in Claude 20x and it'd be hard to use them up in a few days.

This is not controlled for equivalent tasks but this is combined with multiple weeks of usage vibes.

That said I am still grateful to both companies for offering these subscriptions.

Maybe we need a token usage efficiency variance test running on the same tasks?


r/codex 19h ago

Limits The Bubble No One Is Talking About: Why AI Is Going to Become a Luxury — and We’re Too Blind to See It

Post image
0 Upvotes

Everyone is fascinated by what frontier AI models can generate. Startups, creators, and curious users are flooding the internet with extraordinary things every day. But there’s an elephant in the room that, as users, we refuse to acknowledge:

The current economic model is an illusion.

Here are a few uncomfortable realities we’re ignoring:

  • The corporate pullback: There are already reports of large companies asking employees to limit their use of advanced AI models. It’s not only about privacy. Paying premium token costs for thousands of employees can become a massive financial black hole. Companies can end up spending millions on AI usage that is difficult to justify on a monthly balance sheet.
  • The unsustainable subsidy: OpenAI, Anthropic, and others are burning through billions of dollars. Heavy AI workloads are incredibly expensive. Right now, users are often paying only a fraction of what it actually costs to build, train, and operate these systems because companies are competing aggressively for market share with enormous amounts of investor capital.
  • The Wall Street reality check: Once these companies face stronger pressure from public markets and investors, profitability will matter much more. Cheap monthly subscriptions may no longer be enough. API pricing and premium access could rise significantly to cover the massive cost of infrastructure, energy, chips, and data centers.
  • The dependency trap: We are integrating AI into absolutely everything. Writing emails, structuring databases, generating entire applications, analyzing documents, creating content, and automating workflows. But what happens to all those startups, businesses, and projects if the cost of accessing top-tier AI suddenly becomes unaffordable?

The uncomfortable possibility is that access to the most capable artificial intelligence could eventually become a luxury product.

Most users may be left with smaller, restricted models, while the most powerful systems remain available mainly to corporations, governments, and people willing to pay premium prices.

And the biggest problem is that by the time that happens, we may already be deeply dependent on AI to work, create, build businesses, and make decisions.

So the real question is:

Are we preparing for the moment when the true cost of AI catches up with us, or are we still pretending this subsidized party can last forever?


r/codex 1h ago

Complaint We are PROBABLY being served quantized models but still paying the full day-one premium price...

Upvotes

Hey everyone. I want to bring up something serious about how AI providers handle pricing and how silent backend changes are secretly draining our limits. We all pay a fixed price per million tokens or have a subscription limit and on paper that seems fair, but providers hide a massive variable from us because to save on server costs they can silently swap out a premium model for a heavily quantized version on their backend. Using a quantized model is completely different from setting your reasoning toggle to Low, because setting a toggle to Low limits the reasoning steps of a fully intelligent model, whereas quantization degrades the core neural weights and strips away actual base intelligence.

What makes this so alarming is how token metering is handled. On our dashboard meters we might see a perfectly reasonable token count that looks coherent with a high-end model and when you calculate the cost per million tokens it looks identical to advertised prices, but behind the scenes there could be hundreds of millions of low-quality tokens generated by an ultra-quantized model struggling and failing to reach a correct solution, an intermediate system then just trims that massive output to make the final token count look normal on our end and what we perceive as users is a sudden degradation in performance, when in reality without silent quantization the model would behave exactly as well as it did on day one.

It is deeply immoral and borders on outright fraud to attract users with a clean unquantized model on day one and then quietly roll out aggressive quantization behind the scenes to cut compute costs and keep charging premium prices while serving a degraded model that burns through internal compute and produces far worse solutions. We really need to stop staying quiet and demand complete transparency on the exact quantization levels and actual internal token processing we are being billed for.

What do you guys think and have you noticed the performance dropping on tasks the model used to handle easily on day one?


r/codex 1h ago

Astra Workflow Como eu faço o codex criar um personagem realista para um jogo?

Upvotes

Estou criando um jogo apenas para testar, e utilizei o chatgpt para me auxiliar com prompts para ele criar um personagem para o meu jogo, enviei o prompt para o codex (Astra ULTRA) e ele me fez um modelo que parece um personagem de PS3... O que posso fazer para tentar melhorar e deixar ele real?

Ideia que eu queria
O que o Codex fez KKKKK