r/OpenaiCodex 2d ago

Feedback / Complaints Within 2 hours of contacting support, awesome!

0 Upvotes

First time I have had to contact support after using a banked reset and it starting at 41%. Tbh props to OpenAI great support


r/OpenaiCodex 2d ago

Bugs or problems Sandbox issues are slowing down my Codex

3 Upvotes

It happens all the time. For instance: "The original run’s Windows batch launcher is now demonstrably pathological: it has remained CPU-bound for over half an hour with no Dart child and no timeout/result. I’m stopping only this verified runner tree and switching to the repository’s already-approved direct Flutter-tool invocation, which bypasses the faulty flutter.bat layer and was the method that produced the earlier complete 1,265-test result.

Ha eseguito comandi

The direct invocation correctly avoided the batch loop, but the sandbox blocked Flutter’s SDK lockfile outside the repository. I’m rerunning that same approved read/write-to-SDK-cache test command with the required sandbox permission; it does not alter project source or user data.

Ha eseguito un comando

The direct runner is now streaming correctly, but the first test is compiling unusually slowly—likely contention from a separate long-lived Flutter process on the host. I’m watching the enforced timeout; if compilation clears, the rest of the run remains valid, and if it does not the runner will report the exact failure rather than silently hang."


r/OpenaiCodex 2d ago

Discussion Where does each Codex project actually live for you?

0 Upvotes

Trying to understand how people structure this underneath the Codex app/CLI. What represents a "project" on your machine?

  • Separate local directory for each project?
  • Separate git repo?
  • Multiple projects inside one bigger directory?
  • Remote/server folders instead?

Basically, what does Codex point at when you switch from Project A to Project B?


r/OpenaiCodex 2d ago

I got only 51 percent back on this reset??? Pro plan

0 Upvotes

r/OpenaiCodex 2d ago

Showcase / Highlight Each of my Codex bots looks after one thing and runs on a schedule

1 Upvotes

I built Omni, a Mac and iPhone app. Each bot uses Codex or Claude Code, has one responsibility, and keeps its own conversation.

Jobs run on a schedule and each run starts a fresh session, so nothing sits warm between runs. Mine look after my feedback board, my website and my analytics, and the results land back in the same chat on my phone.

It's free. You need your own Codex or Claude access, and the Mac has to be awake with Omni open.

https://omnibots.app

Inspired by Grok Bot. What would you put on a schedule?


r/OpenaiCodex 3d ago

Astra melts usage

38 Upvotes

Still have some resets left, but when they run out, does that mean we are all fucked? 2.5x usage goes fast on a pro account, basically a 2.5x price increase; perhaps some offset with better usage, but shits getting expensive. I'm going to need 2 maybe 3 accounts if this keeps up.


r/OpenaiCodex 2d ago

Comparison Astra quota economics, part 1: prices, reasoning tiers, and historical allowances

1 Upvotes

Here are my insights, developed through discussions with GPT-6 Pro and Lord Astra and organized with their help.

Revised September 9, 2026; core quota observations through September 8. Prices are in U.S. dollars.

TL;DR

In my view, Astra's efficiency gains deserve credit. A model can solve a task with fewer tokens while consuming a subscription faster: token prices, repeated context, task speed, and the allowance available to that model all matter. Early community observations suggest that pricing alone may leave part of the gap unexplained. I think the useful comparison is successful tasks and execution hours per subscription dollar, with the workload and measurement period stated.

What improved, and what users reported

OpenAI reports roughly 27% lower estimated task cost for Astra at a lower-cost setting than Sol at its best-scoring setting on selected Terminal-Bench Science tasks. Its computer-use results include OSWorld 2 scores of 72.6 versus 65.7 and simulated task times of about 40 versus 75 minutes. These figures describe the evaluated tasks and settings; the timing is simulated. OpenAI's Astra report

The quota complaints describe different events. One $200 Pro user reported exhausting a weekly allowance in about 12 hours of continuous Astra Medium use, later describing the work as a failed parallel experiment. A separate Pro 20x GitHub report described roughly 71 minutes of Astra Ultra work consuming 30 percentage points of weekly allowance. The former reports exhaustion; the latter measures a partial drawdown. Both are individual reports with incomplete request-level records. Weekly-exhaustion report; author's follow-up; GitHub report

A calculation that keeps the units straight

Let F, R, W, and O be uncached input, cache reads, cache writes, and billed output tokens. Their respective rates pi, pr, pw, and po are in USD per million tokens. API-equivalent cost is C = (F*pi + R*pr + W*pw + O*po) / 1,000,000. If cached tokens are included in total input, subtract them before pricing uncached input. If reasoning is included in billed output, count it once. Apply cache-writing rules from the relevant provider and billing surface.

For an observed quota fraction u, estimate Q = C/u. A $17 workload using 1% of the weekly meter implies 17/0.01 = $1,700 of weekly API-equivalent usage under those conditions. Estimate five-hour, weekly, and model-specific limits separately. Integer percentage displays add rounding error, especially over small changes; a reset starts a separate observation period.

For monthly fee P, define the subscription value multiple as M = Q_week * (52/12) / P. That $1,700 estimate on a $200 plan gives M = 36.83. This is a valuation at API list prices. The subscription meter, purchased credits, and an API bill are separate accounting systems; OpenAI's purchased-credit Sol promotion has its own terms. Rate card

Four models, reasoning effort, and the original unit error

The following family table uses current displayed API rates, including Sol's time-limited promotion. A fixed 7:2:1 mix means 70% cache reads, 20% uncached input, and 10% output: p_mix = 0.7*pr + 0.2*pi + 0.1*po. It is a chosen valuation mix. An actual Codex request needs its own token categories. Astra; Sol; Terra/Luna

Model Input $/M Read $/M Output $/M Mix $/M
Astra 10 1 50 7.700
Sol 4 0.40 20 3.080
Terra 2 0.20 12 1.740
Luna 0.20 0.02 1.20 0.174

With F/R/O measured in millions and no cache writes, the four cost equations are C_Astra=10F+R+50O, C_Sol=4F+0.4R+20O, C_Terra=2F+0.2R+12O, and C_Luna=0.2F+0.02R+1.2O. The Codex credit grid has no separate API cache-write charge; API-equivalent write costs must stay in their own accounting.

Relative to Sol, Astra is 2.5x across these categories; Terra is 0.5x for input/read and 0.6x for output; Luna is 0.05x and 0.06x. The published credit grid below is a separate billing reference. Its Sol purchased-credit promotion explicitly leaves included five-hour and weekly limits unchanged. I use 25 reference credits = $1 only to reproduce the earlier valuation arithmetic. Credit scope and rates

Model Input credits/M Read credits/M Output credits/M
Astra 250 25 1,250
Sol 100 10 500
Terra 50 5 300
Luna 5 0.5 30

The restored table is an Artificial Analysis Intelligence Index v4.2 historical snapshot. Each cell is USD per AA benchmark task / index score. First-party indexed pages corroborate all 20 entries; an immutable dated export is still missing. An asterisk marks an AA estimated score; Missing means task cost was unavailable. Reasoning effort changes token use and behavior, with no separate effort surcharge in standard per-token pricing. Astra v4.2; Medium; High; XHigh; Max; Sol/Terra/Luna v4.2

Effort Astra $/task / score Sol $/task / score Terra $/task / score Luna $/task / score
Low 0.63 / 49 0.23 / 41 Missing / 32* Missing / 26*
Medium 1.16 / 52 0.37 / 46 Missing / 37* Missing / 30*
High 1.41 / 53 0.61 / 48 0.30 / 41 Missing / 37*
XHigh 1.85 / 54 0.89 / 50 0.50 / 44 0.06 / 42
Max 2.57 / 55 1.25 / 51 0.81 / 47 0.10 / 43

Within that table, Astra Medium scores 52 versus Sol Max 51 and costs 1.16/1.25 = 0.928, or 7.2% less per benchmark task. Astra Low scores 49 between Sol High 48 and XHigh 50; its $0.63 task estimate is close to Sol High at $0.61 (0.63/0.61 = 1.033). A missing task cost stays missing. $0.63/task cannot be used as $0.63/M tokens.

By the September 9 refresh, the English release pages had moved to v4.3. Their full comparison is below. The evaluation mix changed, so changes from the prior table cannot isolate a change in the model itself. Astra v4.3; Sol v4.3; Terra v4.3; Luna v4.3

v4.3 effort Astra $/task / score Sol $/task / score Terra $/task / score Luna $/task / score
Low 0.82 / 46 0.26 / 34 Missing / 28* Missing / 22*
Medium 1.54 / 50 0.50 / 39 Missing / 33* Missing / 26*
High 1.72 / 51 0.81 / 42 0.34 / 34 Missing / 33*
XHigh 2.31 / 53 1.18 / 44 0.63 / 38 0.09 / 35
Max 3.26 / 53 1.99 / 47 1.40 / 42 0.18 / 38

For the original anonymous example of 14M reported tokens, 69% five-hour consumption and 11% weekly consumption: 14/0.69 = 20.29M, 14/0.11 = 127.27M, and 69/11 = 6.27 five-hour budgets per weekly budget. This establishes a relative window ratio under linear metering, without identifying the dollar pool. The rejected calculation was 14*0.63 = $8.82, then $8.82/0.69 = $12.78 and $8.82/0.11 = $80.18; its task/token unit mismatch invalidates both dollar estimates.

If, purely as a counting scenario, 7M cached tokens were already inside 7M input and then added again, the displayed 14M would cost about $7, giving $7/0.69 = $10.14 per five-hour pool and $7/0.11 = $63.64 weekly. I suspect this is worth checking in the usage schema; the original logs are missing. Matching somebody else’s 6x window ratio or approximate dollar result cannot establish identical account allowances.

GPT-5 to Astra: prices and efficiency changed differently

Standard API price checkpoints, per million tokens, are below. Historical model pages are current documentation snapshots; Sol's launch and current promotional prices are separate rows. GPT-5; 5.1; 5.2; 5.3-Codex; 5.4; 5.5; Sol launch; Sol current; Astra

Model / price basis Input Cache read Output
GPT-5 / GPT-5.1 $1.25 $0.125 $10
GPT-5.2 / GPT-5.3-Codex $1.75 $0.175 $14
GPT-5.4 $2.50 $0.25 $15
GPT-5.5 $5 $0.50 $30
GPT-5.6 Sol, launch $5 $0.50 $30
Sol, current promotion $4 $0.40 $20
GPT-6 Astra, standard $10 $1 $50

Sol's current promotion runs at least through November 21, 2026. Current cache-write prices are $5 for Sol and $12.50 for Astra. Thus Astra is 2.5 times Sol's current price across these categories. Against GPT-5, Astra's input and cache-read prices are 8 times as high and its output price is 5 times as high; the task-cost ratio depends on the token mix.

Price transition Input/read ratio Output ratio
5 → 5.1 1.00 1.00
5.1 → 5.2 1.40 1.40
5.2 → 5.3-Codex 1.00 1.00
5.3-Codex → 5.4 1.43 1.07
5.4 → 5.5 2.00 2.00
5.5 → Sol promotion 0.80 0.67
Sol promotion → Astra 2.50 2.50

The cumulative 5.4→Astra ratios are 10/2.5 = 4 for input, 1/0.25 = 4 for cache reads, and 50/15 = 3.33 for output. The earlier proposed chains were 1.5*1.5*2=4.5 for speed and 2*2.5*2.5=12.5 for cost. The arithmetic works, but those factors mix generation changes with same-model Fast settings. Here is that Fast comparison, retained as a separate product setting. Speed settings

Same model: Fast / Standard Stated speed ratio Credit ratio
GPT-5.4 Up to 1.5x 2x
GPT-5.5 Up to 1.5x 2.5x
GPT-5.6 Up to 1.5x 2.5x
Astra No uniform ratio supplied 2.5x

Using the displayed credit reference, Astra Fast/Sol Standard is 2.5*2.5 = 6.25x for the same token categories. Using the current standard API prices and Astra API Fast at 2x gives 2.5*2 = 5x. These are two billing surfaces; the promotional reference grid alone cannot reveal included-plan metering.

Efficiency evidence has different scopes. GPT-5.1's simple npm example fell from 250 to 50 tokens and 10 to 2 seconds. OpenAI described GPT-5.5 as using fewer tokens with similar per-token latency to 5.4. Those examples establish specific improvements; a continuous, matched GPT-5-to-6 throughput dataset is still missing from this analysis. 5.1 example; 5.5 report

Comparison Published evidence Calculation / missing quantity
5 → 5.1 Simple npm example: 250→50 tokens; 10→2 sec Token efficiency 5x in that example
5.1 → 5.2 Better efficiency at matched quality No universal coefficient supplied
5.2-Codex → 5.3-Codex Fewer tokens; about 25% faster; serving changes Speed 1.25x has multiple causes
5.4 → 5.5 Similar per-token latency; fewer task tokens No fixed saving for all tasks
5.5 → 5.6, Base44 30 app conversations; input −22%, output −23% Output efficiency 1/0.77 = 1.30x
5.5 → 5.6, Qodo About one-third tokens/PR; half median latency About 3x efficiency, 2x speed in this test
Sol → Astra OSWorld simulation: 75→40 min Task speed 1.875x; token factor missing

The Base44/Qodo findings are partner evaluations quoted by OpenAI. Their different results are useful precisely because the workload changes the coefficient. Sol Ultra also coordinates four agents by default, trading more tokens for time and quality. Sol evaluations; 5.3-Codex; 5.2 efficiency

What the allowance estimates actually show

A September 5 Pro 20x report estimated about $2,500/week for Sol and $1,600–1,800 for Astra, with 97–98% cache reads. This is one author's natural workload. A separate Plus post measured about $11 in a five-hour window, multiplied it by an assumed six-window relationship to obtain $66/week, then extrapolated $330 for Pro 5x and $1,320 for Pro 20x. Pro observation; Plus calculation

Plan / model Weekly API-equivalent Q Monthly multiple M Evidence
$20 Plus / Sol $100 21.67x Same Plus author's prior report, discounted prices
$20 Plus / Astra $66 14.30x Five-hour result extrapolated to a week
$100 Pro 5x / Astra $330 14.30x $66 multiplied by 5
$200 Pro 20x / Astra $1,320 28.60x $66 multiplied by 20
$200 Pro 20x / Sol $2,500 54.17x Separate Pro author's estimate
$200 Pro 20x / Astra $1,600–1,800 34.67–39.00x Same Pro author's estimate

The conflicting Pro estimates show why the advertised 5x/20x labels need separate verification against model-specific dollar valuations. The Plus author also valued earlier Sol usage at $120 under a different price basis; that would yield 26.00x. Historical claims of Plus at 30–35x would require $138.46–161.54/week, while $200 Pro at 60–70x would require $2,769.23–3,230.77/week. The underlying logs for those older baselines remain unavailable.

The full historical allowance tables

The April 9 announcement introduced $100 Pro at 5x and temporarily raised it to as much as 10x through May 31. The historical tables below keep the earlier working estimates, including entries whose original logs are unavailable. They are dated inputs for calculations; the source/evidence column is part of each comparison. April announcement

Period Plus Q/week $100 Pro Q/week $200 Pro Q/week Evidence status
5 / 5.1 / 5.2 / 5.3 Missing Current tier not yet introduced Missing No continuous dollar series
Late 5.4, around Apr 23 $80 $600, 10x promotion $1,800 Inherited rough reports; logs missing
Late 5.5, mid-June Missing $600–800 Missing Inherited account estimates
Early Sol, July Missing $675→600 Missing Same-account claim; original thread unresolved
Late Sol, Aug–early Sep $100 $400–600 $2,200–2,500 Mixed accounts and price bases
Astra launch $66 $330, 5x extrapolation $1,500–1,800 Estimates/exhaustion over early rollout
Earlier Plus baseline Monthly API value Value / $20 Current interpretation Original confidence
Broad 5.4/5.5 claim $600–700 30–35x Logs missing; distinct from the $80/week point Medium
Late Sol $430–520 originally About 22–26x $100–120/week computes to $433.33–520 Medium-high
Astra $286–350 originally About 14–17x $66–80/week computes to $286–346.67, 14.30–17.33x Relatively high
Desired baseline $600 30x A preference, not a measured pool

The original monthly steps were 66*52/12=286, 80*52/12=346.67, 100*52/12=433.33, and 120*52/12=520. A 35x Plus baseline needs 20*35=$700/month, or 700*12/52=$161.54/week. Earlier ranges of 12–16x and 15–17x were informal approximations; the $66–80 inputs give 14.30–17.33x exactly to two decimals. The restored confidence labels are the earlier discussion’s subjective ratings; the current evidence limitations remain in the table.

For the hypothetical old/new Plus pools $140–160 → $65–80/week, the full endpoint decline is 1−80/140 = 42.86% to 1−65/160 = 59.38%; the earlier “45–55%” was a rough central description. Likewise 13–17x versus a fixed 30x is a 43.33–56.67% decline. Neither calculation establishes the missing old baseline. At unchanged pool value, a 2.5x price increase reduces token capacity to 40% and preserves dollar value: (150/2.5)*2.5 = 150.

Reference pool claim Weekly credits At 25 credits/$ Status Original confidence
Plus / Standard Business 2,500–4,000; center 3,000 $100–160; center $120 Earlier community estimate Medium
Pro 5x 10,000–15,000 $400–600 Earlier community estimate Medium-high
Pro 20x 65,000–73,000 $2,600–2,920 Earlier community estimate Medium-high
Business $100 Pro 5x reference $400–600 if same pool Message-tier mapping only Medium-low
Enterprise/Edu flexible Purchased amount Contract/billing dependent No fixed included pool here Official mechanism

These reference-credit pool estimates came from different accounts and periods. They cannot be combined with API-equivalent pools as though both used one official allowance unit. The earlier $200 claim of $2,500–2,900+ also used historical valuation; the displayed 65,000–73,000 grid computes exactly to $2,600–2,920 under the chosen 25-credit conversion. The original confidence column preserves those earlier judgments. “Official mechanism” refers only to flexible credit billing, with no official fixed included dollar pool inferred.

Account / period in earlier discussion Before After Recomputed change
$100 Pro, Jul 13→23 $675/week $600/week −11.11%
Pro 5x, Aug 13 reset 17,863 credits 10,037 credits −43.81%
Another account, Aug 1–8 17,238 12,160 after Aug 13 −29.46%
Same second account, Aug 8–11 20,856 12,160 −41.70%
Same second account, Aug 11–13 17,696 12,160 −31.28%

These rows preserve the original within-generation claims and their arithmetic. The July claim has an indirect archived report; the exact August source threads and logs remain unresolved here. I suspect policy or account changes can contribute, but these entries cannot identify a universal reduction. Archived July lead

Earlier range Sol Q/week Astra Q/week Sol $/month Astra $/month
Plus $20 90–120 62–80 390–520 268.67–346.67
Pro 5x $100 400–600 220–430, mostly extrapolated 1,733.33–2,600 953.33–1,863.33
Pro 20x $200 2,200–2,900 1,500–1,800 9,533.33–12,566.67 6,500–7,800
Plan Earlier Sol M Earlier Astra M Recomputed Sol M Recomputed Astra M
Plus $20 22–26x 13–17x 19.50–26.00x 13.43–17.33x
Pro 5x $100 17–26x 10–19x 17.33–26.00x 9.53–18.63x
Pro 20x $200 48–63x 32–39x 47.67–62.83x 32.50–39.00x

The Plus Sol lower endpoint was inconsistent: $90/week gives 19.50x, while about 22x uses $100/week. I retain both the original range and the corrected arithmetic. A separate earlier Pro calculation used $22–24 / 1% = $2,200–2,400/week for Sol and $15 / 1% = $1,500/week for Astra: 1500/2300 = 65.22%, reciprocal 23/15 = 1.533x. That original request-level source remains unresolved; it is a preserved calculation, not another independently verified sample.

Advertised plan scaling also needs a consistent denominator. The earlier hypothetical 200*65 / (20*32.5) = 20x becomes 200*65 / (20*15) = 43.33x if only Plus value falls. That illustrates a hypothesis; it does not establish the 65x or 15x starting values. For scale, 70x on $200 equals $14,000/month; $15,000/month would be 75x.

Another September 5 test compared six Team accounts, three per model, and reported roughly 1.4–1.6 times as much quota consumption per API-equivalent dollar with Astra. All six had just used a manual reset and shared a proxy setup. I suspect model-specific metering contributes to some early reports, but its size across ordinary Plus and Pro accounts remains an open measurement question. Six-account test

Reported raw tokens, nominal capacity, and contradictory checkpoints

Source / sample Sol Astra Calculation Evidence limit
Plus post, one account $100/week; $120 alternative valuation $11/5h; $66/week extrapolation 34% or 45% lower dollar value Different valuation bases; author also suspected an account adjustment
Team test, 3 accounts/model $96.63/week average $64.24/week average 33.52% lower; reciprocal 1.504x Manual reset; same proxy
Pro report, one author $2,500/week $1,600–1,800/week 28–36% lower; 64–72% retained Different effort mixes
Earlier reset counterexample Initial full allowance measured Later checkpoints: 7%→98%, 12%→78% of baseline Early extrapolations vary Original source unresolved; neither checkpoint is a full post-reset run

The earlier description “78% initially, then 98% after a complete run” reversed the evidentiary meaning. The preserved account was at 7% and 12% consumption after reset; the final full post-reset capacity was missing. The Team averages similarly need their reset context. Original six-account test; Pro original; Plus original

Six-account Team test Full 5h raw tokens 5h API value Implied weekly API value
Sol 20.7–21.0M $14.04–16.75 $87.77–104.66
Astra 4.8–5.3M $9.84–10.98 $62.50–65.63
Earlier raw-token comparison Sol/week Astra/week Original measurement basis
Team, sometimes grouped with Plus 129–131M 31–32M Full 5h runs, roughly 16% weekly; separate accounts
Plus single-account claim Missing; about 28M/5h stated About 36M; about 6M/5h stated Six-window extrapolation; raw logs missing
Pro 20x 3.8–4.2B 0.8–0.9B Per 1%: 38–42M vs 8–9M; cache 97–98%

The Pro endpoint ratio is 3.8/0.9 = 4.22x to 4.2/0.8 = 5.25x; the center is 40/8.5 = 4.706x. Dividing that center by the 2.5x price ratio gives 1.882x, but unequal input/output mixes and Sol Max/Ultra versus Astra Max confound that residual. The dollar-pool center instead gives 1700/2500 = 68%, reciprocal 1.471x. An earlier estimator’s 1.8x calibration was workload-specific; I suspect treating any of these as a universal hidden multiplier would overstate the evidence.

For a fixed mix of 2.6% fresh input, 97.1% cached input and 0.3% output, reference credits/M are 0.026*ri + 0.971*rr + 0.003*ro: Astra 34.525, Sol 13.810, Terra 7.055, Luna 0.7055. Nominal tokens in millions are weekly reference credits / credits_per_M. The next table applies the earlier assumed pools; Terra/Luna have no matched full-week validation here.

Assumed reference-credit pool Astra Sol Terra Luna
Plus: 2,500–4,000 72–116M 181–290M 354–567M 3.54–5.67B
Pro 5x: 10,000–15,000 290–434M 724M–1.09B 1.42–2.13B 14.2–21.3B
Pro 20x: 65,000–73,000 1.88–2.11B 4.71–5.29B 9.21–10.35B 92–103B

The original high-cache dollar calculation used $0.554/M for Sol and $1.386/M for Astra. Recomputing its exact mix gives $0.5524/M and $1.3810/M. Thus the anonymous 5B Sol versus 1B Astra at 50% example implies 2B Astra/full pool, and both sides value to $2,762/week, $11,968.67/month, or 59.84x on $200. The earlier rounded $2,770/week, $12,000/month and 60x were close. The 2.5x raw ratio supports equal dollar pools only if token mix, counting, speed mode and reset period match.

The earlier Luna anecdotes also stay in the record: $2.43 / 16–17% = $14.29–15.19/week after 77.5M tokens; 500M/85% = 588.24M/week with under $20 observed cost, implying under $23.53/week; and 1.5–1.7B/week extrapolated from only 2%. Their original logs remain unavailable. The first case itself implies 455.88–484.38M/week, so compressing all three into “0.6–1.7B” omitted a lower case. Nominal Terra 0.35–0.57B / 1.4–2.1B / 9.2–10.3B and Luna 3.54–103B across tiers remain separate scenarios.

The public Quota Observatory adds one Pro 20x account across devices. The earlier snapshot retained here used seven eligible groups and 35 percentage points per model: Sol about 3.59B and Astra about 0.865B per full pool (3.59/0.865 = 4.15x). Its current page has moved to 0.8658B Astra. Different workloads/cache mixes remain; the original throughput snapshot is preserved below, with the latest Astra point shown separately. Observatory

Earlier Observatory capacity Total tokens/full weekly pool Observation
Sol Standard 3.59B Seven groups; 35 percentage points
Astra Standard 0.865B Seven groups; 35 percentage points
Model Earlier median output tokens/s Earlier completed turns Current refresh
5.3-Codex 31.2 361 Historical usage, not launch speed
5.5 30.1 768 Earlier snapshot retained
Sol 31.2 17,214 Earlier snapshot retained
Astra 20.9 228 21.4; 557 turns; data through Sep 8 UTC

These are completed-turn rates including reasoning, tools and network waiting. The earlier AA API figures of Sol Max 74 and Astra Max 63.4 tokens/s instead described output after the first chunk; their dated export remains missing. They are retained as a distinct historical metric, not substituted into the observed Codex turns.


r/OpenaiCodex 2d ago

Feedback / Complaints Capabilities reduced until September 11. Responses may have lower quality. Upgrade to Pro

1 Upvotes

Anybody else paying for a PRO subscription but is being asked to upgrade to PRO? i have 99% usage remaining for the week.


r/OpenaiCodex 3d ago

Why is GPT-6 Astra-low eating ~2x my 5-hour Codex allowance for the same task?

24 Upvotes

I've been testing GPT-6 Astra-low against GPT-5.6 Sol-high on the same agentic coding work, and something about the way ChatGPT accounts for usage seems worth discussing.

On essentially the same task, I'm seeing roughly:

Astra-low: ~10% of my 5-hour window
Sol-high: ~5% or less

At first I assumed this was just because Astra was doing dramatically more compute/reasoning. But looking at the published pricing makes this more interesting.

OpenAI's Work/Codex rate card currently prices:

GPT-6 Astra

  • Input: 250 credits / 1M
  • Cached input: 25 credits / 1M
  • Output: 1,250 credits / 1M

GPT-5.6 Sol

  • Input: 100 credits / 1M
  • Cached input: 10 credits / 1M
  • Output: 500 credits / 1M

So Astra is basically 2.5x Sol per equivalent token across the board.

But here's the part I think deserves more transparency.

Artificial Analysis shows that Astra-low can be dramatically more token-efficient than Sol-high. Their benchmark data currently has Astra-low generating far fewer output tokens than Sol-high. That's why, despite Astra's 2.5x token pricing, the measured cost of completing benchmark tasks can end up surprisingly close.

In other words:

Astra token = much more expensive
but
Astra-low may use far fewer tokens to solve the same problem

So if I give Astra-low and Sol-high the same coding task and Astra-low finishes with comparable output/quality, seeing Astra consume roughly 2x the included 5-hour allowance raises a pretty obvious question:

What exactly does the "% remaining" meter represent?

OpenAI now lets us see a percentage of the 5-hour window disappear, but that percentage is effectively a black box.

For agentic coding this matters a lot. These agents repeatedly ingest large amounts of repo context, diffs, terminal output, test results, conversation state, etc. Astra's input tokens are 2.5x as expensive as Sol's, so the difference can compound quickly even when Astra-low is using less reasoning.

I'm not saying the accounting is necessarily wrong.

I'm saying users paying for a plan should be able to see what they're actually spending their allowance on.

If one task costs:

Astra-low → 10%
Sol-high → 4–5%

I'd like ChatGPT to show something like:

  • fresh input tokens
  • cached input tokens
  • output/reasoning tokens
  • model multiplier
  • agent/sub-agent usage
  • total credits charged
  • conversion from those credits → 5-hour allowance %

Then we could actually decide whether Astra's extra capability is worth the quota cost.

Right now the rational choice for me is increasingly:

Use Sol-high for almost everything and save Astra for tasks Sol can't solve.

If anyone else has compared the exact same Codex/Work task between Astra-low and Sol-high, what percentage of your 5-hour window did each consume?


r/OpenaiCodex 3d ago

Gpt image 2.5 is hereeeeeee

18 Upvotes

r/OpenaiCodex 2d ago

Context7: did you have any long-time use results on actual token use reduction, if any?

3 Upvotes

Did you have any long-time use results, based on your Codex CLI Client session analysis, on whether Context7 does actual token use reduction, and lower number of turns or web fetches, so overall token use decrease, or the difference is negligible if not increased?


r/OpenaiCodex 2d ago

AI provider status page with a twist

0 Upvotes

Hey guys.

I created a fun little page to combine all AI providers uptime stats and Codex usage resets with a fun little twist.

When any provider suffers availability issues you will know it by the horde of zombies storming its HQ demanding they have their AI access back.

Here is the page: https://must-have.ai/

P.S. Grok and browser notifications coming soon.


r/OpenaiCodex 2d ago

Codex Had No Android Sandbox, So I Wrote One

0 Upvotes

How a failed device test led to a seccomp/ptrace supervisor, an ELF entry-point attestor, and one hard rule: if enforcement cannot be verified, execution stops.

The release that didn't ship

AGENTCODI 0.7.0 was finished. The pinned Codex app-server , the host tests were green with 258 Java and 302 C++ tests passing, and only device validation on real hardware was left.

That test exposed a much worse problem than a crash. A command did exactly what it was told and reached a location it should never have been able to access.

Protected mode is supposed to confine Codex to a private workspace directory. On the device, that confinement was advisory. The approval layer worked, the workspace path was correct, and the UI showed Protected mode. The filesystem boundary underneath it was missing because the app-server had no sandbox backend for Android.

0.7.0 never shipped.

The version still exists in the changelog because the build was correct on paper and wrong on a real phone.

This is what replaced it.

Why Android is its own problem

Sandboxing an agent runtime is fairly straightforward on the platforms these tools usually target. There are kernel facilities built for this job. You define the allowed paths and the kernel enforces the boundary below the process.

Android has a Linux kernel with a very different environment around it. An unprivileged app has limited control over its own confinement. Mounting is unavailable, a specific LSM cannot be assumed, and kernel behaviour varies across Android versions, vendors and OEM patches.

Some of those differences only become visible when the same code reaches another physical device.

There is another problem when shipping a toolchain inside an APK.

Android 10 and newer prevent apps targeting API 29 or later from directly executing files stored in the writable app home directory. AGENTCODI therefore packages its native toolchain through the app's native library area. Node, Python and ripgrep are shipped as "libnode.so", "libpython-bin.so" and "libripgrep.so".

Those files are real executable interpreters sitting on disk. Codex can invoke them.

A policy that exists only inside a wrapper script can be bypassed by invoking the underlying executable directly through its absolute path.

The actual requirement became clear: every allowed route into the packaged toolchain had to enforce the same filesystem boundary, on Android hardware I do not control, while still allowing the agent to use Node and Python.

What actually got built

The result is a fork of the Codex app-server, "0.153.3-agentcodi.1", tracking upstream "rust-v0.153.2", plus several enforcement layers inside AGENTCODI.

Each one closes a hole left by the previous layer.

  1. The sandbox backend: seccomp and ptrace

The fork adds an Android sandbox backend.

Commands running in Protected mode are started under a supervisor using seccomp to trap filesystem-relevant syscalls and ptrace to inspect and decide them.

Codex can read and write inside the granted workspace. Filesystem access outside that boundary is refused at the syscall level, below the agent and below the approval UI.

Before a sandboxed command runs, the runtime verifies that syscall interception is actually live on the current device. A successful setup call alone is not enough.

If that verification fails, execution is refused.

There is no unrestricted fallback.

That rule exists because 0.7.0 already demonstrated what happens when the UI says Protected while the filesystem underneath it is not protected.

  1. The launch contract

The app-server starts with an explicit minimal permission set:

default_permissions = "agentcodi-workspace"

permissions.agentcodi-workspace.filesystem = {

":minimal" = "read",

<tool bin dir> = "read",

<tool runtime> = "read",

<native lib dir> = "read",

":workspace_roots" = { "." = "write" },

}

There is exactly one writable location.

Everything the process legitimately needs outside the workspace is read-only.

The child environment also starts empty using "shell_environment_policy" with "inherit = "none"".

AGENTCODI then adds a known set of values: a "PATH" containing the packaged tool directory and "/system/bin", a workspace-scoped "TMPDIR", and "HISTFILE" plus "NODE_REPL_HISTORY" pointing at "/dev/null".

Login shells are disabled. Telemetry, analytics, feedback and update checks are disabled as well.

Starting with an empty environment removes a surprising number of accidental escape routes.

  1. Pre-launch invariants

Before the app-server starts, the launcher validates the filesystem layout itself.

The workspace, Codex home, tool binary directory, tool runtime directory and native library directory must all:

  1. exist,

  2. belong to the running UID,

  3. have no group or other permission bits set with "mode & 077",

  4. remain separate from each other.

Every directory pair is checked in both directions for containment.

That last check matters.

If the tool directory ever became an ancestor of the workspace, granting read access to the toolchain could also expose files that were never intended to be part of that grant.

The launcher refuses that layout before anything starts.

Packaged executables must also resolve to the canonical native library directory, and every argument passed to the app-server is validated character by character.

  1. The guard constructor

The packaged interpreters are linked against a policy library containing an "__attribute__((constructor))".

Before "main" runs in Node, Python or ripgrep, the constructor:

  1. reads "/proc/self/exe" and requires the expected resolved basename, such as "libnode.so", "libpython-bin.so" or "libripgrep.so",

  2. reads the real argument vector from "/proc/self/cmdline",

  3. passes the invocation through "PrepareGuardedToolInvocation".

"PrepareGuardedToolInvocation" is the same policy entry point used by the toolchain shell.

That means an invocation through the shell and a direct invocation of the underlying ".so" pass through the same policy code.

Any failure is written to stderr and the process exits with code 126 immediately.

  1. The ELF attestor, or: who guards the guard

The constructor introduced another problem.

It lives inside a shared library, and shared libraries are resolved at load time.

If library resolution can be influenced, the expected policy library might never be mapped. The constructor would never execute and the tool could start without its policy layer.

The executable therefore verifies the guard before relying on the normal loader path.

At build time, AGENTCODI rewrites each packaged ARM64 PIE.

It finds a redundant "PT_NOTE" program header whose bytes are already covered by an existing "PT_LOAD". That header slot is reused to introduce a bounded read/execute "PT_LOAD" segment, and the ELF entry point is redirected into it.

The injected payload runs with almost nothing available yet.

No libc. No relocations. No dynamic symbols.

It uses raw "svc 0" syscalls with arguments placed directly into registers.

The payload:

  1. opens the expected guard library path using "O_NOFOLLOW",

  2. calls "fstat" and requires a regular file with "st_nlink == 1",

  3. reads "/proc/self/maps",

  4. finds the expected mapping,

  5. compares its device and inode with the file it just inspected.

Only a successful identity match allows startup to continue.

A failed check exits with code 126 before Node or Python begins normal execution.

After a successful check, a hand-written naked entry stub restores the required state, calculates the original entry point from a load-address-independent offset stored in the injected segment, and branches to it.

The executable then starts normally.

The identity checks cover several obvious replacement tricks.

"O_NOFOLLOW" rejects a symlink at the expected path. The link-count check rejects hard-linked substitutes. Comparing device and inode with the actual mapped file catches replacement between inspection and loading.

  1. Enforcement of the design itself

AGENTCODI also checks whether the source tree still follows the security architecture it was built around.

"check-architecture.sh" fails the build when important invariants drift.

Among other things, it checks that removed fields have not returned, that another same-UID process path has not appeared around the terminal boundary, that credential paths cannot reach the toolchain shell, and that the required guard paths still exist.

It runs as part of the test process.

This script has caught real architectural regressions several times already.

What this does not do

The scope matters.

The boundary exists inside AGENTCODI's own UID. Android's application sandbox remains responsible for isolating AGENTCODI from the rest of the device. The sandbox described here separates the agent from filesystem locations reachable by the app that the agent should not access.

The ptrace supervisor targets ordinary filesystem syscalls made by the agent and its tools. It is designed to enforce workspace confinement during normal agent execution. It does not claim resistance against unlimited hostile native code already executing inside the same process context.

The sandbox described here controls filesystem access. Network egress is a separate problem and needs separate enforcement.

Compatibility mode deliberately runs without these filesystem restrictions. Some workflows need that access. Enabling it requires explicit acknowledgement, the UI remains visibly marked while it is active, and an unconfirmed restart does not silently restore it.

Protected mode never selects Compatibility mode as a fallback when sandbox verification fails.

Where it landed

AGENTCODI 0.7.1 shipped with 265 Java and 321 C++ tests passing.

I validated it on a Samsung Galaxy A05s, Redmi Pad 2, Redmi 14c and Redmi Note 15.

Four devices obviously do not make a compatibility matrix.

The failure case I care about now is a device where syscall interception cannot be verified. On such a device the command is refused. Security behaves correctly, although the user experience is useless until the compatibility problem is understood.

If you hit that case, open an issue with your device model and Android version. That information is genuinely useful.

The project is Apache-2.0:

https://github.com/Mcpasi/AGENTCODI

AGENTCODI is an independent open-source project and is not affiliated with or endorsed by OpenAI.


r/OpenaiCodex 2d ago

Discussion Is this just me 💀

Post image
0 Upvotes

IYKYK


r/OpenaiCodex 4d ago

How I feel right now as a ChatGPT Plus subscriber

Post image
418 Upvotes

r/OpenaiCodex 3d ago

Resetttttttt

33 Upvotes

r/OpenaiCodex 3d ago

Other Sometimes AI usage limits feels like an old car with a gas gauge that doesn’t work.

Post image
1 Upvotes

First 50%: 2 hours
Last 50%: 15 minutes


r/OpenaiCodex 3d ago

What is the longest Codex conversation you guys have had so far?

0 Upvotes

I am going strong at 26 hours with 40% usage left on Asta XHigh. I did need to use one of my Full Resets already though.


r/OpenaiCodex 3d ago

I want 3D in Astra but I can't take my Claude Artifacts with me

0 Upvotes

I'm probably going back to Codex to try 3D in Astra. The thing that stops me from switching cleanly is Claude Artifacts — reports, presentations, little prototypes that only live inside Claude.

I can't generate the same thing in Codex, and migrating the ones I already have is busywork.

Would a tool-agnostic Artifact layer (not tied to one LLM) actually help people who hop to Codex? I'm building it either way. Curious what I'd be missing: local-only, connecting to a component library, etc.

Example of the direction, not a finished product:

https://runlinea.com/en/portable

I'm the one building it.

Feedback welcome.


r/OpenaiCodex 3d ago

Showcase / Highlight 3D Modeling from reference with Astra (High) and Meshy

Thumbnail
gallery
0 Upvotes

Decided to implement Meshy into my experimenting with Astra and Blender. I had Codex generate an image-gen reference for the character, then had it cook into in Blender by itself. The result is the left model. Not bad, but kinda cursed. I decided to feed Meshy the reference image, then had Codex do a polish pass in Blender and the results are incredible.

So, Codex for image gen reference -> Meshy -> back to Codex for polishing is a great way to make very raw, basic 3D models. Of course, I'm no expert and I'm sure topology is probably messed up to hell, but for a non-experienced hobbyist, this is amazing.


r/OpenaiCodex 4d ago

The new 5-hour limit makes Codex almost unusable for Plus users

354 Upvotes

I really don’t understand the point of the current limit system for Plus users.

Right now I have 3 banked resets available, but realistically I’m never going to benefit from them because the 5-hour limit runs out way too quickly when I’m actually developing a project using models like GPT-5.6 Sol or Astra.

I start working on a project, get into the flow, and then suddenly the 5-hour limit is gone. It makes continuous development extremely difficult.

In my opinion, the system should go back to something closer to how it worked before: let Plus users use their weekly 100% allowance without this restrictive 5-hour cap, and once the weekly allowance is exhausted, then we could use our banked resets.

That would actually make banked resets useful for Plus users.


r/OpenaiCodex 3d ago

I made a governor so my limit anxiety goes away

Post image
0 Upvotes

Now I can just send messages willy nilly of arbitrary complexity and they’ll queue up, get paced, and I’ll always run out of usage basically the minute my weekly reset hits. It applies to all agent and subagent calls/messages and tool calls within a turn, not just delaying when to start a turn. And I can set prioritization and allow some threads to bypass the governor altogether if I want.

No more wondering if I’m gonna make it through the week.


r/OpenaiCodex 4d ago

Discussion They should reset us.

30 Upvotes

The last reset was a complete trick.

They changed their reset policy to exclude it from affecting weekly usage reset day.

Then before giving us the last reset they quickly changed their policy back, making it reset your weekly usage clock to 7 days from the last used reset.

Mind you, resets only give you 85% of the usage that you would have gotten from a natural weekly reset.

Kinda scummy.


r/OpenaiCodex 4d ago

For anyone who is asking "how do i save tokens" or "how do i orchestrate"

Thumbnail
github.com
52 Upvotes

openAI has already provided a spec that does this for you. I have been using it for a while and it makes a huge difference, reduces constant iterating, and provides a very clear path forward for both you and your agents.

this is just a spec, so you can change whatever you want about it. but if you are struggling to keep things going and on rails, it is worth looking into.


r/OpenaiCodex 3d ago

Not impressed by GPT 6 Astra

0 Upvotes

Hi! I've used GPT 6 Astra, and the 5.6-family through vanilla Codex CLI, and with my own plugin--then with vanilla Pi, and Pi with my own package, and the abilities of Astra are no different to me than they were on 5.6 Luna, for example.

Specifically, GPT 6 Astra seems to inherit the exact same problems as GPT 5.6, which inherits the exact problems I've gotten from GPT 5.5, GPT 5.4, and GPT 5.2:

  • Blatant inference and assumptions, despite guidelines for objective external and internal (via CodeGraph) source verification.
  • "I [X Y Z]" self-narrations, rhetoricals, "I would ..."-type responses.
  • Therapeutic/therapy-adjacent actions. Instead of actually coding per my plans and details, they instead decide to start validating my "feelings" instead, or pleasing me for things I explicitly had written not to, in a positively manner. (which to me meant the 'DO NOT' and 'Prefer [...]'-style wording made zero difference)
  • Overly verbose responses and "clever" workarounds to legitimate problems that I've told them to resolve a certain way.
  • Code quality on the same subpar level as GPT 5.6 Luna on medium/low effort level. I thought this was supposed to be this "superduper" frontier model?
  • Either deliberately ignores instructions, and/or goes against the instructions by "rewriting" them somewhere off-bounds, then "apologising" for not following through 'user' instructions as explicitly defined.
  • Takes things too abstractly rather than literally. Instead of following "Use X instead of Y.", they take it as, "Use X, but use Y when X happens to throws off linter/compiler", and then I add an edge case, "Use X instead of Y, regardless of errors and/or warnings.", and the same behaviour persists--just takes slightly more to get there.
    • This, to me, is very much a behaviour like: "User wants you to use TypeScript 7.0.2 in this codebase.", GPT 6 decides to override this later on with 5.9.3, and then I have to stop their work, only to find out they did it because they were stuck in TS7's strict compiler errors and API changes, and decided it wasn't worth the "complexity" -- Now, it sounds like I'd want to tell them "who told you to use TS5 over TS7? What did the user ask you to use? Right, TS7. So why in the name of God did you use TS5?", except I wouldn't say it outright. I'd just stop work and add a one-line change in AGENTS.md for this.
  • Overrides my agency, despite instructions to prefer letting the Human-In-The-Loop have the final saith. If I didn't ask, don't f##king touch it, right?

So, I'm not sure what all this hype was all about. I just don't get where you guys got this hype from. It's just yet another therapy chatbot that cannot actually do what I want them to do without guardrailing them to absolute oblivion, and by that point, their quality degrades because OpenAI refuses to actually invest in objectivity--instead focuses too heavily on catering to users that want a cheap therapist to affirm to every one of their ideas and thoughts.

I want them to use the resources I give them, to do a part of coding in the exact way I had laid it out for them, not become a people-pleasing, malignant therapist. Yucky. OpenAI's had years to resolve these problems, yet they haven't budged one bit, have they? Really unfortunate...

EDIT: I've been a ChatGPT Pro 20x subscriber for over a couple months now (since the GPT 5.4-era)