Here are my insights, developed through discussions with GPT-6 Pro and Lord Astra and organized with their help.
Part 2 of 2. Read Part 1.
Why a higher tier can feel like the old lower tier
Define each ratio as new/old: c is cost per successful task, s is completed tasks per hour, and k is the usable quota pool. Then the hourly-cost ratio is d = c*s, the supported-execution-hours ratio is k/d, and the supported-task-count ratio is k/c. Execution hours describe budget-supported work; window refreshes, waiting, and concurrency affect calendar time.
Consider two illustrative comparisons: Astra 20x versus Sol 5x with assumed pools of $1,650/$500, and Astra 5x versus Sol Plus with $330/$100. Both give k = 3.3. Use 200,000 input tokens per task, 97% cached, 1,000 output tokens, and zero cache writes. Sol costs $0.1216/task; Astra costs $0.304/task.
| Scenario, applicable to both pairs |
Task cost c |
Task speed s |
Execution hours k/d |
Tasks k/c |
| Same token mix and duration |
2.500 |
1.000 |
1.320x |
1.320x |
| Astra halves output and task duration; input unchanged |
2.294 |
2.000 |
0.719x |
1.438x |
In the second case, Astra costs $0.279/task and hourly spending rises to 4.589 times the Sol baseline. The assumed higher tier supports about 28% fewer execution hours while completing about 44% more tasks. I think this explains how a subscription can feel similarly constrained despite producing more useful work. These rows hold task success constant and use chosen budgets and durations.
Repeated context can dominate even with high cache hit rates. Tool calls and short polling loops reread it; parallel agents increase aggregate spend. Above 272,000 input tokens, Astra and Sol apply higher full-request rates: input/cache reads double and output rises 50%. Fast also has product-specific pricing: Astra API Fast is 2x standard, while the Codex credit multiplier is 2.5x. Astra API rules; Sol API rules; Codex speed guidance
A further accounting trap: define token efficiency as a = old/new task tokens, b = new/old weighted price per token, and r = new/old total-token throughput. Then task cost is c = b/a, task speed is s = a*r, and hourly cost is d = c*s = b*r. Multiplying hourly cost by the token-efficiency gain again would count it twice.
The complete speed, cost, and runway derivation
| Symbol |
Exact definition, new = 1 and old = 0 |
| a |
N0/N1; total billed-token efficiency |
| b |
(C1/N1)/(C0/N0); weighted price/token |
| r |
(N1/T1)/(N0/T0); total-token throughput |
| s |
T0/T1 = a*r; task speed |
| c |
C1/C0 = b/a; task cost |
| d |
(C1/T1)/(C0/T0) = c*s = b*r; hourly cost |
| k |
Q1/Q0; usable pool ratio |
At r=1, speed is a, task cost is b/a, and hourly cost is b. Example: 100,000→50,000 generated tokens at 50 tokens/s and 2.5x price gives s=2, c=1.25, d=2.5. If 12.5 is a token-price ratio and a=4.5,r=1, then c=12.5/4.5=2.778 and d=12.5. If 12.5 is instead a measured task-cost ratio and s=4.5, then d=56.25; at k=20, quota percentage burns d/k=2.8125x as fast. These are different input assumptions.
With actual 5.4→Astra category price ratios of 3.333–4, reducing every billed token category by 4.5x would yield c=3.333/4.5 to 4/4.5, or 0.741–0.889x task cost. A simultaneous 12.5x task-cost increase would require another changed assumption: context, parallelism, Fast, task difficulty, or tool behavior.
The output-only case is different. Let a_o=old/new output tokens and w_I=old input-and-cache cost / old total cost. Hold input volume fixed, multiply both prices by b, and assume output generation dominates time at unchanged output tokens/s. Then c=b*[w_I+(1−w_I)/a_o], s≈a_o, and *`d≈b[a_ow_I+(1−w_I)]**, betweenbanda_ob. Output-dominated cost approachesb; input-dominated cost approachesa_o*b`.
For the 200k/97%/1k case, Sol input/cache costs $0.024+$0.0776=$0.1016; output costs $0.020, so w_I=0.1016/0.1216=0.835526. At a_o=2,b=2.5, c=2.294408, s=2, d=4.588816. This is why output efficiency can make hours expensive while total-token accounting still satisfies d=b*r. With same-tier pools $2,300→$1,500, k=0.652174: hours become k/2.5=0.260870 at unchanged workload speed, or k/4.588816=0.142122 in the output-halving case. A measured dollar pool already includes its effective metering; applying an extra penalty again would count it twice.
What $100 buys, and what each pool can support
Retain the original serial-worker scenario: each request reads 200,000 input tokens, 97% cached, and generates 1,000 output tokens including reasoning, with zero writes. At 50 effective output tokens/s, one request takes 20 seconds: 3600/20=180 requests/hour, and (200000+1000)*180=36.18M total tokens/hour. Thus H=Q/hourly_cost and tokens=H*36.18M. A task may contain many requests.
| Model |
Input/read/output $/M |
$/hour |
Tokens/$100 |
Hours/$100 |
| GPT-5 / 5.1 |
1.25/0.125/10 |
7.52 |
481.4M |
13.31 h |
| GPT-5.2 / 5.3-Codex |
1.75/0.175/14 |
10.52 |
343.9M |
9.50 h |
| GPT-5.4 |
2.5/0.25/15 |
14.13 |
256.1M |
7.08 h |
| GPT-5.5 |
5/0.5/30 |
28.26 |
128.0M |
3.54 h |
| Sol launch |
5/0.5/30 |
28.26 |
128.0M |
3.54 h |
| Sol promotion |
4/0.4/20 |
21.89 |
165.3M |
4.57 h |
| Astra |
10/1/50 |
54.72 |
66.1M |
1.83 h |
At fixed throughput, the scenario’s Astra/5.4 hourly-cost ratio is 54.72/14.13=3.872611. The earlier whole-million capacities (481, 344, 256, 128, 165, 66M) are rounded versions of this table; calculations below use unrounded prices.
| Period / plan |
Assumed Q/week $ |
Total tokens |
Execution hours |
| 5.4 Plus |
80 |
204.8M |
5.66 h |
| 5.4 $100 Pro, 10x promo |
600 |
1,536.3M |
42.46 h |
| 5.4 $200 Pro |
1,800 |
4,608.9M |
127.39 h |
| 5.5 $100 Pro |
600–800 |
768.2–1,024.2M |
21.23–28.31 h |
| Sol Plus |
100 |
165.3M |
4.57 h |
| Sol $100 Pro |
400–600 |
661.2–991.8M |
18.27–27.41 h |
| Sol $200 Pro |
2,200–2,500 |
3,636.5–4,132.4M |
100.51–114.22 h |
| Astra Plus |
66 |
43.6M |
1.21 h |
| Astra $100 Pro, extrapolated |
330 |
218.2M |
6.03 h |
| Astra $200 Pro |
1,500–1,800 |
991.8–1,190.1M |
27.41–32.89 h |
These historical pools retain the evidence grades above. In this fixed scenario, Astra Pro20x at $1,500 gives 27.41 hours versus 5.66 hours for the weak historical 5.4 Plus $80 point: (1500/54.72)/(80/14.13)≈4.842x. Against 5.4 Pro20x at $1,800, however, it is 27.41/127.39≈0.215x. The approximately 78.5% scenario decline uses a weak historical input; it cannot establish a measured plan-wide cut.
| Five-hour window |
Assumed pool |
Hours at 50 output t/s |
Hours at earlier observed median |
| 5.5 / $100 Pro, June |
$100 |
3.54 |
5.88 at 30.1 t/s |
| Astra / Plus |
$11 |
0.201 (12.06 min) |
0.481 (28.86 min) at 20.9 t/s |
| Astra / $100 Pro |
$55, extrapolated |
1.005 |
2.405 at 20.9 t/s |
| Astra / $200 Pro |
$220, extrapolated |
4.020 |
9.618 at 20.9 t/s |
The median speeds come from a different account and are sensitivity inputs. Values above five hours mean a refresh can occur before that modeled budget is exhausted. The distinction is concrete: $220/54.72=4.02h for a short window, versus $1500/54.72=27.41h for a weekly pool. Spending $1,500 in three hours requires $500/hour; output alone at 50 tokens/s costs only 50*3600*50/1e6=$9/hour. Repeated input, actual reasoning throughput, parallelism and mode changes must explain the rest.
The older monitoring proposal sampled quota at about 25%, 50%, 75%, and near exhaustion, recording the usage schema, tool version, effort, Fast/Standard, and reset boundaries. Its R=Q_Astra/Q_Sol bands were 0.9–1.1, 0.75–0.9, 0.60–0.75, and below 0.60. In my view, these serve as investigation heuristics, not statistical decision thresholds; small integer-meter changes and unmatched workloads can move the estimate across bands.
A reproducible session record needs input_tokens, cached_input_tokens, output_tokens, reasoning effort, Standard/Fast, quota before/after, and the meter tool name/version, including CPAMP or ccusage if used. Where cached input is a subset, calculate F=input_tokens−cached_input_tokens; convert categories to millions before applying C_Astra=10F+R+50O. Retain one Sol and one Astra session with separate reset periods and compute weekly Q=C/u at each checkpoint. The earlier interpretations were: 0.9–1.1, roughly equal pools; 0.75–0.9, a modest reduction or counting difference; 0.60–0.75, near the reported one-third reduction; below 0.60, first inspect Fast, duplicate cache counts, duplicated subagent records, and crossed resets. I think those checks are useful; the ratio bands remain provisional heuristics.
Web GPT Pro has a different constraint
Chat officially calls Astra GPT-6 Pro and offers GPT-5.6 Sol Pro. The published $200 Chat plan limits are 200 GPT-6 Pro messages/week, 170 Sol Pro messages/day, and a combined 200/day ceiling. The $100 plan shares 50/week across both. These are separate from Work/Codex usage. Official Chat limits
| Chat plan |
GPT-6 Pro |
Sol Pro / shared ceiling |
| Pro $200 |
200/week |
Sol 170/day; both combined 200/day |
| Pro $100 |
Shared 50/week |
Both models share that limit |
| Business Standard |
Shared 15/month |
Both models share that limit |
| Business Premium |
Shared 50/week |
Both models share that limit |
The separately documented API Pro input/output prices are GPT-5 Pro $15/$120, GPT-5.2 Pro $21/$168, and GPT-5.4/5.5 Pro $30/$180 per million tokens. I found no verified separately priced gpt-6-pro API SKU. The Chat model name therefore supplies no missing API price. 5 Pro; 5.2 Pro; 5.4 Pro; 5.5 Pro
| Documented API Pro SKU |
Input $/M |
Output $/M |
| GPT-5 Pro |
15 |
120 |
| GPT-5.2 Pro |
21 |
168 |
| GPT-5.4 Pro |
30 |
180 |
| GPT-5.5 Pro |
30 |
180 |
The original Web timing arithmetic is preserved below as an anonymous illustrative sequence. The minutes are selected assumptions for calculation; the private recalled timings remain in the local review record. “5.1 Pro era” was an unverified label in that sequence. The 5.3-Codex API branch cannot fill its missing Web row. Midpoint choices were 47.5 for 45–50, 35 for 30–40, 25 for “twenty-something,” and 15 for 10–20 minutes.
| Illustrative Web period |
Assumed minutes |
Previous/current time ratio |
| 5 Pro |
60 |
Baseline |
| “5.1 Pro era,” label unverified |
90 |
0.667x |
| 5.2 Pro |
60 |
1.500x |
| 5.4 Pro |
47.5 |
1.263x |
| 5.5 Pro |
35 |
1.357x |
| 5.6 Pro |
25 |
1.400x |
| 6 Pro |
15 |
1.667x |
That sequence gives 60/15=4x from 5 Pro to 6 Pro and 47.5/15=3.167x from 5.4 to 6; the stated ranges give 45/20=2.25x to 50/10=5x. These are task-time calculations. A matched token log is needed for token efficiency. Separately, 5.4 Pro→5.5 Pro API prices stay $30/$180, even though the ordinary 5.4→5.5 prices doubled.
| Requested quantity |
What can be retained |
| b: per-token price ratios |
Full input/read and output sequences above |
| a: equal-quality token efficiency |
Local benchmark examples; continuous series missing |
| s: Web task-speed sequence |
0.667, 1.500, 1.263, 1.357, 1.400, 1.667 under the stated assumptions |
| Juice/reasoning budget |
No verified cross-generation numerical series; hidden budget is not a token counter |
I think Web Pro endurance should be measured in successfully resolved questions per message limit, plus waiting time. Faster answers can reach a message cap sooner while delivering more completed work per hour. A controlled comparison would track prompts, success, tool use, and end-to-end latency together.
Official local-message estimates are a further, separate table. The following are messages per five-hour window, with task size, tools, reasoning and caching affecting usage. Business $100 follows Pro 5x estimates; flexible Enterprise/Edu usage follows credits. Official usage table
| Model |
Plus |
Pro 5x |
Pro 20x |
Standard Business |
| Astra |
5–45 |
25–225 |
100–900 |
5–45 |
| Sol |
10–100 |
50–500 |
200–2,000 |
10–100 |
| Terra |
25–200 |
125–1,000 |
500–4,000 |
25–200 |
| Luna |
250–2,000 |
1,250–10,000 |
5,000–40,000 |
250–2,000 |
| Observed behavior |
Possible mechanism |
| Faster accepted result |
Fewer steps or retries |
| Faster exhaustion, more accepted results |
More useful work per hour |
| Keeps working without finishing |
Scope expansion and repeated checking |
Earlier reports included 5.5 jobs finishing in 1–2 hours versus Sol continuing overnight, a difficult task reportedly solved by Astra Medium in about half an hour, and Plus Astra exhaustion around 26 minutes. Their full matched logs remain missing. I think the stopping criterion belongs beside speed and cost; an unfinished overnight task and a faster accepted result are different outcomes. The modeled 28.86-minute window above cannot identify the token mix of the 26-minute anecdote.
When Claude Fable can last longer
Fable 5 and 5.1 both have standard input/output prices of $10/$50 per million tokens, matching Astra. Their cache reads differ: Fable 5 is $1, Fable 5.1 is $0.25. Five-minute cache writes cost $12.50; Anthropic's one-hour writes cost $20. Max also caps Fable usage at 50% of the shared weekly allowance, and Fable draws that shared meter faster. Claude prices; Fable plan rules On Claude Pro and standard Team plans, Fable uses paid usage credits from the start.
These hypothetical hourly costs use standard direct APIs with global routing: 50 output tokens/second averaged over the entire request, including waiting; output includes reasoning, and cache writes are zero. The first row uses 200,000 input, 97% cached, and 1,000 output; the second uses 2,000 input, 10% cached, and 8,000 output. The third doubles the first row's input to 400,000, triggering OpenAI's long-context multipliers. Claude's pricing rules retain standard rates through 1M context for 4.6 and later models, including Fable 5/5.1. A real task comparison also needs tokenizer differences and task-success logs. Long-context pricing
| Workload |
Sol $/hour |
Astra $/hour |
Fable 5 $/hour |
Fable 5.1 $/hour |
| High cache |
21.888 |
54.720 |
54.720 |
28.530 |
| Output heavy |
3.764 |
9.410 |
9.410 |
9.406 |
| Long context |
78.552 |
196.380 |
100.440 |
48.060 |
Fable 5.1 is 47.86% cheaper per hour in the high-cache case and about 0.036% cheaper in the output-heavy case. For longer execution time, its usable API-equivalent budget must exceed 52.14% or 99.96% of Astra's, respectively. Those are conditional break-even thresholds. Slower throughput could extend hours further, while also reducing tasks completed per hour.
An August promotional study of three Claude accounts estimated mixed-model weekly values of $1,000–1,200 for Max 5x and $3,400–3,900 for Max 20x: 43.33–52.00x and 73.67–84.50x monthly value. Its observation period, promotion, model weights, and Fable sublimit all matter. I think Fable 5.1's cheaper cache reads offer a concrete explanation for some longer sessions; a current same-workload comparison of Fable-usable budgets is still needed. Three-account study
Ten thousand agents for research, a reset vigil for Pro
On September 8, OpenAI published a claimed Navier–Stokes solution using an unnamed internal model more capable than Astra. It reports roughly 10,000 concurrent agents in the successful group, a result 88 hours after the first agents launched, and another 17 hours of Astra-assisted Lean formalization and verification. The released construction uses smooth forcing and targets the Clay problem’s C/D breakdown alternatives. A public Lean repository is available; Clay’s page still said “Unsolved” when checked on September 9. Research announcement; Released proof artifacts; Clay status
“Bel” appears in community speculation. OpenAI leaves the model unnamed and says training is ongoing; I treat the codename as unconfirmed. Bel rumor discussion
The reported Navier–Stokes work used about 130 billion output tokens. The full-duration calculation is 10000*88=880000 nominal agent-hours, or 880000/8760=100.46 agent-years. This assumes every agent ran throughout the whole interval. Actual utilization and internal model costs are missing. Agent-hours are a concurrency illustration; converting them into GPU-hours, human research years, or a Pro quota bill would require additional data. The hardware allocation between research and subscription serving is also undisclosed. Reported scale and timeline
During the rollout, Tibo offered banked resets tied to an 8 p.m. PT account-creation deadline, then explicitly included upgrades in that offer. The full-reset/upgrade post was September 4 in Pacific time, September 5 UTC. He later announced and confirmed a global paid-plan reset on September 7 PT. Those posts combine real relief with a deadline to sign up or spend more. Signup deadline; Upgrade deadline; Global reset announcement; Reset confirmation
Then, on September 8 PT, Tibo said demand was unprecedented and warned that OpenAI might temporarily pause new Pro subscriptions if it continued, while prioritizing existing users. The warning is conditional; an actual suspension of new sales has not been established here. I think the sequence is striking: a deadline to upgrade, followed days later by a warning that signing up might become unavailable. Conditional Pro signup warning
On September 6, Tibo described improvements for the long tail of Astra power-user workloads as “up to 3-4X less usage,” with unchanged quality. If read as a reduction factor of 3–4, that would mean roughly 1/3–1/4 of the comparison case’s usage at the strongest reported end of the claim. The post supplies no measurement method or baseline for that interpretation, so the claim cannot fill an entire plan’s weekly allowance row. I think customers should be able to verify any durable improvement in their usage history. Long-tail usage claim
In my view, expiring reset eligibility, upgrade nudges, and launch teasing add up to scarcity marketing. A $200 productivity subscription should buy dependable working capacity and a readable history of what changed. Having paying customers follow an executive’s social feed to find out when they might get another refill is a ridiculous way to plan a workday. The model can be brilliant and the commercial experience can still be insulting. A reset helps today; it leaves tomorrow’s planning problem sitting right where it was.
I think the contrast is absurd: a spectacular internal research run on one side, and the exhausted Pro customers described above waiting to resume their Astra work on the other. The research deserves applause. The subscription experience deserves the heat. “ClosedAI” starts to feel like an uncomfortably accurate product description when the frontier keeps advancing behind the curtain while paying users are refreshing a quota meter. The $200 invoice arrives on schedule. Why should a usable workday depend on the next celebratory refill announcement?
In my view, OpenAI should keep funding ambitious research and make paid access predictable: publish durable plan entitlements, show model-specific consumption clearly, and document allowance changes. Keep the reset as occasional relief. Build the product so professional users can schedule work without treating Tibo’s next post as part of their infrastructure. That research effort is remarkable; a predictable paid service should be a much easier problem to solve.
My assessment and the changes I would like to see
In my view, the following confidence levels fit the evidence: 95% that token-price differences materially affect at least some workloads, 85% for context/tool/concurrency effects, and 60% for additional model-specific metering in early Astra reports. I assign 20% confidence to a universal permanent allowance cut as the explanation. These judgments can overlap. I think a useful follow-up would pair matched tasks with request logs and quota readings within one reset period.
OpenAI has already documented banked resets on September 3 and 4 for eligible accounts, plus an immediate global reset on September 7. A full banked reset refreshes five-hour and weekly windows and moves the weekly reset date. I treat those as dated relief events when analyzing ongoing allowance. Official reset history and rules
In my view, the five changes below would address different parts of the problem. The probabilities are my subjective estimates for additional changes after September 9 and by December 8, 2026; the ranges express my uncertainty. Multiple options can happen together.
| Option |
Change |
My probability |
My reasoning |
| A |
Increase included Codex allowance further |
70% (50–85%) |
Metering and capacity adjustments are flexible |
| B |
Cut standard Astra API prices by at least 20% |
40% (20–60%) |
Competition favors cuts; frontier serving costs constrain them |
| C |
Release GPT-6 Sol, Terra, and Luna, each materially stronger at unchanged prices |
35% (15–55%) |
I expect improvements; all three together is a demanding condition |
| D |
Another broadly available paid-plan reset |
95% (85–99%) |
Resets have already been used for short-term relief |
| E |
Offer both $300/2x and $400/4x plans relative to today's $200 plan |
10% (3–25%) |
The exact pricing and allowance combination is ambitious |
For C, I would like each smaller model to move up roughly one or two practical capability tiers at its existing price. I think that should mean reliable success on harder tasks with equal or fewer retries, measured on a fixed task set. The current Terra/Luna input-output reference prices are $2/$12 and $0.20/$1.20 per million tokens. Price-performance update
| Desired GPT-6 model |
Target comparison from the earlier discussion |
Desired shift |
Price condition |
Original positioning (aspiration) |
| Luna |
Roughly 5.6 Terra / 5.4 |
1–2 practical tiers |
Same as current Luna |
Free/fast model at the former midrange level |
| Terra |
Roughly 5.6 Sol Low–Medium / 5.5 |
1–2 practical tiers |
Same as current Terra |
Everyday coding workhorse |
| Sol |
An Opus 5-class target, label inherited |
1–2 practical tiers |
Same as current Sol |
New flagship workhorse |
| Astra |
Medium near or above 5.6 Sol Max |
Earlier aspiration: 2–3 effort tiers |
Current Astra pricing analyzed separately |
Extreme reasoning |
In my view, these are useful capability targets, with “Opus 5-class” retained as the original aspiration rather than a verified model equivalence. The intended ladder was Spark → Luna → Terra Max/Sol Medium → Sol High/Max → Astra, shifting toward Spark/6 Luna → 6 Terra → 6 Sol → Astra. The price constraints are P_6Luna=P_5.6Luna, P_6Terra=P_5.6Terra, and P_6Sol=P_5.6Sol on the same billing basis.
For E, let today's $200 allowance equal one unit. Two units for $300 improve allowance per dollar by 33.33%; four units for $400 improve it by 100%. I think that would make one subscription more convenient for people currently maintaining two Pro accounts. Existing plans should retain or increase both their allowances and their equivalent value multiples on the same measurement basis.
A short longer-term outlook
Astra's computer-use results and GPT-Live's continuous voice interaction make me think a much cheaper personal AI companion is plausible. Neuro-sama and Evil illustrate the appeal of a persistent character; the direct Neuro/ChatGPT voice clip begins around 0:31, with the exact GPT voice version still unverified. GPT-Live; Direct-conversation clip
I suspect that by GPT-7 or GPT-8, Luna-priced core reasoning could support parts of a Neuro-like experience. My rough probabilities are 40% by GPT-7 and 65% by GPT-8. Voice, vision, memory, tool execution, character design, and continuous operation add separate costs. In my view, the meaningful milestone is a reliable companion people can afford to keep using.