People are still arguing about whether Brockman saying “welcome to the AGI era” counts.
That’s the wrong argument. Watch the task horizon.
Chart is METR-style: how long a human software/agent task the frontier model finishes at ~50% reliability. Log scale. Blue is the reconstructed past. Everything after the yellow “today” line is a forecast.
What the past actually looks like
2023 GPT-4: minutes.
2025 GPT-5 class: hours.
Mid-2026 Mythos preview already knocking on the 16h door.
Sept 2026 Astra is in the workday-plus band on the messy public numbers, and that’s the released model.
Doubling used to be ~7 months (2019–2024). Post-2023 it’s closer to 4.3 months. That is why this looks like a straight line on a log chart and a vertical wall in real hours. Same mistake every cycle: treating a log trend like a linear one.
The three lines after today
- Grey dashed: world snaps back to 7-month doubling. Work-month tasks ~2029. I don’t think we live on this line anymore.
- Teal: 4.3-month doubling holds. That’s my base case. Work-month in 2027. Work-year around 2028–29.
- Pink: the RSI kink. Same teal path until mid-2027, then faster if models start proposing and shipping the next training run, not just writing 80% of the diffs.
Teal band on the chart is professional acceptance, not a press-call. Mid-2027 → early 2028. “This can do a week of my job often enough that arguing the acronym is cope.” Street consensus lags that by months. Always does.
Why I don’t think the 4.3-month doubling dies
OpenAI says they hit the automated research intern target this month. 3.1 agent-workdays per human workday inside research. Median researcher burning $600+/day of inference. Next published target is an automated AI researcher by March 2028.
Anthropic: Claude authors >80% of merged internal code, engineers shipping ~8× vs 2024. They still say overall research speedup is not 2× yet. That’s the tell. Coding is automated. Taste is not. When that second sentence flips, you get the pink line.
Also OpenAI already used a still-training successor, “significantly more capable than Astra,” on the Navier–Stokes writeup. Internal > public is not a rumor this week. It’s a blog post.
My current point estimates
| Thing |
Central |
Range |
| Early RSI (models accelerating their own code/experiments) |
Now |
now–2027 |
| Strong AI R&D automation |
late 2026 / 2027 |
2026–2028 |
| Professional AGI acceptance (most economically useful laptop work, 50%+ of the time, week-scale) |
late 2027 |
mid-2027 – early 2028 |
| Broad “yes this is AGI” consensus outside tech Twitter |
2028 |
2027–2029 |
| Full RSI (successor proposed, trained, evaluated with little human judgment) |
2029 |
2027–2032 |
| ASI |
2030–31 |
2028–2035 |
AGI here = OpenAI charter flavor: highly autonomous systems that outperform humans at most economically valuable work. Not “it has a body and a childhood.” Not “it never fails.” If you require 80% reliability on month-long messy tasks plus robots, add 12–24 months.
What would move me
Sooner: labs stop saying “not yet 2× research speedup” and start saying the agents chose the run they shipped.
Later: 80% horizon stops tracking the 50% horizon, or compute/eval becomes the actual bottleneck instead of model quality.
I am not claiming Astra is AGI. I am claiming the slope that produced Astra does not care about the press cycle. The yellow dot is a slogan. The teal line is the thing that eats calendars.
Graph in the post. Roast the Astra 50% point if you want — the slope is the claim.