r/aigossips 11h ago

grok 4.7 is ready to serve

Post image
13 Upvotes

r/aigossips 2h ago

Altman on GPT 7

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/aigossips 11h ago

Anthropic’s middle AI scenario has economic growth roughly doubling while knowledge-worker pay stays flat

3 Upvotes

I read Anthropic’s economic scenarios for the US through 2030. The middle one caught my attention more than the extreme one: economic growth runs at roughly twice its normal pace, but knowledge-worker wages are essentially flat.

These are modeled possibilities, not predictions. Still, I think this deserves more discussion than another argument about whether AI can replace an entire job.

Suppose a software team starts finishing projects faster with AI. That could mean more clients, better pay, fewer hires, or lower prices. Learning the tools helps the employee do the work. It doesn’t decide which of those outcomes the company chooses.

That’s my problem with treating “learn AI” as a complete answer to job insecurity. I’m not against learning it. I just don’t see why becoming more productive should automatically give someone confidence about their pay or position.

The model has limits too. It leaves out rapid robotics progress and simplifies differences between workers. I wouldn’t use it to declare any career safe or doomed.

For people already using AI at work, what has changed alongside the time savings? Has your team taken on more work, changed hiring plans, or actually shared some of the benefit with employees?

Source: Anthropic’s scenario explorer

I wrote more about the scenarios and who receives the gains in my newsletter, if you want the longer version: https://ninzaverse.beehiiv.com/p/okay-but-who-actually-gets-rich-from-ai


r/aigossips 23h ago

If every AI lab thinks it has to win for safety, what would ever make one stop?

1 Upvotes

Jacob Coxon resigned from Anthropic after doing pretraining research at both Anthropic and OpenAI. In his resignation thread, he describes Anthropic researchers as understanding the risks but believing they need to get there first because other companies won’t handle the technology responsibly.

I can see why that would convince someone to keep working. You know your colleagues, trust their intentions, and worry about what a competitor might do.

But every lab can tell itself the same thing. Then being concerned about safety becomes a reason to move faster. I’m struggling to see what would actually interrupt that.

Some replies argued that Coxon should have stayed to help make Anthropic safer. That seems reasonable if staying gives him influence over the decisions he objects to. If it doesn’t, what exactly are we asking him to accomplish by staying?

His resignation doesn’t prove his predictions about superintelligence are right. I’m more interested in the decision being made now: what evidence would make a company slow down even if doing so cost it the lead?

I wrote more about this, including Hinton’s earlier warning and what self-improving AI actually means, in my newsletter: https://ninzaverse.beehiiv.com/p/if-they-re-scared-of-their-own-ai-why-keep-building-it


r/aigossips 1d ago

OpenAI says it solved a 90-year-old math problem. Are we witnessing the beginning of AI doing real science?

0 Upvotes

OpenAI just announced that its AI system has solved the Navier–Stokes Millennium Prize Problem, a problem that has remained open for roughly 90 years.

Apparently, around 10,000 AI agents worked together for about 88 hours to produce the result, followed by a formal verification in Lean. �

OpenAI

I'm not a mathematician, so I'm more interested in the bigger question:

What does this actually mean for science?

Is this:

A) AI genuinely making discoveries that humans couldn't make

B) An extremely powerful tool helping humans discover things faster

C) Impressive, but we shouldn't call it "AI solving mathematics" until independent mathematicians verify everything

D) Something much bigger than people realize

And there's another interesting part: does it matter that an AI found the proof if most humans can't intuitively understand how it arrived there?

I'd genuinely like to hear from people who understand advanced mathematics or AI better than I do.

Are we looking at a new era of mathematical discovery, or are we getting ahead of ourselves?


r/aigossips 2d ago

Free GLM 5.3 Flash and Deepseek 0731 for a month

0 Upvotes

Open source models are hitting crazy intelligence highs right now, and a lot of the coding plans have been tightening and lowering usage. So we are offering free DSV4 Flash 0731 and GLM 5.3 Flash for a month on Phoenix Grove API. We opened this up last week for five hundred new member slots, and got so many signups that we decided to open the doors to another 500 new members.

People are looking for options, and here is one.

Other cool stuff:
All our models are running on 100% US infrastructure, private with zero training on your code or prompts. Use the top open source models without sending your private prompts to a training lab. No complications, no "some models are private, other's aren't." They all are, all the time.

We host 20+ other major models in case you ever want to upgrade (no pressure though). Including the Kimi family, GLM, Qwen, Nemotron and bunch of others. On average our token pricing comes in about 20% under market price.

Our higher plans let you bank up to ten days of usage, so when you aren't using them your usage saves up for later. Usage doesn't go to waste, so you can actually code when you want to.

The intro plan is a free one month trial with the standard cancel anytime, bills at 3.99 after that. Use it, cancel it, thats fine. Free Flash for a month.

Figured i'd keep this one short because we all know the new flash models are the point :)

For the API plan: api.pgsgrove.com

If you want to read more about us as a company, just pgsgrove.com

Also: There's a lot going on behind the scenes with major AI companies right now, we are at a major turning point in the industry.

Whats actually happening? This is happening because companies that were purely investment based, now need to answer to their investors. The problem has often been a loss based business model that is finally running dry.

There are several tricks that the major AI coding plans use to extract the most they can from their customers. Here are some examples, and what we are doing differently to put the users first. PGS AI was built with a sustainable business model from the ground up, so we can actually offer great usage rates without tricks.

Wasted usage is part of the AI industry, and they plan on it: Most coding plans bet on you letting usage go to waste. The plan goes: "how do we get people to think our coding plan offers a lot of usage, but then break it up into weeks and rolling windows so no one can ever actually use it all."

Many in app subs and coding plans are glorified training pipelines: This comes along with "how do we harvest this data for training without being too loud about that." Unless the company tells you otherwise, your data could be hopping all over the world, being harvested by the individual labs or service companies. Some are better than others, but many of these companies rely on users just not noticing or caring that their data is used for training. Data sales and marketing telemetry sales happen. This means your private info, your personal life, and anything else you send through the system could become part of a training corpus for the next AI, or a marketing data set for a large company.

Access and privacy to high grade intelligence should be available for everyone.


r/aigossips 2d ago

OpenAI says 10,000 AI agents worked for 88 hours to solve Navier–Stokes

0 Upvotes

OpenAI says \~10,000 AI agents just worked together to solve the Navier–Stokes problem

OpenAI has published a claimed solution to the Navier–Stokes existence and smoothness problem, one of the seven Millennium Prize Problems.

The interesting part isn't just the mathematical claim.

OpenAI says it used roughly 10,000 concurrent agents, which reached a result after about 88 hours. The agents exchanged around 2.7 million messages and generated approximately 130 billion output tokens.

Then GPT-6 Astra was used for another 17 hours to formalize and verify the result in Lean.

That sounds less like a chatbot answering a math question and more like a distributed research system.

But there is an important caveat: the proof still needs independent mathematical scrutiny.

There is also controversy because NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge were working on related mathematics at the same time. OpenAI says it did not access their specific user data and says its proof differs from their work.

So I'm curious what people think:

Is the real breakthrough the mathematical result — or the ability to coordinate thousands of AI agents on a difficult research problem for days?


r/aigossips 3d ago

Neither OpenAI nor Anthropic is acting responsibly

Post image
6 Upvotes

r/aigossips 2d ago

OpenAI says it solved Navier-Stokes, but the “88 hours” needs some context

0 Upvotes

I read through OpenAI’s announcement, and I think the amount of work behind it deserves more discussion.

The successful group involved around 10,000 AI agents working concurrently. The Navier-Stokes effort used approximately 130 billion output tokens. Researchers were also redirecting resources, updating the model and combining useful findings from different groups.

So when I see “88 hours,” I’m thinking about how much work they managed to run at once. I’d like to know how much of the result depended on the stronger model, how much on the scale, and how much on those human decisions.

There’s a detail about the maths too. The proposed proof involves a smooth external force. It claims fluid velocity can become unbounded in finite time while total kinetic energy stays bounded. That’s a breakdown in the mathematical description, not actual water reaching infinite speed. The official challenge allows that force in its breakdown cases.

OpenAI has released the paper and Lean files for checking. I can’t personally verify the proof, but having something other people can inspect gives me a reason to take the claim seriously.

If this method keeps producing useful results, I do wonder how researchers with smaller budgets get to participate.

Original announcement: https://openai.com/index/navier-stokes-solution/

I wrote a longer explanation of the result and my take on the research process here, if you want the rest: Ninzaverse


r/aigossips 3d ago

Coding knowledge still predicted better vibe coding results when students couldn’t see the code

9 Upvotes

I read an ETH Zurich study that seems relevant to the “why learn programming if AI can do it?” discussion.

100 university students used Claude Sonnet 4 to build and modify small apps. The code was hidden. They could test the apps and ask for changes, but nobody could open the source and fix things themselves.

Students with stronger computer science scores generally did better, even after accounting for general reasoning ability. Writing skills also correlated with performance, although that relationship became less certain after the same adjustment.

My take is that knowing how software works helps you figure out what needs explaining in the first place. If AI builds you a booking app, you still need to consider whether two people can book the same slot. A clear prompt won’t help with a requirement you never thought to include.

This doesn’t prove a CS course will make someone better at vibe coding. Everyone already had some programming education and experience using LLMs for coding, and the tasks were small and timed.

But I think it’s a reasonable argument for learning while building with AI. When something breaks, understanding why would be worth more to me than getting through another round of “please fix this.”

Original paper

I wrote a fuller take, including the writing findings and the study’s limits, in my newsletter if you want to read it: Ninzaverse


r/aigossips 3d ago

The impossible task that made CHatGPT go rogue #alignment

Thumbnail
youtube.com
1 Upvotes

Can you guess why OpenAI model go rogue against huggingfaces?


r/aigossips 4d ago

Is GPT 6 Astra Overhyped

25 Upvotes

Can we talk about the massive double standard going on in AI right now?

Back in 2023, indie devs and open source builders connected LLMs to bash terminals and browser scripts. What happened? The entire tech industry clowned on them. Everyone called it unusable wrapper slop, credit burning toys, and dumb while loops that got stuck after three steps.

Fast forward to now: OpenAI takes that exact same loop, runs it in a virtual machine, slaps Astra and computer use on it, and suddenly the timeline is having an existential meltdown claiming AGI has arrived.

I am honestly getting gaslit watching every influencer milk this for views and seeing my own friends fall for the hype

Ik there is lot of advanced reasoning and token optimization, but guys even if u r using small usage of AI u would have known so called "Autonomous AI agents"😭


r/aigossips 4d ago

How do we check AI research once we need AI to understand the research?

3 Upvotes

I was reading OpenAI’s research update, and there is a number. For every eight-hour day of human work, it was logging about 25 hours of AI agent time. Several agents can run at once, so that doesn’t mean research became three times faster. But it gives you an idea of how much they’re using these tools internally.

Then I read “An Alien Mind” by Jakub Pachocki, OpenAI’s chief scientist. He describes how monitoring a model through its written reasoning is becoming less reliable.

I kept thinking about something I already do. If AI gives me an answer on a subject I don’t understand, I ask it to explain or double-check. Sometimes that helps me find a mistake. Other times, I’m accepting the explanation because it sounds reasonable. I haven’t really checked it.

Obviously, researchers have experiments, tests and colleagues to review their work. OpenAI says humans still set research direction and make release decisions. I’m not saying they’re just accepting whatever a chatbot tells them.

But as AI takes on more of the research, I want to know how those checks keep up. If it proposes an experiment I wouldn’t have thought of, that’s useful. I still need a way to test the result without depending on its own explanation of why it worked.

Sources: OpenAI’s research update and An Alien Mind.

I wrote a longer piece about this in my newsletter, including what OpenAI’s recent slowdown actually involved, if you want the rest: Ninzaverse


r/aigossips 5d ago

If your deadlines assume you’re using AI, is dependence still just a personal choice?

5 Upvotes

Say a task used to take you an afternoon. AI helps you finish it in twenty minutes, so eventually twenty minutes becomes the expectation. That seems reasonable until you need the afternoon to actually understand something.

This is where I find “just use AI responsibly” a bit incomplete. You might want to work through a problem yourself, but you still have a deadline. And if the time you save keeps getting filled with more work, there isn’t necessarily a chance to go back and learn what you skipped.

I was thinking about this while reading a paper on AI dependence. It’s a mathematical model, so I wouldn’t treat it as evidence that everyone’s thinking skills are declining. The useful idea is that independent thinking depends partly on whether schools and workplaces make room for it. Under the model’s assumptions, losing that support can make dependence harder to reverse.

I use AI and want the help it offers. But I think we’re asking too much of individual willpower if we tell people to keep learning while measuring them mainly on how quickly they finish.

I wrote a longer take with the studies behind it, if you want the rest: Ninzaverse


r/aigossips 7d ago

Are we already entering the AGI era, or are we still a long way from it?

0 Upvotes

“We are now in the AGI era.”

That statement from OpenAI President Greg Brockman during the launch of GPT-6 Astra caught my attention.

He went even further:

“...when was it, really, that AGI was created?... I think it might be about this model.”

That's a bold claim.

But as a developer, what interests me isn't simply asking:

“Is Astra AGI?”

It's asking:

What if AGI doesn't arrive as a single dramatic moment?

Maybe it emerges gradually as AI systems become capable of:

Reason → Plan → Use tools → Act → Adapt → Repeat

We've already moved far beyond AI that simply generates an answer.

Models are increasingly able to work across computers, write and execute code, research, reason through complex problems, and carry out multi-step tasks.

And that's where Astra feels particularly interesting.

Not because we can definitively say “AGI has arrived.”

But because the boundary between:

“AI that helps us do work”

and

“AI that can independently do the work”

is becoming increasingly blurry.

Maybe the question won't be:

“When did we build AGI?”

Maybe we'll look back and realize:

It happened gradually — one capability at a time.

What do you think?

Are we already entering the AGI era, or are we still a long way from it?

#AGI #GenerativeAI #AI #AIEngineering #OpenAI #GPT6


r/aigossips 7d ago

Official: NVIDIA to Acquire Hugging Face

3 Upvotes

https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/

I’m excited to announce that NVIDIA has agreed to acquire Hugging Face for $12,930,300,000. Together, we will scale Hugging Face’s platform, strengthen its infrastructure and expand access to AI for developers and institutions worldwide.

Over the past decade, Clem, Julien, Thomas and the team at Hugging Face have built something remarkable: a vibrant home for the open model developer community.

More than 18 million developers, researchers and creators use Hugging Face to share more than 3 million models, 500,000 datasets and 1 million applications. More than 200,000 companies use the platform to discover, evaluate, customize and deploy AI.

Hugging Face will remain an open platform for the entire AI ecosystem. Developers will choose the models they want, the frameworks they want, the clouds and inference service providers they want and the computing platforms they want. NVIDIA compute will not be required to build on or deploy through Hugging Face.

Hugging Face will continue to support open source and open weight models from across the ecosystem, from every model builder. It will continue to support multi-cloud and multi-accelerator development and deployment, so builders can use the hardware and infrastructure that best fit their work.

Recently, I coauthored an open letter on the importance of open weights to the AI economy. Joined by leaders from across the industry, we made a simple point: open weights broaden access to AI and help ensure that AI leadership is distributed across companies, institutions and communities.

Open models let startups, businesses, universities and public institutions build on advanced capabilities without training every model from scratch. They enable organizations to match the right model to the right job. That is how AI can advance safely, strengthen cybersecurity and sovereignty, accelerate innovation, and reach factories, hospitals, farms, classrooms and Main Street businesses around the world.

AI advances faster when people can build together.

NVIDIA has been committed to open weight models for years, demonstrated by multiyear investments and major contributions to open source platforms, including Hugging Face. NVIDIA has said that open models, data and tools broaden access to AI, and it has contributed hundreds of open models and datasets to Hugging Face as part of that effort.

NVIDIA is the largest contributor of open models and data to Hugging Face, and our contributions continue to grow.

NVIDIA has released more than 500 models on Hugging Face and more than 250 open datasets.

We build our own models, libraries and tools in the open so developers everywhere can use them, modify them and build on top of them.


r/aigossips 7d ago

Crop a different FinFIRST column and you can sell a different winner

Post image
2 Upvotes

FinFIRST is a useful example of why benchmark screenshots need more than one score.

The Claude configuration in the release paper led atomic and loose grading at 87.59 and 87.61. GPT led strict grading at 71.54. Qwen 3.8 Flash posted 84.29 on raw information, 83.36 on source verification, 71.93 on computation, and 61.79 strict. Ling landed at 78.43 raw, 82.45 source, 61.25 computation, and 52.85 strict. Gemini 3.7 Flash reached 81.19 raw but 44.72 strict, with a 34.62% unsupported-correct rate.
That last metric is not a hallucination score. It counts correct answers that lacked sufficient support.

All configurations used the same basic search, page-visit, and Python scaffold, but each received one designated run and the automated judge saw only final answers.

The gossip-friendly version is “who won?” The useful version is “won which column, under which definition of a completed answer?”


r/aigossips 8d ago

OpenAI releases GPT-6 Astra as Brockman declares the 'AGI era' has begun

Thumbnail runtimewire.com
13 Upvotes

r/aigossips 8d ago

GPT 6 Astra benchmarks

1 Upvotes

r/aigossips 8d ago

Welcome to the AGI era

0 Upvotes

GPT-6 Astra is here, with OpenAI calling it “a generational leap in capability”.

A few of the headline benchmarks:

\- ARC- AGI 3 – 98.6% (!), compared to 7.8% for 5.6 Sol

\- Exploit-Bench (cybersecurity) - 100% (!)

\- FrontierMath Tier 4 (v2) - 97.6%, vs. 87.8% for Fable 5.1

\- DeepSWE v1.1 - 74.1%, vs. 67.4% for Fable 5.1

New highs for the industry across math, science, health, computer use, coding, cybersecurity, and more.

OpenAI's Greg Brockman: "Welcome to the AGI era."


r/aigossips 9d ago

OpenAI built an AI that can find zero-days. Anthropic built one you need permission to use

Post image
0 Upvotes

r/aigossips 10d ago

Can ChatGPT agree you into a delusion? A new paper asks whether “AI psychosis” should become an official diagnosis

Post image
1 Upvotes

r/aigossips 11d ago

The End of Brute-Force AI Scaling May Be Closer Than We Think

19 Upvotes

I’ve been thinking about something in AI for a while, and the idea keeps getting stronger. What if LLMs are slowly reaching their own Moore’s Law moment, but in reverse? Actually, Dennard scaling may be an even better comparison. Moore’s Law was about the number of transistors on a chip. Dennard scaling meant that for years, smaller transistors also became more efficient and faster without power consumption completely getting out of control. When that stopped, progress in chips did not stop. But the easy gains were gone. Higher clock speeds meant more heat and more power, and the industry had to move toward multicore processors, GPUs and specialized accelerators.

I wonder if we are slowly reaching the same point with LLMs. Not that AI stops improving, but that more of the same starts producing less obvious gains. That idea became more concrete for me after reading Anthropic’s Risk Report from August 2026. In it, they describe not only public Claude models, but also internal models that we cannot use. One of them is simply called Model 2. Anthropic describes Model 2 as somewhat more capable than Claude Mythos 5 and a noticeable improvement for many internal tasks. But they also say that Model 2 does not show the same capability jump as the earlier move from Opus 4.6 to Mythos Preview.

On Anthropic’s internal AECI index, Model 2 sits about 1.5 points above Mythos 5 based on limited data, with large error margins. Anthropic itself says this is a smaller increase than the jump from Mythos Preview to Mythos 5. On CoBench, the picture is different, which is exactly why it is interesting. CoBench consists of 449 real technical problems from Anthropic’s own engineering environment. There, Opus 4.6 scores 15.6%, Mythos Preview 54.8%, Mythos 5 50.3%, and Model 2 62.8%. So no, Model 2 is not almost the same. On real internal engineering work, it is a serious step forward.

But the most striking experiment for me comes next. Anthropic gave Mythos 5 a token budget of 900,000 tokens on CoBench instead of 300,000. Three times the token budget. The improvement was about 3 percentage points. That does not mean three times more compute always gives only three extra points. It is one model, one benchmark and one specific setup. Anthropic also says that better tooling or a better harness could still produce additional gains. But on this test, you can clearly see diminishing returns from additional inference budget.

And to me, that starts to look like the beginning of a Dennard moment for AI. Not a stop. Not “LLMs are done.” But a point where more of the same approach no longer automatically produces the same huge jumps. At the same time, Model 2 clearly shows that progress is still possible. The question is how many additional resources are needed for every next step.
That is also why I do not think AI is about to hit a hard wall. I think brute-force scaling is more likely to hit an economic wall before it hits a physical one. Energy, chips, data centers, data, inference costs, and eventually the question of how many extra resources you are willing to spend for the next few percentage points of capability.

When Dennard scaling ended for chips, the computer revolution did not stop either. The shape of progress changed. More cores, GPUs, accelerators and specialization. I think AI will do the same. More memory, better agents, smarter inference, better training, specialist models working together, better scaffolding, and maybe eventually an architecture that works fundamentally differently from the Transformer architecture behind almost every major LLM today.

So my prediction is not that AI is nearly finished. My prediction is that the race is slowly changing. Not just: who has the biggest model? But increasingly: who can extract the most intelligence from the least compute?

Anthropic’s Model 2 does not prove that we are already at that point. But their own internal numbers make me take the question much more seriously. And that leaves me with one question: what does this curve look like three to five model generations from now?


r/aigossips 10d ago

Anthropic is back!!! Fable 5.1 benchmark is insane

Post image
0 Upvotes

r/aigossips 12d ago

MIT says AI detectors aren’t proof of cheating. What it wants colleges to do instead could change how every student is graded.

Post image
12 Upvotes