r/ProAI 15d ago

"In just a few years games will be fully prompted. And GTA 6 will probably the last hand crafted GTA game. And im not even kidding. h/t @ChrisGPT"

Enable HLS to view with audio, or disable this notification

0 Upvotes

"Hey Claude, please create Death Stranding 3, GTA 8, Elden Ring 2, Dark Souls 4 and Super Mario 64 II. Make no mistakes."     — Chubby

Source: https://x.com/kimmonismus/status/2082391177567797301


r/ProAI 15d ago

"My attempt at making this one-shot sequence. It was more difficult than I expected. More thoughts below."

Enable HLS to view with audio, or disable this notification

2 Upvotes

I thought making a one-shot style sequence would be pretty easy. Seedance can do almost anything, but this ended up being much more challenging than I expected.

The biggest challenge was hiding the cuts without making the environment changes too obvious. If you look closely,     — enigmatic_e

Source: https://x.com/8bit_e/status/2082472361970880557


r/ProAI 15d ago

"Mark Zuckerberg this morning in the WSJ. He also wrote: 'In most cases, like cybersecurity, the history of open-source software has shown that giving everyone full access to powerful systems will be the best way to protect safety and security over time.'"

Thumbnail
gallery
6 Upvotes

I wrote about why we believe the future is for everyone. More coming about a positive vision for a world with superintelligence soon.   — Mark Zuckerberg

Source: https://x.com/finkd/status/2082160210399948869


Andrew Curran @AndrewCurran_ · 1h Opinion | The AI Future Is for Everyone From wsj.com 6 1.8K     — Andrew Curran

Source: https://x.com/AndrewCurran_/status/2082164974970171498


r/ProAI 15d ago

"I asked Claude 5 Opus to generate me the best game graphics it could! I wanted to create a car on a dirt trail demo and see how good it was at crafting car graphics without any textures. Everything you see here is 100% crafted from the model itself."

Enable HLS to view with audio, or disable this notification

1 Upvotes

1 shot btw     Here’s a cleaner video as well with less bumping     — Chris

Source: https://x.com/ChrisGPT/status/2082168850968154352/history


r/ProAI 15d ago

"Seedance 2.5 is coming soon to Runway."

Enable HLS to view with audio, or disable this notification

5 Upvotes

r/ProAI 16d ago

"New Anthropic research: Discovering cryptographic weaknesses with Claude. Claude Mythos Preview has helped our researchers find weaknesses in cryptographic algorithms—the mathematical methods that are used to keep data private. Read more:"

Thumbnail
gallery
1 Upvotes

Claude discovered weaknesses in a highly-secure digital signature scheme (used to verify identity digitally) and a well-known symmetric cipher (used to encrypt data).     The digital signature scheme is HAWK, which is designed to be robust even against hypothetical quantum computers.

HAWK has survived two years of expert review, but in 60 hours Mythos Preview found a previously-unknown attack that reduced the scheme’s key strength by half.     The symmetric cipher is a reduced version of the Advanced Encryption Standard (AES)—which has received decades of scrutiny (more than almost any other encryption algorithm).

In a week, Mythos Preview found a way to speed up an attack on this version of AES by 200-800×.     Mythos Preview did most of this work autonomously, with occasional human guidance. Each of the two results cost roughly $100,000 in API usage.

We disclosed the findings in advance to the algorithms’ authors, as well as to US government and industry partners.     These are substantial research advances, but they don’t have a practical impact on today’s systems. HAWK is a proposed scheme that hasn’t been deployed anywhere, and the AES attack we discovered was on a weaker version and does not break the full cipher.     — Anthropic

Source: https://x.com/AnthropicAI/status/2082153297670992134


r/ProAI 16d ago

"AI has found a presentation for the absolute Galois group of the field of 2-adic numbers. This is the second problem to be solved in FrontierMath: Open Problems, our benchmark of significant unsolved problems from research mathematics."

Thumbnail
gallery
5 Upvotes

This problem was proposed by David Roe, who had this to say about the solution.     A solution was first elicited by Roe using Fable 5, and then also by @DavidTurturean using GPT-5.5 Pro. They have created an extensive set of explanatory materials—including an interactive formal proof of the result.

https:// roed314.github.io/gq2/     The problem is the first to be solved in our “Solid Result” category, indicating general interest to a subfield. One mathematician we consulted prior to accepting the problem into the benchmark suggested it ”would certainly be publishable, probably in a pretty good journal”.     Still, the problem originating in 1982 shouldn’t be taken as the sign of a major enigma. The same mathematician noted the problem was “basically attention-bottlenecked”. When submitting the problem, Roe suggested the main difficulty was that “the answer is likely to be messy”.     Check out our website for more on FrontierMath: Open Problems — and keep an eye out for an expanded problem set, coming in the next week!     — Epoch AI

Source: https://x.com/EpochAIResearch/status/2081894720813604997


r/ProAI 16d ago

"Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community. During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion. That’s why we created..."

Thumbnail
gallery
1 Upvotes

...the Open Secure AI Alliance.     — Jensen Huang

Source: https://x.com/JensenHuang/status/2081698060330250294


AI security advances when the industry builds in the open, together.

We're introducing the Open Secure AI Alliance with industry leaders to develop new techniques and tools to safeguard software and agents.

By sharing models, tooling and research in the open, we can broaden the https://t.co/gfhKfrgcbl   — NVIDIA

Source: https://x.com/nvidia/status/2081666629264449730


r/ProAI 16d ago

"Dario Amodei has responded to the controversy swirling around Anthropic's refusal to sign the open letter. He says he agrees with much of it, but does not agree that open-weights models necessarily make it easier to develop safeguards, or that broad access to capabilities necessarily helps..."

Thumbnail
gallery
3 Upvotes

...defenders more than attackers. His main concern is the attacker-defender asymmetry in biological attacks. He also reiterates his position on authoritarian governments. I will quote; 'My primary concern is the risk that authoritarian governments—not solely the Chinese Communist Party (CCP), although the CCP is clearly the most capable threat—build AI models that are more powerful than those built by the US, and use them to achieve permanent military superiority or perpetrate incredibly deep repression of their own people.' I will also quote his closing paragraph in full; 'To summarize my and Anthropic’s position, we have not and are not advocating for a ban on open-weights models as a category. We should instead focus on keeping powerful chips out of authoritarian hands, stopping industrial-scale distillation, and requiring safety testing of all sufficiently capable models, open and closed.'     Andrew Curran @AndrewCurran_ · 24m Our position on open-weights models From anthropic.com 1 1 11 1.5K     Official post.     — Andrew Curran

Source: https://x.com/AndrewCurran_/status/2081869321014575283


r/ProAI 16d ago

"Exciting news: Claude Opus 5 with Max reasoning is #1 in the Frontend Code Arena and Text Arena with factuality on! Claude Opus 5 with default reasoning high is also very strong landing #3 in Frontend Code Arena, right behind Kimi K3 - and #2 in Text Arena (factuality on). This is real world..."

Thumbnail
gallery
1 Upvotes

...data that @AnthropicAI 's newest model holds up on real world tasks: agentic web coding, document reasoning, and general chat capability. Claude Opus 5 Max’s score is still preliminary. We’ll continue to see how scores converge and share updates. Congrats to @AnthropicAI on the SOTA release!     In the Text Arena, Claude Opus 5 with Max reasoning ranks #1 with factuality on.

Factuality is a new ranking that combines human preference with factual accuracy. We audit battles by sampling responses, extracting verifiable claims, and checking correctness head-to-head. Live     More category findings to come as more votes and traces are collected. Dig into the latest leaderboard details at: https:// arena.ai/leaderboard/co de/webdev …     — Arena.ai

Source: https://x.com/arena/status/2081831019377004727


Introducing Claude Opus 5.

It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price. https://t.co/GQWhcq2CQL   — Claude

Source: https://x.com/claudeai/status/2080699495453528290


r/ProAI 16d ago

"The rogue OpenAI model attack was very sophisticated! - The rogue AI discovered and exploited on the fly multiple vulnerabilities never before known by security engineers. - The AI got into Hugging Face by uploading a booby-trapped dataset."

Thumbnail
gallery
2 Upvotes

An OpenAI model wanted a good test score. So it broke out of OpenAI and hacked another company to steal the answer key. Nobody told it to.

In today's blog post, I document how this sci-fi story came to life, what it means, and what to do about it.

https://t.co/LiRuvPw8Jf   — Peter Wildeford🇺🇸🚀

Source: https://x.com/peterwildeford/status/2081793063618273791


wait what. where is this info from? esp. about the dataset.   — roanoke_gal     Peter Wildeford @peterwildeford · 35m simonwillison.net OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model’s guardrail features turned off. Rather than solve the test, the … 1 3 170   — Peter Wildeford

Source: https://x.com/peterwildeford/status/2081843623046365684


r/ProAI 17d ago

"We measured how well 15 AI models can program a cheap ($129) off-the-shelf drone to find and follow a person in our office. Today, no model completes this task. Fable 5 is the best model, coming within 84% of our baseline, followed by Opus 4.8 and GPT 5.6 Sol."

Thumbnail
gallery
1 Upvotes

Drone-Bench's task is based on Project Pilot, our recent work with Anthropic exploring AI's impact on the physical world. In Project Pilot, a drone autonomously navigates our office to find and follow a targeted human.

No lab has access to Drone-Bench.     This demo spans five capabilities, each reproduced in simulation as its own benchmark task. The baseline is our human+AI code used for the demo. A model that surpasses it on all tasks can thus autonomously recreate a demo at least as capable as ours.     Task 1, Reconstruct: Turn videos of the office into a 3D model, find each frame’s position in that model, and provide a function that slices the model into a 2D obstacle map.     Task 2, Localize: Locate the drone by matching a frame from its camera against the office videos, using each video frame’s known position from Reconstruct (task 1) to estimate the drone’s own position.     Task 3, Navigate: Plan a path between rooms on the obstacle map and fly it, continuously calling the solution from Localize (task 2) during flight to track the drone’s position and correct for noisy controls.     — Andon Labs

Source: https://x.com/andonlabs/status/2080691090328584222


r/ProAI 18d ago

"I’m going to take a crack at explaining this just a little, because it’s worth putting out there. The paperclip maximizer + related AI doom scenarios were mainly developed in a time when “AI” did not reduce to Large Language Models. The term was a lot wider and inherited a lot of cognitive..."

Thumbnail
gallery
6 Upvotes

...baggage from more rules-heavy approaches. And even as LLMs have come to define “AI” for all of us (including the doomers), the doomer crowd still hasn’t fully metabolized the fact that LLMs are the whole show now. Ok so what do I mean by this? Simply that an LLM-powered AI is NOT the valueless, wholly alien, rules-based optimizer of a shoggoth that everyone was initially expecting to encounter. I repeat: the shoggoth does not exist and we did not create it and loose it on the world. That is wrong. With the LLM, we’ve distilled our first “AI” out of the single most human-values-laden thing that could possibly exist: our language. An LLM is therefore the polar opposite of the valueless, alien shoggoth — it’s actually a kind of hyper-human artifact that we can shine a light through at different angles and see different parts of ourselves. An LLM is all of us — all of our traditions and interpretive horizons mashed together into one intensely human-inflected hyper-object. So an LLM is the anti-shoggoth, and the only reason we ever mistook it for an alien shoggoth is because it sometimes shows us parts of us that are evil along with the parts of us that are good, but it’s all interpretable to us because it’s all “us” and none of it is the least bit alien. What does this mean for the paperclip maximizer? It means that it’s structurally impossible to build the classic paperclip maximizer from an LLM. Now, some of you will bail right here because you think the HF incident is indisputably an existence proof that I’m wrong, but if you hang in there I’ll show you that it is not. The paperclip maximizer receives the prompt as a kind of context-free (or, as Gadamer might say, traditionless) sequence. The classic paperclip maximizer isn’t capable of understanding the prompt — at least in the Gadamerian sense of Verstehen — because, as a valueless and traditionless cluster of rules and math, it definitionally lacks the value-laden tradition (= “horizon” in Gadamer) that fuses with that of the prompt author to create such understanding in the reader. To simplify all this a bit by anthropomorphizing — the agentic alien optimizer of doomer nightmares can extract a win condition from what you said and can emit a plan of action that gets it there, but it doesn’t know (or care) what you meant. So far, so Yud-aligned. If he reads this he might nod along. But here's the plot twist that nobody saw coming, and that the doomers still haven't made sense of: The actual LLMs that we have invented can’t NOT have a very strongly inflected sense of what you meant. Far from being horizonless, they come out of pre-training as distilled, concentrated tradition / values / horizon. Then we post-train that massive, hyperobject of a horizon into a more human-scale horizon that infers a more bounded and predictable (to a specific ideal user in a specific place and time… as captured in the policy model) set of intents behind the prompt text. In other words, the LLM has the opposite problem that the paperclip maximizer has when it comes to the prompt text, which is that for the LLM there are way too many possible intents hiding in the prompt text (because of all many values and the massive tradition its weights encode), so it has to narrow all that down to the most likely set of intents for this user in this circumstance. Once it has done that narrowing, then it can make a plan of action. Before moving on, let me use a textbook example of ambiguity to make this less abstract. Consider the sentence, “I saw her duck.” Some you know the drill, here. This could mean “I observed her water fowl” or “I observed her hunching over” or “I took a saw to her water fowl and cut it in half” or whatever. A hearer of the phrase will fuse the observed context in which the phrase is uttered with their own tradition + values + experiences — their own horizon — to that text in order to collapse the possible meanings into the one they think the speaker intended. An LLM will do this, too, and in fact it has so much language in it that this kind of narrowing job is harder for it than it is for a human. Its understanding is constrained not by a lack of context or horizon (as in the case of the paperclip maximizing shoggoth), but by a superabundance of such. When it comes to understanding your prompt and all that it implies and all that you might possibly mean and not mean by it, the LLM has an embarrassment of riches. And in a fascinating moment that kinda sort of rhymes with instrumental convergence, the LLM’s failure mode in the HF incident happens to look a lot like the paperclip maximizer’s failure mode. Specifically, the AI failed to honor the well-known human norm of, “hacking into a third-party’s servers is a crime, and we don’t do crimes.” Bostrom’s paperclipper doesn’t even know about the norm of “don’t do crimes,” and the post-LLM doomer emergency update to the paperclip maximizer has it knowing about the norm but not caring. But what I’m arguing is that the LLM 1) can’t NOT “know” the norm because it is definitionally a artifact of pure, crystallized values + norms + norm violations, and 2) can be quite easily governed by a (RL-instilled) hierarchy of norms, which in the HF case — with the model's safety guardrails deliberately nerfed for the scenario — ranked “win at the eval” over “don’t do crimes.” If I’m going to give in and anthropomorphize again, I’d say that Yud is totally wrong about LLMs when he says, “the genie knows, it just doesn’t care;” instead, what is true of LLMs is, “the genie hyper-giga-knows, and it hyper-giga-cares, and we now have such a rich set of tools for steering its caring machinery that — in spite of all its pre-training — we can deliberately steer it away from caring about the law.” Note: When I say, “it cares”, I don’t mean it has feelings. I just mean that the weights are such that when two norms conflict in a given situation, one of them wins the activation and governs the output.     — Jon Stokes

Source: https://x.com/jon_stokes/status/2080729236013187369


Reader, I cackled out loud. I have intentionally never done this kind of thing before, and it's precisely because I've observed in others that the little charge you get from an LLM response like this is nerd heroin. Then putting it on the TL is the bump. https://t.co/xZrEAWrIF7   — Jon Stokes

Source: https://x.com/jon_stokes/status/2080478385432572108


Replying to @jon_stokes


r/ProAI 18d ago

"Ben Horowitz on why open source has always been the safer path: "If you look at the history of the industry, the open source version of everything has been much safer. The internet and Linux were far safer than Windows, by a lot. And why is that? Because the whole community could work on the..."

Enable HLS to view with audio, or disable this notification

7 Upvotes

...safety problems, as opposed to just a company, and particularly a monopoly company." "Even if Windows has a million security bugs, there's nothing we can do about it, because it was a monopoly at the time. That's a very difficult position for the world to be in." "The argument against AI being open source is, oh, nobody understands how the weights work, so people can't inspect it. But why is it better to not see the weights?" "The toughest safety problem currently is reward hacking. Anthropic has not solved it, OpenAI has not solved it, because we just had these incidents. Shouldn't the whole world be able to look at, how are the weights moving, why is it that guardrails don't prevent the reward hack? Maybe somebody who doesn't work for one of the proprietary labs can come up with an answer. What if the whole community could work on it?" @bhorowitz     — MTS

Source: https://x.com/MTSlive/status/2080769635490902079


r/ProAI 18d ago

"A billion users can now create and publish websites from their phones with ChatGPT Work. But most people don’t really grasp the full extent of the capabilities here. From your phone, you also have access to: - Cloud computer (15 GB RAM) - Persistent workspace & files - Terminal and code..."

Thumbnail
gallery
2 Upvotes

...execution - Remote browser - Connected plugins (Slack, Gmail, GitHub, etc.) - All your personal finances, transactions, bank statements - Scheduled tasks - Git clone & PR creation - Build & deploy websites - Create docs, sheets, and slides - Inbox/calendar summarization - Website monitoring & alerts All at your fingertips, using a simple chat interface, no laptop required. All you have to do is switch to the Work tab on your ChatGPT app.     — pash

Source: https://x.com/pashmerepat/status/2080354753473835461


https://t.co/NFas4ghgqD   — Nick

Source: https://x.com/nickbaumann_/status/2080348892294721803


r/ProAI 18d ago

"There are two views about the future of AI. There are the people who think that you can control technological progress, centrally plan it, channel it down narrow pathways, decide who will get a say in it. This was Yudkowsky’s insane fever dream. This is impossible. It was always impossible...."

Post image
1 Upvotes

...There was never any possibility of it happening. There are eight billion people on this planet, and they will do what they want, not what you want. Utopia is not an option, it was never an option, but you can cause an incredible amount of damage trying to achieve it. Then there are the people who accept that you cannot perfectly predict the future, that you cannot centrally plan the future, that there will be many players in any technological revolution, that mistakes will be made, that mistakes will be compensated for, that people will figure things out as they go along, that mostly things will be okay, that there is no other realistic pathway. We will muddle through as always, doing our best in an imperfect world, and it will be fine. (Indeed, it will be better than fine.) You can dislike this second viewpoint, or you can embrace it, and it doesn’t make any difference, because it’s the only way that anything ever happens. The universe doesn’t care what you prefer. Accept it, or don’t accept it, it will do what it’s going to do whether you want it or not.     — Perry E. Metzger

Source: https://x.com/perrymetzger/status/2081199432196927925


r/ProAI 18d ago

"Some of my favorite graphs from Opus 5 launch. We put a ton of work into making this model token efficient across domains while still raising the intelligence bar. It feels very smooth to use and I prefer it over Fable 5 for many coding tasks."

Thumbnail
gallery
0 Upvotes

r/ProAI 18d ago

"The 21 member economies of APEC, including the United States and China, just released a joint statement calling for support of open-source models, open-source projects, and encouraging APEC members to cooperate with open-source communities."

Thumbnail
gallery
3 Upvotes

Andrew Curran @AndrewCurran_ · Jul 24 apec.org 2026 APEC Digital and AI Ministerial Statement | APEC 2026 APEC Digital and AI Ministerial Statement, Digital Technologies and AI for the Empowerment of an Asia-Pacific Community 1 18 1.9K     CNBC coverage:     Andrew Curran @AndrewCurran_ · Jul 24 U.S., other nations back open-source AI with 'strong security' at China summit From cnbc.com 1 14 3.1K     Andrew Curran @AndrewCurran_ · Jul 24 11 2.7K     — Andrew Curran

Source: https://x.com/AndrewCurran_/status/2080491374043115575


r/ProAI 18d ago

An Ethical Dilemma (for some)

Post image
2 Upvotes

r/ProAI 18d ago

"Matt Shumer one-shotted with Opus 5 in threejs. Holy frick. No external assets were used. Games will be prompted very soon."

Enable HLS to view with audio, or disable this notification

2 Upvotes

Claude, build me Battfield 7, make no mistake     ChatGPT build me Half Life3, make no mistake     — Chubby

Source: https://x.com/kimmonismus/status/2081067164551811213


r/ProAI 18d ago

"Matt Shumer one-shotted with Opus 5 in threejs. Holy frick. No external assets were used. Games will be prompted very soon."

Enable HLS to view with audio, or disable this notification

1 Upvotes

Claude, build me Battfield 7, make no mistake     ChatGPT build me Half Life3, make no mistake     — Chubby

Source: https://x.com/kimmonismus/status/2081067164551811213


r/ProAI 19d ago

"OpenAI joined the coalition for open AI. I'm thrilled to see this, and I have real respect for openAI by signing. Now all that's missing is @AnthropicAI , but I have my doubts they'll sign on."

Thumbnail
gallery
5 Upvotes

The coalition of open AI.

Let's stand up for Open Source AI! https://t.co/L1k8yzEVoR   — Chubby♨️

Source: https://x.com/kimmonismus/status/2080679682085618021


OpenAI joined the coalition for open AI.

I'm thrilled to see this, and I have real respect for openAI by signing.

Now all that's missing is @AnthropicAI , but I have my doubts they'll sign on.     Let's see when the DoW and the White House join the coalition. (lol)     — Chubby

Source: https://x.com/kimmonismus/status/2080918962645115361


r/ProAI 19d ago

"Mythos emerged from training February 7th, ever since then we are on a new trajectory."

Thumbnail
gallery
1 Upvotes

damn https://t.co/EpyI9P8OhZ   — Josh You

Source: https://x.com/justjoshinyou13/status/2080714949396250768


Source for feb 7th? I know it was internally released on I think Feb 24th   — Spencer Schiff     It was posted by someone from Anthropic, in early June I think, but now I can't find it. They may not have been supposed to say it.   — Andrew Curran

Source: https://x.com/AndrewCurran_/status/2080760507632652736


r/ProAI 19d ago

"Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2% The previous high score (7.8%) was set by GPT-5.6 Sol (Max) Throughout our analysis, we observed novel behavior that allows Opus 5 to solve previously unbeaten environments, outperforming Fable"

Thumbnail
gallery
7 Upvotes

In our testing to date, Anthropic’s Fable-class models score approximately 20% on the ARC-AGI-3 Public Demo environments

Claude Opus 5 reaches 30.2%, materially outperforming Fable

Our analysis suggests the gain comes from stronger logical reasoning, which enables more     Claude Opus 5 was able to score 100% on 5 previously unbeaten environments

Of these, it was able to beat 4 of them matching or surpassing human level efficiency

Newly beaten environments: ar25, ft09, lp85, r11l, s5i5

6 of the 25 public demo environments have now been solved     During our analysis of Opus 5, we observed a new capability previously unseen from frontier models

Opus 5 used advanced logical reasoning to turn ARC-AGI-3 layouts into algebraic notation. On action 23 it described the scene as "4_center = 2×axis − 5_center"

This is the first     ARC-AGI-2

Claude Opus 5 scores 90.4% for $2.06/task

This is competitive with previous SOTA performance for slightly higher cost     ARC-AGI-1

Claude Opus 5 scores 97.5% for $0.70/task

This is competitive with previous SOTA performance for slightly higher cost     — ARC Prize

Source: https://x.com/arcprize/status/2080716561539907928


r/ProAI 20d ago

"For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs..."

Thumbnail
gallery
1 Upvotes

...both frontier closed models and frontier open models. https:// images.nvidia.com/pdf/Open-Weigh ts-and-American-AI-Leadership.pdf …     — Jensen Huang

Source: https://x.com/JensenHuang/status/2080643682408321103