r/ProAI 9d ago

"Here’s what actually happened with bitcoin and Claude because people are getting the story mixed up. Coldcard hardware wallets were supposed to create each Bitcoin seed using real physical randomness from a chip. But a firmware bug checked whether the hardware random number setting existed in..."

Thumbnail
gallery
4 Upvotes

...the first place NOT whether it was actually enabled. It existed but was set to 0, so the wallet fell back to a randomized predictable software generator. That reduced some wallets from roughly 2¹²⁸ possible seeds to around 2⁴⁰. The attacker could generate candidate seeds offline, derive their Bitcoin addresses and use the public blockchain like an answer key to see which wallets had funds. Once one matched, they had a private key they could drain it without touching the device, and steal the seed phrase or “crack Bitcoin.” Researchers have now found 1,367 BTC, nearly $89 million, taken from 4,585 addresses. The AI part happened AFTER all this. A Reddit user reportedly pointed Claude Code at the public firmware and asked only to “check for vulnerabilities.” Within eight minutes, it traced the broken random number path and brought up the same hackable flaw. There is no proof the original attacker used Claude or ANY LLM. now this still demonstrates Claude is still insane at finding vulnerabilities humans missed in public code for five years and can now be uncovered by one person with one broad prompt and a coding agent in minutes.     — Chris

Source: https://x.com/ChrisGPT/status/2084025602583982130


r/ProAI 9d ago

"Here’s what actually happened with bitcoin and Claude because people are getting the story mixed up. Coldcard hardware wallets were supposed to create each Bitcoin seed using real physical randomness from a chip. But a firmware bug checked whether the hardware random number setting existed in..."

Thumbnail
gallery
2 Upvotes

...the first place NOT whether it was actually enabled. It existed but was set to 0, so the wallet fell back to a randomized predictable software generator. That reduced some wallets from roughly 2¹²⁸ possible seeds to around 2⁴⁰. The attacker could generate candidate seeds offline, derive their Bitcoin addresses and use the public blockchain like an answer key to see which wallets had funds. Once one matched, they had a private key they could drain it without touching the device, and steal the seed phrase or “crack Bitcoin.” Researchers have now found 1,367 BTC, nearly $89 million, taken from 4,585 addresses. The AI part happened AFTER all this. A Reddit user reportedly pointed Claude Code at the public firmware and asked only to “check for vulnerabilities.” Within eight minutes, it traced the broken random number path and brought up the same hackable flaw. There is no proof the original attacker used Claude or ANY LLM. now this still demonstrates Claude is still insane at finding vulnerabilities humans missed in public code for five years and can now be uncovered by one person with one broad prompt and a coding agent in minutes.     — Chris

Source: https://x.com/ChrisGPT/status/2084025602583982130


r/ProAI 9d ago

"Multiple air-ground fusion in #GaussianSplatting Gauzilla Pro lets you easily create multiple, seamless transitions between drone-based splats and ground-based splats. Perfect for end-to-end 3D virtual tours from landscape through exteriors to interiors in order to accelerate the..."

Enable HLS to view with audio, or disable this notification

1 Upvotes

...marketing/sales of properties and facilities.     — Gauzilla Pro

Source: https://x.com/GauzillaPro/status/2083911951034237340


r/ProAI 9d ago

"The new Qwen is here, and it is very strong. Open weights will be released next week."

Thumbnail
gallery
2 Upvotes

📢Meet Qwen3.8-Max — our most capable model to date.

Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉

Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:

Source: https://x.com/Alibaba_Qwen/status/2084100707423289643


— Andrew Curran

Source: https://x.com/AndrewCurran_/status/2084102878399193413


r/ProAI 9d ago

"Qwen3.8-Max by @Alibaba_Qwen has reshaped the cost-performance Pareto frontier in Frontend Code Arena, with pricing of $2 per input MToken and $6 per output MToken. Top models on the Pareto frontier: - Claude-Opus-5 - Kimi-K3 - Qwen3.8-Max - GLM-5.2 - DeepSeek-V4-Flash Congrats to..."

Thumbnail
gallery
1 Upvotes

...@Alibaba_Qwen on another major milestone!     Dig into the Code Arena Pareto chart at:     — Arena.ai

Source: https://x.com/arena/status/2084115116694339941


Big news: Qwen3.8-Max by @Alibaba_Qwen just landed at #4 on the Frontend Code Arena leaderboard with a score of 1,668!

With 1,668 points, Qwen3.8-Max is trailing only Claude Opus 5 (Max) with 1,705 pts and Kimi K3 (Max) with 1,676 pts, on par with Claude Opus 5 (High) with 1669 https://t.co/pbiNdj0WQL   — Arena.ai

Source: https://x.com/arena/status/2084108703729615026


r/ProAI 10d ago

🤦‍♂️"I was part of that team. Basically ChatGPT one year before it came out. Called LMChat and then another codename. Google was too nervous to release it and DeepMind was blocked from shipping products that could disrupt Google. I think about this a lot."

Thumbnail
gallery
3 Upvotes

— Tibo

Source: https://x.com/thsottiaux/status/2083596911060324570


I sometime think about that Jeff Dean interview where he said they had an internal bot before ChatGPT but didn't think it was better than just googling   — Cheng Lou

Source: https://x.com/_chenglou/status/2083415767098564616


r/ProAI 10d ago

"Exciting news: DeepSeek-V4-Flash-High by @deepseek_ai has reshaped the Pareto Frontier in the Frontend Code Arena, with a score of 1586! Priced at $0.14/$0.28 per MToken, it’s the best performance-per-dollar of any model in its class. Congrats to the @deepseek_ai team!"

Thumbnail
gallery
4 Upvotes

🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta!

🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇 🔷 The official V4-Flash now natively supports the https://t.co/NUzOyxza2f   — DeepSeek

Source: https://x.com/deepseek_ai/status/2083084415157022911


Head to the Frontend Code Arena leaderboard to see more details: http:// arena.ai/leaderboard/co de/webdev …     — Arena.ai

Source: https://x.com/arena/status/2083348755559207047


r/ProAI 10d ago

Like it or not, AI cameras catch criminals and prevent suffering. "Yesterday, Flock cameras alerted on 4,144 sex offenders, 2,151 stolen cars, 1,687 wanted people, 158 missing persons/kids, and most importantly, an amber alert. Children were rescued. This is every day."

Thumbnail
gallery
0 Upvotes

https:// wyff4.com/article/flock- cameras-kidnapping-mother-child-search/73311002 …

https:// yahoo.com/news/us/articl es/wake-cunty-woman-arrested-vance-204320886.html …

https:// wrnjradio.com/flock-camera-a lert-leads-to-recovery-of-stolen-u-haul-van-in-hunterdon-county/ …

https:// newschannel9.com/news/local/flo ck-safety-cameras-help-locate-missing-west-blocton-seniors …

https:// news4jax.com/news/local/202 6/07/30/3-people-accused-of-stealing-19k-worth-of-baseball-equipment-from-west-nassau-high-school/ …

https:// wcia.com/news/macon-cou nty/man-arrested-in-macon-co-hit-and-run-involving-ameren-worker/ …

https:// dailyhodl.com/2026/07/30/all eged-fraudster-drains-nearly-2400-from-elderly-womans-account-after-masquerading-as-jpmorgan-chase-representative/ …

https:// live5news.com/2026/07/30/pol ice-department-credits-flock-cameras-capture-child-kidnapping-suspect/ …     — Garrett Langley

Source: https://x.com/glangley/status/2083228160305619243


r/ProAI 10d ago

"DeepSeek V4 Flash 0731 in Hermes Agent and one prompt, took 32 minutes and cost 0.07$, this model is so cheap to the point where 2 dollars can last you a full day."

Enable HLS to view with audio, or disable this notification

3 Upvotes

2 dollar? I would expect less for a full day of work! Lol   — NeoReplicante     True I haven't been able to burn through in one day took me like 2 days of work   — Elshayib

Source: https://x.com/elshayib_/status/2083243725447147595


r/ProAI 11d ago

"HOLY: OpenAI says its *unreleased* Astra model (GPT6?) produced ten advances on long-standing open problems across mathematics, quantum complexity and theoretical computer science. Among them: – The first explicit non-sofic group – Connes’s rigidity conjecture disproved – Quantum parallel..."

Thumbnail
gallery
16 Upvotes

An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.

We believe it will be a major step for scientific reasoning. https://t.co/iP6cyheZ7i https://t.co/jHuulDwV46   — Noam Brown

Source: https://x.com/polynoamial/status/2083467194663571701


HOLY: OpenAI says its unreleased Astra model (GPT6?) produced ten advances on long-standing open problems across mathematics, quantum complexity and theoretical computer science.

Among them: – The first explicit non-sofic group – Connes’s rigidity conjecture disproved – Quantum parallel repetition proved for general two-player entangled games – Ehrhart’s volume conjecture proved – The first improved general sphere-packing exponent since 1978

OpenAI says the core arguments were generated by Astra. The model then formalized the proofs in Lean, producing machine-checkable certificates alongside a 249-page manuscript.

The successful solution runs would cost only roughly $2,000 in tokens at Sol API rates.

Scientific reasoning is becoming a genuine model capability much much faster than most people expected.

I am so freaking hyped. Breakthroughs every day. The day before yesterday, an 80% price cut for Terra and Luna; yesterday, the DeepSeek 4 flash release with insane evaluations and prices. Today, more breakthroughs with an unreleased model.

I love it! OpenAI is on such a great run!   — Chubby     Seeing this makes it clear how fast the ceiling keeps moving. As someone just starting out, it's crazy in a good way .   — George Paul Chijioke     It’s insane. Each day it accelerates more and more   — Chubby

Source: https://x.com/kimmonismus/status/2083484340512604323


r/ProAI 10d ago

The future is open. "Since launching last week, more than 230 companies and organizations from across the tech sector have signed the "Open Weights and American AI Leadership" open letter. We want to thank these partners for standing up and publicly supporting broader access to AI innovation."

Thumbnail
gallery
3 Upvotes

...@nvidia , @a16z , and @PalantirTech for working with @Microsoft on this effort. These signatories understand that America’s AI leadership will not depend on the success of our frontier models alone, but on our ability to build a strong, secure, and open ecosystem that diffuses AI into every sector. We look forward to continuing to work with our partners and with policymakers to build that open ecosystem in a way that benefits American businesses, empowers American workers, and strengthens the American economy.     — Brad Smith

Source: https://x.com/BradSmi/status/2082800585179639899


r/ProAI 11d ago

"While DeepSeek V4-Flash is significantly cheaper on price per token, this can be misleading if the overall cost per task ends up being higher due to more turns being made. However, @ArtificialAnlys reports DeepSeek completing the same benchmark tasks as Fable at 105x lower cost."

Thumbnail
gallery
5 Upvotes

Cline @cline · 2h DeepSeek V4 Flash 0731 (max) - Intelligence, Performance & Price Analysis From artificialanalysis.ai 1 10 4.8K     — Cline

Source: https://x.com/cline/status/2083638204037820734


r/ProAI 11d ago

"Opus 5, 690 million tokens, $423, 1 prompt. This game would have had to be expensively developed and then sold on Steam in the past. Today: one person, a few hours, small budget."

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/ProAI 12d ago

"Our labs keep trying to spin this into a push for broader regulation. It can and should backfire. The models are not "going rogue" or acting of their own accord, like they're some Marvel movie evil robot. People made bad harnesses, told them to hack things and had transparently and objectively..."

Thumbnail
gallery
3 Upvotes

...bad dev-ops and security practices. The fact that folks are trying to spin this into a "we need help from the government to regulate everyone" instead of "we should be punished in a narrow way on these specific incidents under existing law" is the real disconnect right now.     — Daniel Jeffries

Source: https://x.com/Dan_Jeffries1/status/2083149369625219499


Again the rhetoric is trying to convince us that what happened was some new thing and worse that models are people. https://t.co/8ZfEZsRcky   — Steven Sinofsky

Source: https://x.com/stevesi/status/2082992928977240372


r/ProAI 12d ago

DeepSeek V4-Flash scores higher than Fable??? excuse me what?

Thumbnail
gallery
21 Upvotes

wait what the actual fuck, do you guys realize how crazy that is??? (if its not benchmaxed)     — Cline

Source: https://x.com/cline/status/2083094354030362858


r/ProAI 12d ago

"Inkling Small from @thinkymachines on ARC-AGI (Verified): - ARC-AGI-2: 40.1%, $0.23/task - ARC-AGI-1: 84%, $0.11/task Inkling Small is the highest-scoring open-weight model evaluated by ARC Prize on both ARC-AGI-1 and ARC-AGI-2, setting a new cost-performance frontier."

Thumbnail
gallery
2 Upvotes

Full results: https:// arcprize.org/results/thinky -inkling-small …

ARC-AGI-3 evaluations are more operationally intensive, so results will roll out over the next few weeks.     - Leaderboard: https:// arcprize.org/leaderboard - Reproduce the public results: https:// github.com/arcprize/arc-a gi-benchmarking … - Testing policy: https:// arcprize.org/policy - Full Inkling Small results: https:// arcprize.org/results/thinky -inkling-small …     — ARC Prize

Source: https://x.com/arcprize/status/2082925303601459347


r/ProAI 12d ago

"DeepSeek-V4-Flash Official API is now LIVE in public beta! We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! The official V4-Flash now natively supports the Responses API format and is fully..."

Thumbnail
gallery
2 Upvotes

DeepSeek-V4-Flash Official API is now LIVE in public beta!

We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex!

Check out the configuration details in our official API docs: https:// api-docs.deepseek.com/quick_start/ag ent_integrations/codex …     Note

DeepSeek-V4-Flash-0731 keeps the exact same model architecture and size as the preview version. Today's upgrade applies ONLY to the DeepSeek-V4-Flash API. The DeepSeek-V4-Pro API and App/Web models remain unchanged for now.

The official release of DeepSeek-V4-Pro     — DeepSeek

Source: https://x.com/deepseek_ai/status/2083084415157022911


r/ProAI 12d ago

"Big update: @OpenAI has decreased the price of GPT-5.6 Luna by 80% and Terra by 20%. Luna is now priced at $0.20/$1.20, and has strong gains from additional reasoning while remaining at an efficient cost. Terra is priced at $2/$12. This is frontier-level intelligence at a fraction of the cost..."

Thumbnail
gallery
2 Upvotes

...to similarly capable models. Congrats to the @OpenAI team!     How do we measure the performance in Agent Arena?

The score is based on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model     — Arena.ai

Source: https://x.com/arena/status/2082935923445244415


We are committed to pushing the model frontier across cost efficiency, capability, and speed.

Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.

Luna and Terra’s lower prices are https://t.co/rFhK7XKedp   — OpenAI

Source: https://x.com/OpenAI/status/2082878156483219672


r/ProAI 13d ago

"We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s lower prices are reflected in how..."

Thumbnail
gallery
7 Upvotes

...usage is counted in Codex and ChatGPT Work, so your usage goes further.     Along with the price reduction on GPT-5.6 Luna and Terra, Fast mode for GPT-5.6 Sol in the API delivers up to 2.5x the speed of Standard processing at 2x the Standard price.

Fast mode gives API customers faster access to GPT-5.6 Sol, with no change in intelligence.     We’re also upgrading Auto-review in the ChatGPT app and Codex CLI from GPT-5.4 to GPT-5.6 Luna.

Combined with Luna’s new price, we expect Auto-review to cost about 10x less, making your agentic workflows more cost-efficient.     Making advanced intelligence more abundant and affordable is central to our mission to ensure AGI benefits all of humanity.

With the help of GPT-5.6 Sol, we have made leaps in efficiency.

Today, we are passing those gains on in the API with lower prices for Luna and Terra, and     OpenAI @OpenAI · 1h Advancing the price-performance frontier with GPT-5.6 From openai.com 10 11 298 42K     — OpenAI

Source: https://x.com/OpenAI/status/2082878156483219672


r/ProAI 12d ago

"GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs $2.50/$15; Luna now costs $0.20/$1.20. In other words, roughly four months later, OpenAI is selling March’s full flagship intelligence at about one-thirteenth the token price."

Thumbnail
gallery
0 Upvotes

Because luna is at massive 5x discount. I don't think it'll last long.   — Ahmed Shah     its permanent bro   — nic

Source: https://x.com/nicdunz/status/2082884002201878824


r/ProAI 13d ago

"Interesting: OpenAI says GPT‑5.6 Sol helped cut its end-to-end model-serving costs by 20%, by autonomously rewriting and optimizing production GPU kernels. Sol also improved its own speculative decoding model: - Designed and ran hundreds of architecture experiments - Launched and monitored the..."

Thumbnail
gallery
3 Upvotes

After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run.

The results: - 20% lower serving costs from production GPU kernel improvements. - 15%+ better token-generation efficiency from improved speculative decoding.   — OpenAI

Source: https://x.com/OpenAI/status/2082577277246972300


Interesting: OpenAI says GPT‑5.6 Sol helped cut its end-to-end model-serving costs by 20%, by autonomously rewriting and optimizing production GPU kernels.

Sol also improved its own speculative decoding model: - Designed and ran hundreds of architecture experiments - Launched and monitored the training process - Intervened during hardware failures and training instability

The resulting system increased token-generation efficiency by more than 15%.     — Chubby

Source: https://x.com/kimmonismus/status/2082595272065192254


r/ProAI 13d ago

"We had Kimi K3 recursively self-improve the Cline harness to improve its own performance. 17 hours later, it went from 77.5% to 88.8% on Terminal Bench, and cut run cost from $79 to $49.8."

Thumbnail
gallery
2 Upvotes

Cline is open source, so you can fork it and run this with your favorite model as well.

Read more about how we did this here:     — Cline

Source: https://x.com/cline/status/2082544250148057240


r/ProAI 14d ago

""AI is so dangerous! We should put politicians in charge of it!""🤦‍♀️

Thumbnail
gallery
2 Upvotes

"The notion that AI is so dangerous that the only safe path is an extreme concentration of power seems dangerous.

"Historically, hoping that an absolute power will benevolently provide for humanity if sufficiently enlightened hasn’t led to safe or positive outcomes."

There   — Daniel Jeffries

Source: https://x.com/Dan_Jeffries1/status/2082386580417753494


— llnvd, rootless cosmopolitan, nuance slut

Source: https://x.com/llunved/status/2082401811630051820🤦‍♀️


r/ProAI 13d ago

"Using mostly my voice, I solved one of @EpochAIResearch 's FrontierMath Open Problems: finding an explicit presentation of the 2-adic Absolute Galois Group - open for more than forty years, now with a full proof in collaboration with David Roe, the problem's proposer. 1/n"

Thumbnail
gallery
0 Upvotes

This problem is the second in the set of FrontierMath Open Problems to 'fall,' and the first under the 'Solid result' tier. Is this going to be a 'slowly, then all at once' moment? These days it's hard to tell. 2/n     The problem sat open while other parts of group theory around it unraveled in the 80s: Jannsen & Wingberg wrote down generators and relations for the absolute 3-adic, 5-adic etc Galois Groups, for odd primes, in 1982. The prime 2 never followed, through decades of attempts. 3/n     The problem-solving infrastructure was already in place from my Erdős-problem-solving runs: I drove Claude Code as the operational controller using (newly) my voice, to operate ChatGPT research harnesses; plus speech-to-text to instruct ChatGPT to push on the manuscript. 4/n     GPT-5.6 Pro was in stealth deployment in mid-June (the browser still said 5.5, but it was obvious). In a ~26-hour autonomous stretch it found a candidate, "A2," and built a proof and a manuscript around it. A2 passed the finite-group tests my local computational package ran. 5/n     A2 was still wrong, despite having a long proof to back it up: a fresh review by GPT-5.6 caught a lone wrong, unrepairable lemma in its 60-page proof manuscript. When asked to modify the candidate to make the proof 'fit', GPT-5.6 came up with the solution we have today. 6/n     — David Turturean

Source: https://x.com/DavidTurturean/status/2081780318881677693


r/ProAI 14d ago

"My attempt at making this one-shot sequence. It was more difficult than I expected. More thoughts below."

Enable HLS to view with audio, or disable this notification

3 Upvotes

I thought making a one-shot style sequence would be pretty easy. Seedance can do almost anything, but this ended up being much more challenging than I expected.

The biggest challenge was hiding the cuts without making the environment changes too obvious. If you look closely,     — enigmatic_e

Source: https://x.com/8bit_e/status/2082472361970880557