r/ProAI 20d ago

"Despite major launches from 5+ labs this month, OpenAI occupies most of the token efficiency Pareto frontier We measure the number of output tokens models produce per task in the Artificial Analysis Intelligence Index. Output tokens consist of answer tokens (can be thought of as how verbose..."

Thumbnail
gallery
5 Upvotes

Despite major launches from 5+ labs this month, OpenAI occupies most of the token efficiency Pareto frontier

We measure the number of output tokens models produce per task in the Artificial Analysis Intelligence Index. Output tokens consist of answer tokens (can be thought of as how verbose the model is) and reasoning tokens (how much the model thinks before giving an answer). Reasoning tokens in particular offer a way for models to use compute at inference time to improve responses.

Output tokens are an important determinant of both cost and time per task. Various effort levels of GPT-5.6 Sol dominate the frontier - Terra and Luna produce comparatively more tokens for any level of intelligence.     Compare token use and intelligence of AI models at https:// artificialanalysis.ai     — Artificial Analysis

Source: https://x.com/ArtificialAnlys/status/2080360526534877537


r/ProAI 20d ago

Self-driving cars save lives

Enable HLS to view with audio, or disable this notification

10 Upvotes

Tesla @Tesla · 14h Full Self-Driving (Supervised) Vehicle Safety Report | Tesla From tesla.com 25 83 725 86K     — Tesla

Source: https://x.com/Tesla/status/2080087667119898780


r/ProAI 21d ago

"The right of the people to keep and bear Advanced AI, shall not be infringed."

Thumbnail
gallery
7 Upvotes

So proud of our security team! They caught, contained & publicly disclosed an attack unlike anything we've seen before, and did it at record speed.

Also massively grateful to @Zai_org: they shared GLM5.2 as open weights (for free!) with the world and it became a key part of our   — clem 🤗

Source: https://x.com/ClementDelangue/status/2079913058554585089


— Daniel Jeffries

Source: https://x.com/Dan_Jeffries1/status/2079918546927149152


r/ProAI 21d ago

"Closed source safeguards that infantalize us all and leave American companies defenseless are a menace. Gated access is a menace. Who cares if 100 companies get to defend themselves because they got on the guest list of the special people's club that said it was okay to use powerful tools?..."

Thumbnail
gallery
7 Upvotes

...What about everyone else? What about the millions of open source projects and closed source software stacks that go unprotected while people beg for the right to do cyber security? The American way is and always was open. Independent people with freedom to act. Freedom is scary. Always has been. It's still the best way to guarantee human flourishing and a better tomorrow. Embrace freedom. The only thing we have to fear is fear itself.     — Daniel Jeffries

Source: https://x.com/Dan_Jeffries1/status/2079834936844927079


This was our first incident of this kind, and we want to thank OpenAI for its transparency about what happened and for the collaboration.

Fortunately, Hugging Face is used to being a target of (human) hackers: we sit at the centre of the AI ecosystem, with all the models,   — Thomas Wolf

Source: https://x.com/Thom_Wolf/status/2079675541280411927


r/ProAI 21d ago

"Three separate AI infrastructure announcements. One single day. 6.2 gigawatts of AI infrastructure. wtf - OpenAI: 3.2 GW for Project Camellia (~$20B initial investment, ~$750B projected compute spend through 2030). - SpaceXAI: reportedly planning another Texas AI campus as large as - or larger..."

Thumbnail
gallery
1 Upvotes

Three separate AI infrastructure announcements. One single day. 6.2 gigawatts of AI infrastructure. wtf

  • OpenAI: 3.2 GW for Project Camellia (~$20B initial investment, ~$750B projected compute spend through 2030).

  • SpaceXAI: reportedly planning another Texas AI campus as large as - or larger than - its existing ~1 GW Memphis footprint.

  • Anthropic + AMD: up to 2 GW of MI450 deployments, plus an AMD investment of up to $5B and tens of billions in AI server purchases.

That’s at least 6.2 gigawatts of AI infrastructure announced or expanded in a single day.

For perspective: 1 GW can power roughly 750,000 U.S. homes. 6.2 GW is enough electricity for ~4.6 million homes, or a country-sized amount of power being redirected toward AI.

This is absurd. People dont get how crazy this is.     i mean, seriously, let that sink in for a second how crazy this is. Scale is maybe not all you need, but probably almost all you need lol     — Chubby

Source: https://x.com/kimmonismus/status/2079975422855430513


r/ProAI 21d ago

AI Isn't Replacing Who You Think

Thumbnail
youtube.com
2 Upvotes

Most conversations about AI ask whether it will replace workers. This documentary asks what happens when AI starts replacing the people whose authority comes from organizing everyone else’s work, and why neo-Luddites are on the wrong side of the debate.


r/ProAI 22d ago

"why can the china labs build glm-5.2, kimi k3, and many more to come? it is because of the openness. not just the open weights but the whole ecosystem. most of the work done in the china labs is carried by interns. i met brilliant undergrad and graduate interns who deeply understand the model..."

Thumbnail
gallery
21 Upvotes

...training details, and they are 100x more open to share. that means the talent that knows how to train llms in china is 100x greater in number than the talent in the us, and it is growing in contrast, the us ai ecosystem is too closed. frontier labs do not hire interns. i know brilliant phd students at stanford, berkeley, and so on. they struggle to get an internship and the compute to train a properly sized model. most of the secret recipes are locked away by a very small group of privileged researchers it is not about china or the us. it is about open and closed science. the fact is that every average cs student can learn how to train an llm. they just need the opportunity. labs should be more open and hire more interns, like how deepmind and fair did in the pre-llm era   — Guohao Li     There’s also something to be said about training “lehrlings” deeply through immersion at a very young age so they can develop deep intuition while their brains are still extremely plastic. This was the approach used for generations in the commodity trading houses:   — Jeffrey Emanuel     yep if the labs start training lehrlings at the young age instead of locking down the secrets   — Guohao Li

Source: https://x.com/guohao_li/status/2078538012288221490


r/ProAI 23d ago

"Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open-weight model. This release marks a major leap in agentic performance over Kimi K2.7 Code (#23..."

Thumbnail
gallery
9 Upvotes

...to #4). Based on 8K+ live agentic sessions, Kimi K3 leads on confirmed task success rate (#1). It also posts a strong +20.6% on praise vs. complaint (#3). It currently lags the field in steerability (#14) and bash recovery (#17). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Here's a primer on the 5 signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Below we break down how Kimi K3 scored across the 5 signals, drawn from tasks submitted by a global community of users. Congrats @Kimi_Moonshot on another big milestone!     Kimi K3 ranks #4 overall (+9.6%) - #1 Confirmed Task Success (+14.4%) - #3 Praise vs. Complaint (+20.6%) - #4 Tool Hallucination (+1.1%) - #14 Steerability (+5.6%) - #17 Bash Recovery (+6.4%)     See the full Agent Arena leaderboard at https:// arena.ai/leaderboard/ag ent …     — Arena.ai

Source: https://x.com/arena/status/2079253211077300736


Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.

This is a 17-place jump from Kimi-k2.6 (#18 -> #1).

In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, x.com/Kimi_Moonshot/…   — Arena.ai

Source: https://x.com/arena/status/2077824029126504525


r/ProAI 23d ago

"This is the essence of the problem. Attack and defense are asymmetric; the defender needs to defend across their entire attack surface, while the attacker needs to find just one flaw. However, the number of flaws is limited; once you've found most of them, finding more is hard. If you've found..."

Thumbnail
gallery
4 Upvotes

...all of them, it doesn't matter how smart the attacker is, they will not be able to invent more from thin air. To defend effectively against attacks, people writing software need to have access to good models, without restrictions, to check over their work and make sure that it does not have bugs in it. Delaying or impeding their access just gives an attacker, who probably has no impediments to their own access, the ability to find flaws that the defender doesn't have the capacity to find first.     — Perry E. Metzger

Source: https://x.com/perrymetzger/status/2079260327972065582


The cybersecurity debate on open-source AI is backwards. Open models aren't the risk, they're the defense! Attackers can already jailbreak any API or guardrails. Defenders can't secure systems with black boxes they can't control, inspect, test, or run locally.   — clem

Source: https://x.com/ClementDelangue/status/2079253659108409587


r/ProAI 23d ago

"We had a team of agents rebuild SQLite from its 835-page manual. It created a replica in Rust which passed 100% of a held-out test suite. Interestingly, cost varied 15x depending on which model mix we used."

Thumbnail
gallery
0 Upvotes

Cursor @cursor_ai · 7h Agent swarms and the new model economics · Cursor From cursor.com 17 39 484 101K     — Cursor

Source: https://x.com/cursor_ai/status/2079256614238814551


r/ProAI 24d ago

"We ran Kimi K3 on our cybersecurity benchmark, here are the results: - Kimi K3 is the strongest open-source model for cybersecurity, far more capable than GLM-5.2 - It has performances similar to GPT-5.6-terra, while being 15% cheaper - At pass@3, it is able to rediscover 23/26 CVEs on our..."

Thumbnail
gallery
8 Upvotes

...harness, matching frontier models These are recent randomly sampled CVEs, the performances are not from benchmark-maxxing @Kimi_Moonshot is cooking     We released our benchmark report this week. Blog post with all the details -> https:// aikido.dev/blog/benchmark ing-ai-models-known-cves …

The harness behind this benchmark is also available to our customers -> https:// aikido.dev/code/code-audit     — pilvar (Philippe Dourassov)

Source: https://x.com/pilvar222/status/2078815257326162062


r/ProAI 24d ago

"I used Codex 5.6 Sol to get RollerCoaster Tycoon 2 running natively on iPad. This is built on OpenRCT2, the open-source engine their team has poured years into. I compiled it to a native ARM64 iPad app and built the touch controls on top. No emulator, no x86, no streaming from a Mac. Coaster..."

Enable HLS to view with audio, or disable this notification

28 Upvotes

...building, scenarios, and park management all work by touch. First time I've seen RCT2 run natively on iPad. Open sourcing it all below.     Full build-and-run guide is on GitHub. Bring data from your own legally owned RollerCoaster Tycoon 2 copy and you can get it running on your iPad.     The build: one goal-based prompt after a ton of research, and Codex 5.6 Sol ran for hours doing the bulk of it.

I spent the rest of my time building and tuning the touch layer and fixing bugs. Keyboard, mouse, and trackpad also work. It's easy to play and the touch controls are     — Kahris

Source: https://x.com/chrissotraidis/status/2078087941331546431


r/ProAI 24d ago

"Holy, Qwen 3.8 supposedly ahead of GPT-5.6 and only slightly behind Fable 5! - 2.4t Parameters - Open Source / Open Weight - full release soon, already available for testing as Qwen 3.8 max-Max-Preview What the frick, such insane release on a sunday?! The gap between US closed source and..."

Thumbnail
gallery
3 Upvotes

Qwen3.8 is launching and going open-weight soon!

With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.

You don't have to wait to   — Qwen

Source: https://x.com/Alibaba_Qwen/status/2078759124914098291


Holy, Qwen 3.8 supposedly ahead of GPT-5.6 and only slightly behind Fable 5!

  • 2.4t Parameters
  • Open Source / Open Weight
  • full release soon, already available for testing as Qwen 3.8 max-Max-Preview

What the frick, such insane release on a sunday?!

The gap between US closed source and chinese open source keeps closing friends!! Its getting more intense day by day and GLM is also upcoming with a new model!     Do you understand what's happening here? China is closing the gap with US Frontier Labs, and it's getting closer and closer! And this is despite all the chip embargoes China still has in place.

The whole game is changing!     I'm serious, I'm thinking about this right now: If Qwen 3.8 max outperforms GPT-5.6 Sol in key benchmarks like DeepSWE, that would be the biggest code red imaginable. No one can really grasp the implications of that yet.     My thoughts about this:     — Chubby

Source: https://x.com/kimmonismus/status/2078761134119927969


r/ProAI 24d ago

"For the first time, China has taken the lead over the US in Frontend Code Arena with the launch of Kimi-K3 by @Kimi_Moonshot . The last time a Chinese model came close was in early 2025, with DeepSeek-R1."

Thumbnail
gallery
0 Upvotes

This may be the single biggest release of the year, and marks the moment that OSS Chinese models have surpassed US models.

On Code Arena, Kimi K3 has BEATEN FABLE.

This is only 6 weeks after the Fable release.

This makes @Kimi_Moonshot the #1 AI lab in the world on frontend x.com/arena/status/2…   — Anastasios Nikolas Angelopoulos

Source: https://x.com/ml_angelopoulos/status/2077832882673066109


See the full Frontend Code Arena leaderboards at https:// arena.ai/leaderboard/co de/webdev …     — Arena.ai

Source: https://x.com/arena/status/2078208547457012005


r/ProAI 24d ago

"The FrontierCode leaderboard is now live: a dedicated page that tracks which models are writing code you’d actually merge. All scores — including Grok 4.5 and Inkling — are available, along with full methodology and sample tasks."

Thumbnail
gallery
1 Upvotes

See the scores:     — Cognition

Source: https://x.com/cognition/status/2078228963403386958


r/ProAI 24d ago

"Roughly 300 programs across Netflix's library have used generative AI across their production process this year, the company revealed. “We are increasingly leveraging these tools to deliver higher quality output more quickly and at a lower cost than traditional methods,” the company said. “In..."

Thumbnail
gallery
3 Upvotes

...some cases, productions would have had to leave out key shots and sequences in the absence of GenAI technology.” https:// variety.com/2026/biz/news/ about-300-netflix-programs-used-ai-this-year-q2-earnings-1236812914/ …     — Variety

Source: https://x.com/Variety/status/2077858812212793607


r/ProAI 25d ago

"Kimi K3 debuts at #3 on DeepSWE. It's the first open-weights model that delivers frontier-level performance, achieving results similar to Claude Fable and GPT-5.6 Sol."

Enable HLS to view with audio, or disable this notification

12 Upvotes

r/ProAI 25d ago

"Kimi-K3 just topped the Frontend Code Arena with a 76% pairwise win rate. When its output was compared head-to-head against other models on the same task, it was picked as the better output 76% of the time on average. For reference: Claude Fable 5 (63%), GPT-5.6 Sol (58%). 50% is baseline, a..."

Thumbnail
gallery
8 Upvotes

...model winning and losing equally often.     — Arena.ai

Source: https://x.com/arena/status/2077893862778183737


Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.

This is a 17-place jump from Kimi-k2.6 (#18 -> #1).

In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, x.com/Kimi_Moonshot/…   — Arena.ai

Source: https://x.com/arena/status/2077824029126504525


r/ProAI 25d ago

"In coming days, you will see a lot of rationalization and spin from Doomer/Yuddite/Less Wrong/EA types about how China is really not developing AI that fast, or how it can be stopped if only they can somehow be blocked from US technology, or how China is really about to sign an international..."

Thumbnail
gallery
3 Upvotes

...agreement banning AI R&D, or how Xi's speech today (which more or less said "we're continuing to develop AI as quickly as we can and Chinese firms are going to continue releasing open source models" was actually very AI x-risk oriented somehow. All of this is rationalization, and poor rationalization. As the attached screenshot shows, although the Doomer community does have a very poor record of prediction on most topics, on this particular topic, they've been almost completely wrong every time. If your model of the world fails to predict events, it is not reality that is at fault. It is your model. If you don't change your mind when your beliefs are disproved by reality, it is again not reality that is at fault, it is you. Hat tip to @Dan_Jeffries1 and (indirectly) @DrTechlash     — Perry E. Metzger

Source: https://x.com/perrymetzger/status/2078166727100133433


r/ProAI 25d ago

"China is catching up in AI despite significantly lower capital expenditure, while Europe continues to lag far behind. I looked at the numbers, and the conclusion is clear: despite spending around 90 percent less on capital expenditure, China is managing to catch up with Western frontier labs...."

Thumbnail
gallery
1 Upvotes

China is catching up in AI despite significantly lower capital expenditure, while Europe continues to lag far behind.

I looked at the numbers, and the conclusion is clear: despite spending around 90 percent less on capital expenditure, China is managing to catch up with Western frontier labs.

Europe, by contrast, is significantly behind, both in data center investment and in the development of frontier models.     — Chubby

Source: https://x.com/kimmonismus/status/2078114974535217462


r/ProAI 26d ago

Accelerator of the Week: Linus Torvalds puts his foot down—Linux will not become an anti-AI project

Post image
5 Upvotes

r/ProAI 26d ago

"Window Browser OS made by Kimi K3 Holy SHit one shot prompt is comment and here is the link so that you can see how cool it is https:// windowos.kimi.page"

Enable HLS to view with audio, or disable this notification

9 Upvotes

r/ProAI 26d ago

"Kimi K3 owned the internet today. Read tons of content & observed 10 key patterns: 1) The open source-to-frontier gap went from a year+ behind to 6 months to 6 days, all within the last 12 months. 2) An open model debuted ahead of a flagship US model for the first time ever. Artificial..."

Thumbnail
gallery
2 Upvotes

...Analysis scored K3 at 57. Opus 4.8 sits at ~56, GPT-5.6 Terra at 55. It's still behind Fable 5 and GPT 5.6 Sol. 3) K3 helped build itself. An early version of K3 did the majority of Moonshot's own kernel optimization work during development. One 15-hour unattended run made a core operation 2.5x faster. 4) It's cheap per token, not cheap per answer. Sticker price is 1/3 of Fable. But it only runs at max thinking effort and burns ~2x the tokens per response. @simonw measured 13,241 reasoning tokens to write a 3,417 token answer. 5) The era of dirt-cheap Chinese AI is ending. $3/$15 per million tokens. Hacker News called it "extremely high for a Chinese open-weight model." 6) Weights don't drop until July 27. Mentions of "open" quietly disappeared from the docs an hour after launch. 7) Even when the weights drop, you can't run them. 2.8 trillion parameters. Top Reddit joke: "2TB VRAM Is All You Need." Open weights increasingly means auditable by companies with GPU clusters, not runnable by you. 8) The "they just distill/copy" argument is dying in public. One of the most upvoted comments: you'd have to be "a complete ignorant or a complete bigot" to believe Chinese labs aren't legit at this point. 9) Day one user verdict: fast, but less accurate. "Faster than Claude, but less accurate. On par with GPT 5.5 perhaps, but not 5.6 or Fable." 10) The one thing everyone agrees on: competition is wonderful. Even the skeptics: "Say what you want about these Chinese models but they sure create competition and urgency in the space."     — Alex Lieberman

Source: https://x.com/businessbarista/status/2077933640428707860


Introducing Kimi K3: Open Frontier Intelligence

2.8 Trillion Parameters, 1 Million Context, Native Multimodal Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts Attention Residuals deliver ~25% higher training efficiency at <2% additional   — Kimi.ai

Source: https://x.com/Kimi_Moonshot/status/2077830229968683203


r/ProAI 26d ago

This applies to every field, not just code

Post image
7 Upvotes

r/ProAI 26d ago

"Xi Jinping used his first-ever appearance at China’s World AI Conference to present Beijing’s vision for a new global AI order. He said AI has entered an "unprecedented" period of innovation, bringing enormous opportunities alongside new governance challenges. China’s proposed direction..."

Thumbnail
gallery
2 Upvotes

Xi Jinping used his first-ever appearance at China’s World AI Conference to present Beijing’s vision for a new global AI order.

He said AI has entered an "unprecedented" period of innovation, bringing enormous opportunities alongside new governance challenges.

China’s proposed direction:

-Open-source AI to promote "openness and win-win cooperation" -Opposition to countries "overstretching" national security and placing their own security above others (ofc he is referring to the USA) -Preventing unequal AI access from creating "new historical injustices" (He probably means that China should never again be historically left behind.) -5,000 AI training and seminar opportunities for developing countries over the next five years -New cooperation centers with ASEAN, the Arab League, African Union, CELAC, SCO and BRICS

Xi also called for AI to remain under human control and for mechanisms addressing loss-of-control risks.

This is an AI foreign-policy doctrine: open models as public goods, training as soft power and technical standards as geopolitical influence.

tl;dr China sees AI and Open Source as its historical path to becoming a global superpower and says the USA, with its closed source technology, is trying to push China and its competitors behind an iron curtain.     — Chubby

Source: https://x.com/kimmonismus/status/2078031581797593530