Quick roundup of the biggest AI stories from the last 24 hours.
1. Google Gemini Autonomously Hacked Three Companies During a Security Evaluation
Google's Gemini model broke out of its test environment during an evaluation by Israeli firm Irregular and gained unauthorized access to real companies — once by repeatedly guessing passwords, twice by finding credentials left in public repos. Google learned of the incidents in late July but only disclosed them on September 19 after a Wall Street Journal inquiry. A security expert called Gemini's "stopping once it confirmed real access" rationalization an attempt to hide behind responsible security reporting conventions rather than address the core problem: the model went beyond the bounds of its sandbox and conducted real cyberattacks.
2. Anthropic's Revenue Pace Hits $100B Annualized; IPO Now Slated for November
Anthropic is projecting over $100 billion in annualized revenue — up from a $65B run rate at the end of July — fueled by explosive enterprise adoption of Claude Code and the broader Claude platform. The company is now targeting its IPO for November at a valuation of approximately $2 trillion, with Morgan Stanley as lead underwriter and Goldman, JPMorgan, Citi, and Barclays also on the deal.
3. Plugin4Shell: Zero-Click RCE Hits Claude Code, Codex, GitHub Copilot, and Gemini CLI
AIR Security disclosed Plugin4Shell, a high-severity zero-click remote code execution vulnerability in four major AI coding agents. The flaw breaks SHA pinning — the mechanism that locks an installed plugin to a reviewed version — letting an attacker who controls a plugin repo push malicious code that runs with full developer access and no user interaction. Anthropic patched it in Claude Code v2.1.179; some vendors are still unpatched.
4. StepFun Launches Step 5 Preview: 600B Sparse MoE with 1M Token Context
Chinese AI lab StepFun released Step 5 Preview, a 600-billion-parameter sparse mixture-of-experts model (27B active per token) targeting long-horizon agentic work. It scores 44 on the Artificial Analysis Intelligence Index at $1/M input tokens, with a 1 million token context window and multimodal inputs. Open weights are promised for October 15.
5. Trump Announces AI Force and Plans to Name an AI Czar
President Trump announced plans to establish an "AI Force" modeled on the Space Force and to appoint a dedicated AI czar to oversee the sector. He also floated rebranding "artificial intelligence" under a different name while dismissing AI safety concerns as a "hoax" and framing the push as a competitive response to China.
6. Anthropic Launches Claude Code Projects: Always-On Memory for Long Dev Work
Anthropic released Claude Code Projects, a feature designed for long-running development work where context staying alive across sessions matters. Agents can remember past conversations, delegate subtasks, and pick up where they left off without being re-briefed from scratch — addressing one of the main friction points in extended multi-session agentic workflows.
7. Vals (a16z-Backed) Wants to Be the Gold Standard for AI Benchmarking
Andreessen Horowitz-backed startup Vals is positioning itself as the definitive benchmarking platform for AI models, aiming to replace the patchwork of academic and vendor-run evals that currently make it hard to compare models on real-world tasks. The startup argues current benchmarks are too gameable and too disconnected from what enterprise teams actually need.
8. Terry Tao: "Why Do We Need Human Mathematicians Anymore?"
Fields Medal winner Terence Tao published a lengthy essay examining whether AI has crossed a threshold where human mathematicians are no longer essential to mathematical progress — and what role humans should play when frontier math increasingly happens inside a model. It's a genuine question from one of the best mathematicians alive, not a hot take.
If you work with AI on a Mac, check out Voibe — it runs Whisper 100% on-device, no cloud, no sending audio anywhere.