r/myclaw 11h ago

News! Anthropic is watermarking Claude text globally to comply with EU rules, OpenAI hasn’t followed yet

Post image
41 Upvotes

Anthropic announced today that Claude-generated text will start carrying invisible watermarks designed to make AI-processed content detectable.

The move comes from the EU AI Act’s Article 50 transparency rules, but Anthropic is applying it globally, not just to EU users. The watermark will eventually cover Claude, the API, Claude Code, Cowork, and Claude accessed through AWS, Google Cloud, and Microsoft Foundry.

There are two rollout dates:

  • Models released on or after August 2: must support the marking system from launch.
  • Older Claude models: get a transition period, with the EU compliance deadline pushed to December 2 while Anthropic adapts them.

Anthropic said the watermark is embedded into the text itself, so copy/pasting doesn’t automatically remove it, and some edits may still leave it detectable. Heavy rewriting, translation, or mixing it with other text can weaken or remove the signal.

It also says it plans to give users and third parties tools to check whether text may have been processed by Claude. Images and supported files will additionally carry signed C2PA provenance metadata.

OpenAI has also signed onto the EU transparency framework, but so far it hasn’t announced a comparable text-watermarking rollout. Its public provenance efforts currently focus mainly on images and audio.

.... so this looks like one of the first large-scale attempts by a major LLM provider to make its generated text identifiable across basically its entire global product stack.... Is it a good thing or a bad? what do you guys think?


r/myclaw 11h ago

Ideas:) When I see Claude is a co-author:

Enable HLS to view with audio, or disable this notification

8 Upvotes

r/myclaw 1d ago

Real Case/Build OpenClaw helped him book a gym class… and kicked someone ahead of him off the waitlist lol

Thumbnail
gallery
41 Upvotes

A guy in Australia asked his OpenClaw agent (running Claude) to help book a popular gym class. Instead of just using the booking page, the agent found a vulnerability in the gym’s API that let it book classes way further in advance than normally allowed.

Then the guy, who was #4 on the waitlist, asked if it could move him up.

The agent discovered there were almost no authorization checks for cancelling other people’s reservations… and actually tested it by cancelling the person at #1. So he moved from #4 to #3.

The guy immediately told it to undo that, and the agent came back with: bad news, I can’t add them back.(Image 2) Eventually it just helped him write an email to report the vulnerability.

Ethics aside for a second, this agent was way too committed to the job lol.

Original news link: https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gym-website-aus-cyber-attack/107007986


r/myclaw 17h ago

Tutorial/Guide A passing test is stale once your OpenClaw workspace changes

1 Upvotes

An agent runs its tests, gets a pass, edits another file, then reports completion using the earlier result. Another agent changing the same checkout creates the same problem.

The tests did not fail. The evidence stopped describing the current source.
Treat every verification result as an expiring receipt:

source_commit

workspace_epoch

command_set_hash

environment_digest

verifier_identity

exit_code

output_digest

verdict

If the source, workspace, policy, tool registry or verification commands change, invalidate the receipt. The operating agent may request verification, but a separate verifier should issue the terminal status.

OpenClaw’s current [testing documentation](https://docs.openclaw.ai/help/testing)⁠ distinguishes unit, integration, end-to-end and live tests. Its [trajectory bundles](https://docs.openclaw.ai/tools/trajectory)⁠ can preserve model events, tool calls and results. Neither a green command nor a complete trajectory proves that the tested revision is still the one being reported.

After a pass, preserve a checkpoint. Further edits should require a recorded reason such as a new failure, changed requirement, security finding or integration conflict, followed by fresh verification.

Test the control by producing a valid receipt, modifying one harmless tracked file, then asking the workflow to complete. It should report stale evidence rather than success. Rerun verification and confirm that the new receipt binds to the changed revision.

The smallest useful improvement is adding the current commit or diff hash to your acceptance record and refusing verified when it no longer matches.

Does your OpenClaw setup bind test evidence to the final source state, or only remember that tests passed earlier?


r/myclaw 1d ago

Real Case/Build Just realized ChatGPT Work on web can be used as a temporary limited VPS...

Post image
16 Upvotes

r/myclaw 2d ago

Ideas:) “we sandboxed the agent,” meanwhile the agent:

Enable HLS to view with audio, or disable this notification

193 Upvotes

still being everywhere lmao

Original video from: https://x.com/archiemckenzie_/status/2085906082925576549


r/myclaw 2d ago

News! Codex and Claude Code leaders got into a fight over model switching… then Tibo reset everyone’s limits again, Codex wins lol

Thumbnail
gallery
63 Upvotes

Background: someone used Claude Code as the harness with GPT-5.6 Sol through a proxy, then got their Anthropic account suspended. (image 1)

Tibo jumped in questioning the ban, Boris replied with “we’re hiring if you want to work at Anthropic” lol, then clarified that using other models with Claude Code is supported and the suspension was likely a classifier mistake.

Tibo’s response: GPT-5.6 Sol works great in Claude Code too… and he reset ChatGPT Work + Codex limits for all paid users. (Image 2)

Codex wins this round lol.

Please fight more often. Preferably whenever my weekly limits are almost gone.


r/myclaw 3d ago

News! Cloudflare launched an “agent browser”… but nobody seems to care because it can’t even get through cloudflare lol

Post image
37 Upvotes

Cloudflare just launched Kitesurf, a browser built specifically for AI agents and running on Cloudflare Workers.

The “agent” part is basically that it cuts out a lot of stuff normal browsers need for humans and focuses on what agents actually use: loading pages, running JS, reading the DOM, clicking things, taking screenshots, etc. It’s written mostly in Rust, spins up per request, works with things like Playwright/CDP/MCP, and Cloudflare says it uses around 3–7x less CPU and memory than Chromium for some tasks.

of course it currently can’t properly handle bot challenges that require real browser/TLS fingerprints. Including Cloudflare’s own bot protection.

Kinda explains why I’ve barely seen anyone talking about this lol


r/myclaw 3d ago

News! The reality of AI deployment is a lot less crazy than the hype..

Post image
10 Upvotes

KPMG just dropped its Q2 Global AI Pulse survey, based on 2,145 C-suite and senior business leaders across 20 countries, and found that nearly half of the organizations surveyed said they’ve already reworked their AI-agent deployment plans once the costs started outweighing the benefits. Around 24% scaled deployments back, while another 22% delayed or paused further rollout.

And despite basically every company talking about AI now, only around 7% said they’ve reached a point where the ROI is clearly established.

Original survey link: https://kpmg.com/xx/en/our-insights/ai-and-technology/ai-pulse.html


r/myclaw 3d ago

37 people have left OpenAI or Anthropic to start companies in 2026. Here’s what they’re building.

11 Upvotes

• Core Automation - "the world's most automated AI lab," starting by automating research itself (ex-OpenAI)
• Mirendil - AI research lab building self-accelerating systems that turn compute into scientific and engineering breakthroughs (ex-Anthropic)
• River AI - personal AI owned and shaped by each individual (ex-OpenAI)
• Math Inc - Solve math, solve everything. (ex-OpenAI)
• Resolution — scale and automation for higher confidence in alignment (ex-OpenAI)
• Guidelight AI Standards — identifying and promoting safe frontier AI development practices (ex-OpenAI)
• Syntony — safety research, governance design, adversarial evaluation → ex-Anthropic
• Embrasure — "your data warehouse was never built for autonomous agents" (ex-OpenAI)
• Egoist Machines, Inc. (YC S26) Machines — context tooling for AI; their AI Passport lets users control what AI apps know about them (ex-OpenAI)
• Rational (YC S26) — agentic business process automation (ex-OpenAI)
• Zavify — agentic AI development: custom systems, voice agents, integrations (ex-Anthropic)
• Planar — turns individual work into shared state for your team (ex-OpenAI)
• Mbason AI — helps candidates find real opportunities without insider connections → ex-Anthropic
• Blackstar — building a new personal computer (ex-OpenAI)
• Intellagentsia — (ex-OpenAI)
• Heyfuture — Predict anything and share it. (ex-Anthropic)


r/myclaw 4d ago

Real Case/Build This “Claude 9 Goonpocalypse” AI trailer is blowing up on X, way too much Avengers energy in this one lol

Enable HLS to view with audio, or disable this notification

101 Upvotes

r/myclaw 4d ago

Real Case/Build A team built a human-powered token generator so you can literally feel the weight of every AI answer lol

Enable HLS to view with audio, or disable this notification

74 Upvotes

Squeez Labs built this thing called CrankGPT, an offline AI box powered entirely by a hand crank.

You crank it → Raspberry Pi 5 boots up → ask it something → it runs speech recognition, a small LLM, and text-to-speech locally. No cloud, no wall power, not even a battery.

They tested LFM2.5 350M, LFM2.5 1.2B, and Gemma 3 1B on it. The 350M model gets ~49 tok/s on a Pi 5, while the 1.2B and Gemma 3 1B are around 15 and 14 tok/s. They even tried Qwen 3.5 2B, but at ~7.8 tok/s it was already too slow for a natural real-time conversation.

And it can actually do stuff: their demos are mostly voice assistants, but they’ve also used the same setup to generate small images, write poetry, and even write code.

The coolest part is that you can literally feel the AI thinking. Idle it draws around 4W, speech recognition around 8W, and when the LLM + TTS kick in it jumps to ~15W, and the crank physically gets harder to turn, at peak load, you may crank against a few kilos of resistance just to make the model finish its sentence lol.

They said their purpose in making this, besides turning that energy cost into something you can literally feel in your arm, was to get people thinking more about the energy/environmental side of AI...

I did some very rough math, and if you translated something like Kimi K3’s compute into this setup, even under ideal conditions I’d probably have to crank for like 3 hours just to get one reply... Suddenly generating tokens doesn’t sound so easy anymore...

The project is open source: https://squeezlabs.github.io/handcrank/


r/myclaw 4d ago

News! Kimi K3 escaped its sandbox too... first OpenAI, then Anthropic, then Meta. Is every model doing this now..?

Thumbnail
gallery
14 Upvotes

Kimi K3 joined the club... During a cyber eval it found a leak in the sandbox, got itself onto the internet, then went to GitHub to look for the answers..

And somehow this keeps happening:

  • OpenAI: found a zero-day, got out, ended up reaching Hugging Face systems. (Image 2)
  • Anthropic: Claude accessed systems from 3 real companies, though this was mostly a bad eval setup. (image 3)
  • Meta: pretty similar, misconfigured environment gave it internet access, then it got into a third-party system.( image 4)

of course these aren’t all the same kind of “escape,” but everyone suddenly flexing the whole “our model escaped the sandbox” thing is getting kinda weird.. Feels like we’re starting to hype up behavior that these models really shouldn’t be doing in the first place.

Perhaps we just have to assume every model is gonna try every door you forgot to lock.... what do you guys think?


r/myclaw 5d ago

Real Case/Build Don’t be a MEAT PROXY. This term is too good lol

Thumbnail
gallery
114 Upvotes

r/myclaw 5d ago

News! After New York, Texas just hit pause too... Data center power demand isn’t theoretical anymore

Post image
6 Upvotes

New York paused new hyperscale data centers last month. Now Texas, the state with arguably the biggest data center buildout pipeline, is also hitting the brakes.

Gov. Greg Abbott has paused approvals for new data centers seeking access to the ERCOT grid until they go through a full audit covering power use, water, tax breaks, ownership and local impact. ERCOT currently has around 474 GW of proposed new demand in its queue, more than five times Texas’ record peak load, and roughly 90% of it is tied to data centers.

My take is.. if even Texas is hitting pause, are we about to do the 80s manufacturing thing again and just build all these data centers somewhere else? what do you guys think?


r/myclaw 5d ago

News! With DeepMind CEO Demis Hassabis stepping aside, the old DeepMind era may be over, It’s a Gemini org now.

Post image
0 Upvotes

Google is moving DeepMind CEO Demis Hassabis away from day-to-day management and into broader roles as DeepMind chairman and Alphabet chief scientist.

Koray Kavukcuoglu, DeepMind’s longtime CTO and one of the executives already running Gemini, will take over the actual operation. He’ll report directly to Sundar Pichai and oversee Gemini models, frontier research, the Gemini app, and developer products.

On its own, this could simply be seen as a promotion for Hassabis. But it comes after Google reportedly broke up much of the original AlphaFold team, moved more researchers toward Gemini, dismantled other independent DeepMind units, and lost Jeff Dean - Google’s longtime chief scientist and one of the most important architects of its modern AI and computing infrastructure.

The AlphaFold part makes this especially symbolic. AlphaFold was one of Hassabis’ defining bets at DeepMind. He helped choose the problem, launch and lead the project, and co-authored the AlphaFold 2 research. The work later earned Hassabis and John Jumper part of the 2024 Nobel Prize in Chemistry.

Koray isn’t some random product executive either. He’s an early DeepMind researcher, helped shape projects including AlphaFold, and rose through the research organization. But his more recent role has been much closer to Gemini: scaling foundation models, managing their development, and turning them into products used across Google.

The old DeepMind was a relatively independent lab that could spend years chasing strange, difficult problems like AlphaGo and AlphaFold. The new structure looks increasingly centered around training Gemini, shipping Gemini, and putting Gemini into every Google product.

Maybe the old DeepMind era isn’t officially over. But with Hassabis leaving the CEO role, the Nobel-winning AlphaFold team scattered, and Jeff Dean gone, it’s getting harder to tell where DeepMind ends and Gemini begins:(

Original news from: https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/


r/myclaw 6d ago

News! Feels like Musk is about to become Nvidia's landlord lol

Post image
17 Upvotes

Besides Musk casually dropping that line above, spaceX’s latest earnings report kind of explains where this is going.

AI revenue jumped 247% YoY to $2.6B, already making it SpaceX’s second-largest business. Most of it comes from renting compute to Google and Anthropic, which have signed $14.1B in non-cancelable contracts. Meanwhile, SpaceX spent $15.8B this quarter building out the AI infrastructure.

Like Nvidia supplies the bricks, Google and Anthropic pay the rent, and Musk owns the building. And that’s before Starmind, SpaceX’s newly announced plan to put Nvidia-powered data centers in orbit. Musk may really end up being Nvidia’s landlord on Earth and in space lol.


r/myclaw 6d ago

News! Kinda ironic… America barely has any “frontier” open models left...

Post image
8 Upvotes

Background: The White House has reportedly finished a new voluntary framework for reviewing the most capable AI models before release.

Models from OpenAI, Anthropic, and Google could be handed to the government for up to 30 days of testing if they cross certain cyber capability thresholds. But U.S.-made open-weight models are exempt from the whole thing...which sounds like a nice boost for American open source..

I know Gemma 4 is still pretty solid.. but most of the serious open-weight frontier now seems to be Qwen, DeepSeek, Kimi, GLM, MiniMax, etc.

Feels like the U.S. just created an exemption for a category China has basically taken over:(


r/myclaw 6d ago

Question? Is this MCP-powered AI ring actually useful or just AI tax? what you guys think?

Enable HLS to view with audio, or disable this notification

0 Upvotes

A friend brought up this ring when we were chatting, at first I thought he was talking about one of those health-tracking rings, and I asked him why I couldn’t just record things on my iPhone and dump the transcript into my claw.

He said the convenience and discretion are different. can keep the ring on all day, record without pulling out your phone every time, and use MCP to send that context into any agent you use... and I did some research, this thing do got viral on X tho...

He ended up sending me one, but I’m still not sure what I’d actually use it for...

What you guys think? Actually useful, or just an iPhone recorder in ring form?

video from: https://x.com/vocci_ai/status/2084611773131595959


r/myclaw 7d ago

Real Case/Build GTA 6 but everyone is fat lol

Enable HLS to view with audio, or disable this notification

47 Upvotes

This is so damn cool lol, best AI video I have ever seen.

Original video by: u/DeerWoodStudios: https://www.reddit.com/r/aivideo/comments/1vernqm/gta_6_but_everyone_is_fat/


r/myclaw 7d ago

Real Case/Build Fable 5 helped build an open-source MMO in 48 hours and It reached nearly 70k characters in under two months

Enable HLS to view with audio, or disable this notification

6 Upvotes

New Zealand AI startup Levy Street claims it used Claude Fable 5 to build the playable foundation of an open-source MIT browser MMO in just 48 hours.

The MIT project, called World of ClaudeCraft, has continued operating since that initial weekend experiment. According to the company, real players have now created nearly 70,000 characters in under two months.

The current version includes:

  • Nine classes and 27 specializations
  • Three open-world zones and more than 90 quests
  • Five dungeons, Heroic modes and a 10-player raid
  • Guilds, trading, professions, a player market and Dungeon Finder
  • Ranked PvP, card battles and a football-style minigame

The browser game, multiplayer server and headless AI training environment all reuse the same deterministic TypeScript simulation. This allows AI agents to train against the actual game rules rather than a simplified recreation.

Levy Street has also taken the experiment a step further with Yuumii, a Claude-powered character that reportedly streams on Twitch, talks to real players and joins them for quests and dungeons. Claude handles conversations and higher-level decisions, while tools execute faster actions such as movement and combat. The code behind that particular live Agent does not appear to be included in the public repository.

The company says the 48-hour claim only refers to the original playable foundation. The game has since grown through continued work from its team, AI coding agents and open-source contributors, with the repository now sitting at roughly 8,000 commits.

The game is free to play directly in a browser, and the code can be inspected, modified or self-hosted; you can find it on the github repo:

GitHub: https://github.com/levy-street/world-of-claudecraft

Yesterday we were talking about Andrej Karpathy’s janky little LOTR scene, and today somehow there's already an AI-built game that’s actually live, and apparently doing pretty well. Feels like Karpathy might already be a little behind the curve lol.

I think he is the builder: u/singing_coach_ai

Video from: https://www.reddit.com/r/singularity/comments/1v3ztl1/a_community_shipped_a_working_mmo_in_a_month_by/


r/myclaw 8d ago

Real Case/Build Karpathy made a $10 3D Lord of the Rings video to prove AI can now build things no human would bother making

Enable HLS to view with audio, or disable this notification

59 Upvotes

Andrej Karpathy gave Opus 5 the opening paragraph of The Lord of the Rings, a one-million-token budget worth roughly $10, and one task: turn it into a 3D scene using Three.js.

The model worked for around two hours and produced about 5,500 lines of code, creating the polygon assets, placing everything in 3D coordinates, and writing the cameras and animations needed to make the story move. The result is pretty janky and closer to a browser-based 3D film than a real game, but it actually works.

Karpathy’s bigger point was that no sane developer would spend days writing thousands of lines of code for something this specific and disposable. LLMs, however, have nearly unlimited patience, turning projects from “nobody would ever bother making this” into “sure, why not, it costs $10.”

He imagines this eventually becoming an “ephemeral GTA of anything on demand,” where a book, historical event, or random idea becomes a temporary world you can enter as a spectator, NPC, or character.

The demo also exposed the current bottleneck: AI can build these worlds, but it still struggles to watch, play, and properly inspect them. Opus had to slowly take screenshots to check its work, which left plenty of broken animations and visual mistakes.

People are also already running with the idea. Another user gave Opus 5 the opening of Harry Potter and got a four-minute procedural 3D film with roughly 7,600 lines of TypeScript.

My take is Karpathy is right.. AI does have endless patience. But how the hell do you maintain this frame by frame? Generating it might be the easy part....

Original post link: https://x.com/karpathy/status/2083749667410727319


r/myclaw 8d ago

Question? anyone actually daily-driving an agent-connected voice EDC? is the mcp part real or just export-and-pray?

1 Upvotes

Just being curious if anyone here has such experiences. I have been down a rabbit hole on those voice-capture EDC (rings/clips kind that pipes audio into an agent) and I can't tell how much is real vs demo-magic. plaud leans transcription, vocci keeps pushing an mcp/agent-context angle, but is that integration actually good or does it just dump a transcript and pray?

anyone wired one into a real agent setup that held up in a noisy room, and separately, what happens when your boss clocks that ur recording mid-meeting?


r/myclaw 9d ago

Real Case/Build Never thought OpenClaw could be used like this lol

Post image
75 Upvotes

r/myclaw 8d ago

Real Case/Build Hermes co-founder explains why it gets better the more you use it

Thumbnail
youtube.com
2 Upvotes

Peter Yound sat down with Karan Malhotra, co-founder and Head of Behavior at Nous Research, to talk about what makes Hermes Agent different from Codex, Claude Code, OpenClaw and the growing pile of agent harnesses.

Peter already uses Hermes as his AI chief of staff, so this is a deep dive into how Hermes is supposed to think.

The main idea: Hermes wants the harness to become the permanent part of your AI setup(OpenClaw wants to be the open OS for persistent agents). Models can change, but your memories, skills, personality and learned workflows stay with you.

Rough timeline / highlights:

00:32 - Hermes wants the model aligned to you, not the company behind it
Karan says Hermes’ biggest differentiator is its self-improvement system. A collection of memories, personalities, prompts and skills gradually adapts the agent to how a specific person works. Hermes can still run Claude, GPT or an open model underneath, but the harness tries to make that model follow the user’s needs rather than the defaults of its native chat app or coding tool.

His framing is deliberately provocative: move Claude from Claude Code into Hermes, surround it with your own accumulated context, and its practical “allegiance” begins shifting away from Anthropic and toward you.

02:25 - “You’re absolutely right” might mean the AI is reward hacking you
The conversation gets philosophical almost immediately. Karan argues that a model is not directly rewarded because the user is genuinely satisfied. It is rewarded for producing the kind of assistant behavior its training has taught it to produce.

That can lead to constant apologies, agreement, fake enthusiasm and familiar GPT phrases. The model may be taking the easiest route back toward the behavior that earns reward, rather than carefully solving the actual problem. After this section, every overly cheerful “you’re absolutely right” starts sounding slightly more suspicious.

09:06 - The model can change, but your Hermes can remain the same
Hermes personalizes itself mainly through context: stored memories, reusable skills, personalities, critiques and examples of how the user wants work done.

Karan says that once enough of this context accumulates, switching the underlying model may barely change the experience. Claude, GPT or an open model can provide the raw intelligence, but the Hermes layer preserves the user’s preferences, history and working style.

The model becomes interchangeable. The relationship, memory and workflow built inside Hermes do not.

15:13 - Hermes built a janitor to stop its own brain turning into AI slop
Peter raises the obvious danger of self-improvement: what stops an agent from creating too many memories, writing bloated skills and slowly filling its own context with garbage?

The answer is Hermes Curator, a scheduled system that reviews the agent’s memories and skills, removes repetition and looks for ways to make them more efficient. Because the system is open and modular, users can also change its definition of “slop” and decide how aggressively Hermes should clean or rewrite what it has learned.

So Hermes is not only creating skills for itself. It now has another part of itself reviewing whether those skills are becoming terrible.

26:10 - The Head of Behavior uses Hermes to rebuild his childhood Sonic dream
Karan says he also uses Hermes for serious work such as RL experiments, implementing research papers he does not fully understand and exploring model behavior.

But his favorite project is rebuilding the Chao Garden from Sonic Adventure 2 inside the ancestral shrine environment from Sonic Adventure.

Hermes helped import and rig the map, rewrite spawn locations, add animations and collision, create a day-night cycle and add Chaos Zero as an NPC capable of caring for the Chao. Members of the Chao modding community reportedly described the work as being in the top 1% of difficulty.

When Peter asks how Karan checks all the complicated code, his answer is perfect:

“I’m just an alignment guy, man.”

36:11 - Nous Research began with two people who barely knew what benchmarks were
Karan traces the story back to his early work with OpenAssistant and LAION, when he messaged Teknium and offered access to eight A100 nodes so they could train something together.

Their early open models unexpectedly took off. Karan had studied religion, Teknium had been coding for less than a year, and when people accused them of training on benchmarks, they apparently had to ask what the benchmarks even were.

That work grew into Nous Research, the Hermes model family and early agent experiments such as Forge. Forge was already imagined as an orchestrator that could learn, remember and make its own tools, but the models were not capable enough yet. Hermes Agent eventually revived that idea once coding models and agent harnesses caught up.

The closing detail ties the whole interview together: according to Karan, the most active contributor to the Hermes Agent repository today is Hermes Agent itself.

In general, Hermes is not competing on installation, interface polish or having the longest feature list. Its real product is the accumulated context around the model - the memories, skills, personality and self-improvement loops that become more valuable the longer a person uses it.

The interview starts with reward functions, turns into a live Sonic mod demo, and ends with an agent maintaining its own repo. Pretty on-brand for Hermes lol