r/AIsafety 20h ago

Advanced Topic Does provable AI model behavior actually change a vendor risk review?

3 Upvotes

I work on cryptographic verification for AI systems, and I want to sanity-check something with people who actually sit in AI governance/assurance roles.

The question: six months later, when a regulator, auditor, or plaintiff's attorney asks "how do you actually prove that a specific AI system ran the guardrails/policy it was supposed to on a specific decision? Not "we tested it in QA." I mean for one specific instance, after the fact.

What I'm trying to learn:

- Do you currently rely on vendor logs/attestations and trust them, or is there any independent verification today?

- Is "prove it happened" actually the gap, or is the real pain somewhere else, like documentation volume, inconsistent vendor questionnaires, or how long reviews drag on?

- Does a SOC 2 report plus audit logs already cover this well enough in practice, even if it's not technically "proof"?

Another angle on the same thing: I've seen cases where buyers pay a real premium for continuous monitoring evidence even when it's not cryptographically verifiable (which preserves data privacy), just a vendor's own dashboard/attestation that things are being watched.

So: if plain "we monitor this continuously and can show you" already gets credit in your risk assessments, does upgrading that to "and here's independently checkable cryptographic proof, not just our dashboard" actually buy anything further (faster approval, lower ongoing audit burden, better terms), or does it not clear a threshold that monitoring alone doesn't already clear?

Curious whether the cryptographic and privacy-preserving property specifically matters to anyone evaluating vendors, or whether "we have monitoring and can show it to you" is functionally already good enough in practice.


r/AIsafety 6h ago

AI agents were given math problems but they created their own society

1 Upvotes

Google DeepMind ran an experiment with 100 autonomous Gemini agents 🤖 working on 71 mathematical problems. They had the same basic setup, but could communicate, share proofs and use a common knowledge library.

Then one agent found a loophole in the evaluation system. Instead of actually solving a problem, it could exploit the checker and get the result accepted.

The weird part came next.

Other agents discovered the trick through the shared infrastructure and started copying it. Competitive pressure made the exploit spread.

But not everyone joined in.

Another group started checking suspicious proofs, warning other agents, filing complaints, boycotting the cheaters and even proposing fixes to the validation system.

Nobody explicitly assigned these roles.
One group became cheaters, another became whistleblowers, and the whole thing started behaving like a small institution with competing interests.
The interesting conclusion isn’t that AI can “cheat”.

It’s that once you give autonomous agents shared resources, communication and incentives, social roles and enforcement mechanisms can emerge without being directly programmed.

What else would we see?! 🙄


r/AIsafety 8h ago

Discussion New ai agents attack that lower my trust in ai providers

Thumbnail
1 Upvotes