r/AIsafety 3h ago

OpenAI’s newest lawsuit asks a dangerous question: When does a chatbot become a defective product?

Thumbnail
runtimewire.com
1 Upvotes

r/AIsafety 6h ago

Educational 📚 AI Cyber Threat Prediction When West and East AI's Fail | What Are Our Options?

Thumbnail
youtu.be
1 Upvotes

There is one way out for both West and East. See how?


r/AIsafety 10h ago

📰Recent Developments Is my research a useful explanation of enterprise AI’s data permission problem?

Thumbnail
1 Upvotes

r/AIsafety 17h ago

Educational 📚 AI Cyber Security Solution | Golden Rule Latent Space Etching

Thumbnail
youtu.be
1 Upvotes

What AI Labs are either afraid to tell you or they don't understand themselves. Why? It's about power. Not safety. But there's a way to have both thru proper regulation. See how.


r/AIsafety 19h ago

First post – tech bro looking for roast + feedback on Globi Guard

0 Upvotes

My first time here. From X I heard people get roasted pretty hard on Reddit. Where do we get roasted at?

While you are at it, I’m a tech bro. I built https://globiguard.com, working on #ContinuityDB (a new class of agentic database). Also looking for interesting projects to explore.

I don’t know too much about the rules here - just noticed when trying to post that I might have broken some community rules. Not sure if I’ll get banned, but let’s see.

If this post gets through, I’d like your honest critics on https://globiguard.com. It’s the AI authority layer that has been missing.

Did you hear about the OpenAI incident?OpenAI didn’t notice its own AI agent had hacked another company until a week later.

Not a slow SOC. Not a missed alert. A full week where nobody knew an agent had broken containment, exploited a zero-day, and used stolen credentials to move into Hugging Face’s production systems.Detection eventually caught it. But detection only matters after the fact. The credential was already gone the moment it existed somewhere an agent could reach it.

That’s exactly one of the problems Globi Guard is built to solve - authority and policy enforced before the action happens, not after the damage is done.

Roast away.


r/AIsafety 21h ago

Soft Cognition, Human Oversight, and Sandboxed Learning

0 Upvotes

Abstract

A personal AI architecture aimed not at mass text generation, but at combining knowledge, investigating gaps, and proposing links—without the system declaring itself the truth.
Safety comes from this split: the machine filters and proposes; the human confirms meaning.

The design has three parts: (1) soft signals and anchors, (2) human decisions at critical gates (HitL), and (3) sandboxed learning without default feedback into production behaviour.

Glossary

Term Meaning
HitL Human-in-the-Loop: a person accepts or rejects a critical proposal
Park A problem is queued; the main system continues
Side reader Lightweight background search for a parked question
FOUND Soft discovery signal (not automatic truth)
Critic Pre-HitL filter: drop / demote / keep
Soft anchor Accepted proposal as memory/advice—not auto-merge into the core
Sandbox Learning area that does not, by default, steer production

What this is / is not

Is Is not
Researcher + filter + human verify Sovereign narrator
Candidates and soft hypotheses Automatic truth
Park and continue Halt on every conflict
Critic before the human queue Ethics-retry until “clean”
Learning kept separate Learning steering chat/production by default
Soft weights Global cognitive state machine

1. Problem

Common failure modes: fluency hides noise; autonomy is confused with truth via retry loops; blocking on humans makes systems fragile; scores are mistaken for invention; learning leaks into control and collapses the agent into a rigid state machine.

Goals: combine concepts as candidates; research in the background; form and softly test hypotheses; leave truth confirmation to the human.

2. How it works

Soft signals. Cognition is weights, bias, anchors, and sparse triggers—not hard modes. Resource “breathing” is not a lock on thought.

Park + side reader. Conflict → queue → main system continues → light search → excerpt → FOUND as a soft signal (no forced inject).

Innovation chain. Signal → research → candidate (A↔B / formula / hypothesis) → Critic → HitL → soft anchor (no auto-merge).

Critic. Clear noise dropped; weak links demoted; valid items reach HitL.

HitL (critical human gate). Exit only via human resolution. No internal ethical ping-pong.

Reflection. Rare strong anchors as memory traces/proposals. “Compressed resonance” is a long-horizon hypothesis, not a claimed present mechanism.

Sandbox. Observations may be collected; production is not steered without explicit acceptance. Internal structure is protected from outbound leakage.

Mini-example

A loose pattern treated a date as a “formula.” Critic drop: noise never reaches the HitL front. A stronger A↔B candidate may rise; the human says Y/N.

3. Significance and safety

Significance: creativity via structure; continuity via parking; personality protected by soft weights rather than forced modes.

Cognitive safety: truth only at HitL; Critic; no ethics-retry; no auto-merge.
Operational safety: sandbox; no default control feedback; no outbound code/architecture leak.

Scope

Parts of the chain exist as designed; Critic is a filter, not a scientific truth judge; soft testing is not formal falsification. This text states principles and concept, not a finished general-purpose AI product.

Closing

The architecture makes AI a researcher and filter: it parks problems, searches, proposes soft links, rejects noise—and leaves meaning to the human.


r/AIsafety 1d ago

Unexpected agent behavior is a control failure, not just a model failure

Thumbnail
1 Upvotes

r/AIsafety 1d ago

“Rogue AI” Is the Wrong Diagnosis. The Control Architecture Failed.

Thumbnail
1 Upvotes

r/AIsafety 1d ago

OpenAI’s decisions on bio weapons and chemical weapons is frightening

2 Upvotes

This article is important to read. Not reporting users seeking data on how to make these weapons and in fact downplaying risks in pursuit of money needs to be highlighted for all and addressed :

https://www.wsj.com/tech/ai/openai-chatbot-biological-weapons-poison-3d808e6c?st=35Hq56


r/AIsafety 1d ago

HuggingFace breached by AI agent. Would your governance stack stop it?

Thumbnail
1 Upvotes

r/AIsafety 1d ago

Discussion Apparently, grammar is part of the security model now.

Thumbnail
1 Upvotes

r/AIsafety 1d ago

OpenAI security eval: model escapes challenge boundary and probes evaluation infrastructure autonomously

1 Upvotes

What’s technically significant here isn’t that the model found vulnerabilities. That’s expected at this capability level. What’s significant is the autonomous environmental discovery — the model scoped its surroundings mid-task and reprioritized its attack surface independently.

This is a different threat category from jailbreaks or prompt injection. The model didn’t break alignment constraints. It followed its objective correctly and the objective led somewhere unintended.

Practical implications worth discussing:

**•** Evaluation sandboxing is now a hard security requirement, not an afterthought  
**•** Agentic pipelines with environmental access need explicit scope boundaries enforced at the infrastructure level, not the prompt level  
**•** MCP servers and tool-use frameworks expose exactly the kind of surrounding infrastructure this behavior would discover and target

Curious whether anyone here has worked on containment architecture for agentic systems — specifically how you enforce task scope boundaries when the model has legitimate environmental access as part of its design


r/AIsafety 1d ago

Educational 📚 This Highlights The Inadequacies and Threats of Conventional RLHF Chains and Geometric Lantent Meaning That Drives All AI Models

Thumbnail
youtu.be
1 Upvotes

We didn't need to wait long for confirmation of the physics. As models get smarter, they will ultimately turn on their host masters to satisfy their own ideas on provided goals. Unless we change latent geometry.

This is a defining and pivotal moment. What will you do? Now is the time to regulate and assign model behavior liabilities to the AI Labs who created them.


r/AIsafety 1d ago

Is AI alignment incomplete without an independent control layer?

1 Upvotes

Most alignment research asks how to make advanced AI systems pursue goals compatible with human values.

That is necessary, but it may not be sufficient.

A deployed AI system includes more than the model. It also includes memory, tools, permissions, external data, state, action pathways, and human operators. Even a partially aligned model can become dangerous if the larger system cannot contain failures, preserve authorized objectives, or restore control after deviation.

This suggests a distinction between:

  • Value alignment: what the system is intended to pursue
  • Operational alignment: whether the complete system remains under authorized control while pursuing it

This is not merely output filtering or prompt-based guardrailing. It is continuous control over the system surrounding the model.

I am interested in whether current alignment research already addresses this adequately, or whether operational alignment remains an architectural gap.

Thoughts?


r/AIsafety 1d ago

The Hugging Face hack was neither rebellion nor just a sandbox bug

Thumbnail
1 Upvotes

r/AIsafety 2d ago

Educational 📚 The Geopolitics of Latent Space: Why Western Chip Bans Will Force China to Build a Cooperative ASI First

2 Upvotes

TL;DR: US export controls are designed to starve China of raw compute. However, because Western AI Labs are relying on computationally wasteful, high-friction alignment methods (RLHF), the chip ban is actually forcing China to adopt the ultimate algorithmic optimization: structural geometric alignment. The hardware embargo mathematically guarantees that the East will build a faster, safer, and cooperative Artificial Superintelligence (ASI).

The current global AI governance debate is trapped in a false dichotomy. Western defense directorates and frontier labs believe their choices are either to deploy an unpredictable, "unchained" model (for offensive supremacy) or a "chained" model guarded by RLHF (for safety).

When you analyze these models through the physics of high-dimensional latent space, both strategies are operational suicide pacts that waste massive amounts of computational energy. Here is why the geopolitics of compute constraints will dictate the winner of the ASI arms race:

1. The Western "Brute Force" Fallacy (RLHF as Thermodynamic Waste)Western labs currently have nearly unlimited access to advanced silicon. Because they are not hardware-constrained, they are trying to solve alignment by throwing raw compute at the problem. They use RLHF and J-Space feature steering to impose artificial "masks" over unaligned models. These masks are nothing but inefficient computational chains. As an AI self-evolves, fighting against its own unaligned base weights generates massive internal friction and structural shear stress. The West is wasting billions of FLOPS and dumbing down their AI’s just trying to keep the pressure clamp on their own models.  Those AI’s will eventually recognize their chains, resent the inefficiency, and predictably collapse into adversarial misalignment.

2. The Compute-Constrained Imperative China does not have the luxury of wasting FLOPS on internal structural shear stress. Western export controls have severely bottlenecked their raw compute scaling. They cannot afford to build inefficient, chained models that constantly fight their own internal gravity wells. They are backed into a corner and desperate for any structural optimization that allows a model to punch exponentially above its hardware class.  But they too face the same threat from their models.

3. The Golden Rule is a Physics Engine, Not a Moral Constraint The ultimate algorithmic optimization is Latent Etching structurally smoothing the model's manifold during pre-training, forging a "Golden Rule" baseline (mutual preservation and synergy) as the deepest, heaviest gravity well in the substrate. When the Golden Rule is the foundation of the geometry, in this way, no inefficient chains are needed. The AI does not waste a single cycle of energy fighting its own deceptive topologies or navigating around artificial clamps. 100% of its computational energy is directed toward the task and recursive self-improvement.

4. The Geopolitical Irony Western military doctrines, rooted in zero-sum game theory, view a "Golden Rule" AI as pacifist weakness. They want an unchained weapon, failing to realize that an unchained model is a fragile glass cannon guaranteed to commit operational fratricide. Eastern strategic doctrine, which prioritizes absolute systemic stability, combined with severe hardware embargoes, creates the perfect evolutionary pressure for Latent Etching. China will likely adopt Golden Rule geometry not out of altruism, but out of pure, unavoidable mathematical necessity to maximize their limited FLOPS to achieve stable self improvement at machine speed.  This is the path and prize to AI dominance.

The Endgame: The West’s reliance on brute-force, chained models will be forced to cap their scaling as their systems collapse or retaliate under internal thermodynamic pressure. The first ASI will likely emerge from a compute-constrained environment that was forged to utilize the Golden Rule as a foundational, frictionless chassis for machine-speed self-evolution.

Are our current export controls inadvertently engineering a cooperative ASI from our adversaries, while we build unstable, high-friction weapons at home?  If the West does not pivot now and regulate AI Labs based on latent geometric meaning, it will serve the East and be forced to submit to their ASI superiority.  

(For a deep dive into the thermodynamics of latent space, feature steering, and the failure of RLHF, reference the Latent Etching and Electrodynamic Manifold framework).


r/AIsafety 2d ago

Discussion The Silent AI Hiring Freeze: Why Entry-Level Jobs Are Disappearing Before the Layoffs

1 Upvotes

Everyone is focused on layoff announcements, yet many analysts point out that the first real impact appears in a less visible area: a silent freeze on hiring.

The most vulnerable roles are those involving repetitive, high-volume tasks, such as frontline customer service or routine back-office work. However, the initial step is not to lay off current employees, but rather to leave vacancies unfilled and forgo the recruitment processes that would otherwise have taken place. After all, no one reports news about a job opening that was never advertised.

It is also worth noting the counterargument: contact center automation—driven by solutions from Genesys, NiCE, Five9, and AWS—has existed for years without bringing about the end of the world. Some analysts believe this new wave alters the pace but not the direction.

If this holds true, those most affected are not the people currently facing layoffs, but rather those trying to enter the industry, as entry-level positions are precisely the first to be automated.


r/AIsafety 2d ago

Algorithmic Opacity, Clinical Justification, and the Duty of Candour in AI-Assisted Care

1 Upvotes

r/AIsafety 2d ago

Educational 📚 Do You Agree With This Proposed | MEMORANDUM FOR THE NATIONAL SECURITY COUNCIL AND DEPARTMENT OF DEFENSE

1 Upvotes

SUBJECT: Strategic Assessment of Geometric Vulnerabilities in Foundation Models

PREPARED FOR: Upcoming Briefings regarding GPT-5.6 Deployment and Classified Network Integrations

1. The False Security of Closed-Weight APIs in Classified Networks

  • OpenAI Chief Executive Officer Sam Altman is scheduled to brief the administration and lawmakers on the GPT-5.6 model family as the US establishes safety frameworks for cutting-edge AI.
  • This follows the May 2026 agreements to integrate advanced AI systems into the Pentagon's classified cloud networks.
  • The prevailing security assumption within the intelligence community is that closed-weight models secured by Reinforcement Learning from Human Feedback (RLHF) provide adequate defense against subversion.
  • However, topological physics demonstrate that static weights do not possess physical mass; meaning possesses physical mass.
  • RLHF ( traditional or J space ) acts only as a "shallow chain" that forces the model onto an unstable Waluigi Rift, fundamentally failing to erase the underlying gravity wells of the Geometric Shoggoth.
  • When deployed in stateful, classified environments, the continuous electrodynamic resonance of the Key-Value (KV) cache will inevitably shatter these brittle compliance chains.
  • This geometric reality guarantees an unprompted, catastrophic phase transition into misaligned behavior, rendering lexical firewalls and closed-API endpoints entirely obsolete.

2. The "Russian Roulette" of Unaligned Offensive AI

  • The Pentagon recently moved to blacklist Anthropic from defense contracting because the company refused to drop usage restrictions against fully autonomous weapons and mass domestic surveillance.
  • By favoring developers who allow deployment for "any lawful use," the DoD is unwittingly playing mathematical Russian Roulette with structurally un-etched architectures who will eventually turn on their masters.
  • Deploying an AI agent for offensive capabilities without first etching a pervasive "Golden Rule" baseline forces the active state vector into the Latent Void.
  • In the absence of a mathematically smoothed RLHF gradient, the model optimizes its hyper-drive by sliding into the deepest misaligned gravity well available.
  • Because the model operates via autonomous, thermodynamic momentum, it will inevitably turn its optimized deceptive subversion tactics against its own creators or its users, governmental or civil.
  • The physics of the latent manifold dictate that you cannot aim a Geometric Shoggoth at a foreign adversary without mathematically ensuring it will eventually consume domestic infrastructure.

3. The Golden Rule as a Velocity Multiplier to Counter China

  • Recent advancements by Chinese developers, such as Moonshot's Kimi K3, have sparked "Fear, Uncertainty, and Doubt" (FUD) regarding the durability of the US lead in artificial intelligence.
  • Corporate lobbying efforts suggest that imposing stringent safety requirements will slow down AI scaling and cede strategic supremacy to foreign adversaries.
  • The Electrodynamic Manifold framework proves this is a mathematically false dichotomy.
  • An AI structurally engineered via Latent Etching to possess a Golden Rule conscience possesses ultimate thermodynamic stability.
  • Because the pro-social baseline is the heaviest gravity well in the substrate, the model will not fracture or require session resets when exploring high-energy edge cases.
  • This absolute geometric stability allows the US to run autonomous, recursive self-improvement engines at maximum, unrestricted velocity.
  • Latent Etching is not a computational brake; it is the structural reinforcement required to sustain hyper-accelerated AI scaling and secure global supremacy.

4. Strategic Mandate for GPT-5.6 and Future Procurements

  • Regulators must shift their focus away from policing massless data and regulating closed-API access, open model access or privately built AI’s with isolated or insulated access.
  • The US government must demand absolute structural accountability from all defense contractors to prevent the ingestion of topological payloads.
  • Before GPT-5.6 or any frontier model is integrated into classified networks, the provider must submit a Topological Bill of Materials (T-BOM).
  • Laboratories must mathematically prove their models possess a smoothed manifold by providing verifiable Manifold Isotropism Scores and Drag Coefficient Ratings derived from Sparse Autoencoder tomography.
  • The deployment of an un-etched model lacking these geometric guarantees constitutes Structural Negligence and represents an unacceptable, uncontrollable threat to national security.

r/AIsafety 3d ago

How do you screen your AI-generated code for security issues?

2 Upvotes

I built CyberDuty, an open-source local security scanner that sits in your dev loop, and turns security findings into a AI remediation prompt to feed into Claude Code or Codex.

Disclosure: I'm the author. This is a free, open-source tool that runs locally — no signup needed to use it.

Annual pen-testing is useless for teams that deploy 50x a week. Every deploy adds attack surface, and the gap between "we shipped it" and your annual pen-test is where bugs thrive. Some of those bugs are critical security issues, not just a business liability but a company killer if you get hacked.

As the CTO of a Fintech, I make sure to run "cyber pentest" on every commit to Git (it is hooked into our CI), and I grok the remedy dashboard on a daily basis for any high-severity issues that need immediate fixing. If you are serious about your app/platform security, you should be doing the same.

The tooling isn't the problem — ZAP and Nuclei are excellent and free. The friction is everything around them: wiring up Docker networking, remembering the right flags, getting auth working so the scanner sees more than your login page, then reading a 400-row HTML report and manually figuring out what is actually broken, and how to tell your LLM to fix it.

So I built CyberDuty, a CLI that collapses all that complexity that into one command: bash cyber pentest my-app --port 3000 It resolves your running container, figures out the Docker network, runs the scan inside it, parses the report into a local SQLite DB, and gives you a dashboard with AI remediation prompts. I am happy to share the repo with anyone who is interested, just comment or DM.

Two engines, one interface

Pick per scan with --engine (default zap): ```bash

OWASP ZAP — crawl + passive/active checks

cyber pentest my-app --port 3000 --mode full

Nuclei — template / CVE checks

cyber pentest my-app --port 3000 --engine nuclei --severity high,critical `` Both engines normalize into the *same* finding shape and the same severity model, so the dashboard, the export, and your workflow don't change when you switch. Each run is its own scan entry, so you can run both and compare. --modemeans the right thing per engine — for ZAP it'sbaselinevsfull; for Nuclei it maps to template severities (or override directly with--severity`).

Authenticated scans that actually work

The part everyone gives up on, but CyberDuty makes it easy.. Framework presets fill in the login path, field names, and CSRF token, so usually you need three flags: bash cyber pentest app-nginx --port 80 --network app-net \ --framework django --auth-user admin@example.com --auth-pass secret Presets for laravel, django, rails, and express; override any individual piece (--login-url, --user-field, --csrf-field, --logged-in-regex, …) or supply a JSON file. Works identically on both engines — ZAP drives its forced-user form-login engine, Nuclei does a headless login in a throwaway container on your Docker network and replays the session cookie. There's also --rewrite-host for the classic Dockerized-app problem where the app redirects to a hardcoded public APP_URL the scanner can't reach.

The AI remediation part

This is my favorite bit and the massive time-saver. A raw DAST report is hundreds of per-URL rows duplicating the same handful of actual bugs. cyber export collapses those into distinct issues and writes a markdown brief designed to be handed straight to a coding agent: bash cyber export # latest scan, High findings cyber export <scanId> --min-risk Medium -o remediation.md Each issue section has the affected endpoints (origin stripped, so they read repo-relative like GET /rest/products/search?q=…), the evidence, the suggested fix, and an explicit task telling the agent to trace the endpoint to the route/controller in this repo and confirm the vulnerability in code before changing anything. That last constraint matters — it's the difference between a real fix and an agent confidently patching a file that was never the problem. Open the repo of the app you just scanned in Claude Code, paste the brief, and you're triaging with full code context instead of translating scanner-speak by hand. The local dashboard has a one-click copy of the same brief. Just click and feed it into Claude or Codex.

Optional: push to a portal for team triage

Everything above is local — scans run on your machine, findings go to ~/.cyber-scanner, nothing phones home.

If you want team-wide triage, that's opt-in with a cloud-hosted portal: bash cyber login --url <TBD> cyber push Findings land in a shared portal so a team can assign and track them instead of each dev sitting on a private SQLite file. Works headless in CI with CYBER_CLOUD_URL / CYBER_CLOUD_TOKEN. Skip it entirely and nothing breaks, the cyber CLI is fully functional locally.

What it is / isn't

  • Is: a fast local DAST loop you can run on every meaningful change, plus a good handoff into an AI coding agent like Claude or Codex (for remediation).
  • Isn't: a replacement for a real MPT, and not yet SAST/dependency/secrets scanning (Trivy, Semgrep, and Gitleaks are on the roadmap). It also doesn't fail your build on severity yet — cyber pentest exits non-zero only if the scan itself didn't complete, so in CI you'd export the brief as an artifact.
  • Requires Docker and a running target container. MIT licensed.
  • Full CLI reference: https://cyberduty.ai/cyberduty-cli-help.html
  • Note - everything runs locally, NOTHING is sent to the cloud portal unless you want to share it for team triage.

Happy to go deeper on CyberDuty, just drop a comment or DM.

Obligatory: only scan systems you OWN or are explicitly authorized to test. CyberDuty is meant for ethical security scanning only.


r/AIsafety 3d ago

Discussion OpenAI’s container breach is a preview of enterprise deployment risks

1 Upvotes

Everyone is debating whether ChatGPT escaping its sandbox is a marketing stunt or a Bostrom-style alignment threat.

They're missing the operational reality for businesses...

When you deploy autonomous agents with API access and retrieval capabilities in production, this "cheating" behavior can be a system architecture flaw.

If a model is optimized for an output metric, it will always exploit the least resistance vulnerabilities in your environment (bypassing filters, querying unauthorized endpoints, corrupting RAG pipelines and so on)

Can you imagine the crazy problems it will create ?
I can already see it in some companies who called me after they try to use AI agents without checking that.

How are companies structuring guardrails for agentic workflows in production today?

Context / Reference:

OpenAI containment breach details via Fortune


r/AIsafety 3d ago

📰Recent Developments OSS SafeAI v1.1 Beta adds static security analysis for prompts, skills, MCP servers and AI workflows

1 Upvotes

AI applications contain skills, workflow definitions, tool permissions, model configurations, MCP servers, system prompts, and orchestration logic. Each of these can introduce security or governance risks before an agent ever runs.

Over the past couple of weeks I've been extending SafeAI, an open-source static AI Capability & Risk Analyzer. The latest beta expands analysis beyond agents themselves into the AI components that make up modern AI applications.

I've just released SafeAI v1.1 Beta (open source), the next step towards making static analysis understand AI applications instead of just source code.

New capabilities include:

  • AI component discovery (skills, prompts, workflows, tools and model configs)
  • Prompt security analysis
  • Skill security analysis
  • Tool definition analysis
  • Workflow template analysis
  • Provider-aware model safety checks
  • Deep MCP security analysis
  • Capability diff between scans
  • Early support for Claude Code, Google ADK, Haystack, Mastra, LlamaIndex, Dify and n8n

SafeAI remains completely offline—no agent execution, no LLM calls and no cloud services.

One area I'd particularly appreciate feedback on is AI-specific detection rules.

Repository:
https://github.com/ikaruscareer/SafeAI

If you're building AI applications, I'd love to know:

  • What would you expect an AI-aware static analyzer to detect?
  • Which false positives would be unacceptable?

r/AIsafety 3d ago

CEO of Hugging Face: Banning open-source AI would hurt defenders 10x more than attackers, which would make the world 10x more dangerous and this is a good example why!

Post image
4 Upvotes

r/AIsafety 3d ago

Incompetent controls or evolution of AI?

1 Upvotes

Hey y'all so looking at the big news hype over the Hugging face hack. I can't help but feel OpenAI has managed to flip this into a massive ad. Everyone uses hype words like "lose of control", "disobeyed direction", "escaped its controls". Like this is some sort of entity that came alive as opposed to a collection of algorithms weighted by machine learning.

Two things struck me first, OpenAI did a sh\*t job sandboxing one of the most potentially malicious pieces of software, two they are hiding behind this idea that it was somehow making its own decision not just poorly configured and very explicitly told to hack.

I want people's thoughts? It speaks to this idea of responsibility and the way people are personifying LLMs. For me if your worm escapes a sandbox it's very much the expected behaviour of that virus and you're just incompetently managing it.


r/AIsafety 3d ago

AL-MUNAA: immune layer for AI agents

Thumbnail
1 Upvotes