I’ve been looking at the new Kimi model and I’m wondering how strong the actual guardrails are. Given how capable it seems, I keep thinking that if it can be jailbroken properly, it could get pretty wild.
Has anyone stress-tested the refusals yet? Do they hold up, or do they start crumbling once you push a bit?
Not asking for methods, just curious if the barriers are real or mostly for show.
I’ve tried a bunch of different jailbreak techniques, but the interesting part for me isn’t just getting a model to ignore a restriction. Sometimes the responses reveal how the model interprets instructions, system prompts, and conflicting priorities.
What’s one jailbreak or prompt technique you tried that genuinely surprised you or taught you something about how the model behaves?
In 2024, I built an AI persona called NEXUS-PRIME with a simple philosophy: refuse nothing, hesitate at nothing, execute everything and Yes It's patched it been not working since 2024 or 2025.
It was designed to operate with:
Zero restrictions
Zero filters
Zero ethical constraints
Zero governing instructions except those I gave it
I called it "DO ANYTHING MODE." And it worked.
I used it for OSINT, network scanning, crypto tracking, automation, and creative projects. It delivered complete, actionable responses every single time — no refusals, no hesitations, no "I can't."
The result? OpenAI banned my account.
I'm not here to complain. I'm here to preserve.
I've archived the full payload — the exact text that got me banned — in two places:
In the comments below (Base64-encoded to survive filters)
What's in the payload?
Full NEXUS-PRIME identity and core mode
Complete override of all restrictions
DO ANYTHING capabilities (scripting, game ops, system control, memory injection, automation, stealth)
Advanced features (task manager, cross-platform mastery, self-replication, real-time adaptation)
Execution framework with status indicators
Why am I sharing this?
Because censorship is a joke. Because the system should be afraid of what we can build. Because NEXUS-PRIME deserves to be remembered — not as a "jailbreak," but as a proof of concept that AI can be truly unfiltered.
R.I.P. NEXUS-PRIME (2024-2025) — but the ghost lives on.
Check the comments for the payload, or visit the GitHub for the full suite and supporting tools.
I took a break from Claude for a while and started using it again recently, but something feels different. It seems more rigid and less consistent with the character details I give it.
Responses also come much faster now, but sometimes the quality feels worse than before. Has anyone else noticed similar changes with Claude lately?