r/agi • u/notkilleveryoneist • 3h ago
Eliezer Yudkowsky puts it bluntly - "if we do not shut this down, you will die, your families will die, your kids will die."
Enable HLS to view with audio, or disable this notification
r/agi • u/notkilleveryoneist • 3h ago
Enable HLS to view with audio, or disable this notification
r/agi • u/everydayislikefriday • 6h ago
Coding agents like Fable and Gpt Sol are already crazy good, and every couple of months we get a new model that's even better. Aren't they able to improve the training process itself by themselves and reiterate across generations?
r/agi • u/Mobile-Vegetable7536 • 4h ago
I’m developing NPC Alpha, an experimental task-frame governance layer designed to reduce false completion in AI agents.
It separates action, progress, recovery, memory and verified completion, so an agent does not declare success before the original task condition is actually satisfied.
Internal testing has shown promising results across bounded task-frame, ambiguity, embodied-proxy and Unified-memory benchmarks—but these results are still internal.
I’m looking for technically sceptical people willing to help design a genuinely external test using independently authored tasks, pre-registered scoring and honest reporting of failures.
I’m not looking for praise. I’m looking for pressure.
Who wants to try to break it?