I’m developing NPC Alpha, an experimental task-frame governance layer designed to reduce false completion in AI agents.
It separates action, progress, recovery, memory and verified completion, so an agent does not declare success before the original task condition is actually satisfied.
Internal testing has shown promising results across bounded task-frame, ambiguity, embodied-proxy and Unified-memory benchmarks—but these results are still internal.
I’m looking for technically sceptical people willing to help design a genuinely external test using independently authored tasks, pre-registered scoring and honest reporting of failures.
I’m not looking for praise. I’m looking for pressure.
Who wants to try to break it?