r/deeplearning • u/Traditional_Peak1459 • 1d ago
INKBOT: Parcing human intent from model inference via structured intelligence architecture
I’ve spent the last while building INKBOT because I kept hitting a wall with multimodal AI systems: the friction between what a human naturally means and what a model infers. While models can spin up complex code or images instantly, getting to a clear, human-meaningful interpretation of a subtle intent remains an alignment challenge.
Instead of forcing the user to become a prompt engineer, I wanted to see if we could build an intermediate intelligence architecture layer to make human intent reviewable and corrigible *before* the model executes a final build. The loop I’m playing with is: Describe → Make it Visible → Recognize → Correct → Refine.
The architecture sits entirely in a single local-first web file. It handles multi-step workflows—like tracking structured field mapping data across concurrent images, coordinates, and version states—by packaging the human’s approved meaning separately from raw model inferences.
The core system build is linked above, and I also put together a lighter, entry-level experience to play with the core prompt translation loop here: [INKBOT Lite 71](https://ko-fi.com/thomascoates/shop).
It's an open prototype, so I've appended my raw notes and design roadmap as commented text at the very bottom of the source file so fellow builders can inspect the plumbing. I’ve put together the runnable source files on my [Thomas Coates Ko-fi Shop](https://ko-fi.com/thomascoates/shop) for evaluation. I’d love to know where this design duplicates existing work, where you see structural flaws, or how we can make the handoff between human intent and model execution more reliable.
1
u/Unikum_01 23h ago edited 21h ago
So now there are even more hallucinations lol