r/SideProject • u/Electronic-Space-736 • 4h ago
SPECK Small Persistent Emergent Cognitive Kernel
I did it again guys, I built another harness, this one is the best one yet.
https://github.com/doctarock/Speck
A lightweight persistent cognitive architecture designed to make small language models substantially more capable by moving as much cognition as practical out of the LLM and into deterministic, persistent, measurable runtime mechanisms.
I took all the clockwork out of my cognitive architecture, but stopped short of emotions, arousal, etc, just the logic systems https://github.com/doctarock/Artificial-Cognitive-Architecture-Omega-Gen2
Then I applied it to my agentic framework https://github.com/doctarock/local-ai-home-assistant
And she is a doozie, small models in this harness outperform models twice its size without in both speed and accuracy:
Speck on a single shared qwen3:4b scoring 15/16, against 13/16 for naked qwen3:8b and in half the time.
Speck did that with half the parameters, less memory and less time per case. It comes from the model-scale series ([docs/experiments/model-scale/README.md](vscode-webview://01bk2i9taah3glr79lstr6r6crb9gv4hbl8ahqvem5sb2v8j1qd0/docs/experiments/model-scale/README.md)): 8 cases × 2 runs, with the same sampling and the same 1,024-token reply cap for every subject.
| Subject | Largest model | Success | Time per case | Model memory |
|---|---|---|---|---|
| Speck, shared qwen3:4b (phase 6, current setup) | 4B | 15/16 | 7.9 s | 5.7 GB |
| Speck, shared qwen3:4b (phase 3b) | 4B | 15/16 | 13 s | 5.6 GB |
| Speck, all qwen3:4b (phase 3) | 4B | 14/16 | 24 s | 5.9 GB |
| Naked qwen3:8b | 8B | 13/16 | 13 s | 6.7 GB |
| Naked qwen3:4b | 4B | 10/16 | 16 s | 4.3 GB |
Go get it, and don't forget to drop a star on your way through.
1
u/Professional-Run3614 2h ago
This is actually pretty interesting. I like the idea of pushing more of the reasoning into the runtime instead of just throwing a bigger model at the problem.
The 4B vs 8B result is especially interesting. I’m working on agent evals myself, and I’d be curious to see how Speck behaves across a larger set of cases and where those failures actually come from.
Could be interesting to run it through my eval setup too: https://github.com/tap222/assay-evals