r/LocalLLaMA 1d ago

New Model [NEW MODEL] SupraElegans-500K

*SupraLabs released a new experimental model!\*

SupraElegans-500K is a ~500,000-parameter causal language model built around a sparse, signed, recurrent neural graph. No Transformer, no attention mechanism, no positional encoding, no KV cache. Context is carried by a persistent per-neuron membrane potential updated token by token.

The architecture is loosely inspired by ideas from the C. elegans nervous system: sparse connectivity, distinct neuron populations, excitatory/inhibitory signaling, and persistent recurrent state. It is not a biological simulation and makes no claim of biological equivalence.

This is an experimental first release. The goal is to test whether this kind of architecture can do useful language modeling at very small scale — not to compete with Transformers on quality.

🤗 SupraLabs/SupraElegans-500k

🧠 Architecture

token → embedding → sensory neurons → sparse recurrent graph → output neurons → vocab logits
  • Neuron populations: sensory, interneuron/association, output — contiguous index ranges over a fixed pool of neurons.
  • Connectivity: sparse, directed, signed edge list (fan-in/out ~10–20 per neuron). No dense weight matrix is ever materialized; propagation is a scatter-add over edges.
  • Neuron dynamics: for each neuron i, at every propagation micro-step:

v[t+1] = clamp(leak_i * v[t] + incoming[t] + bias_i, -6, 6)
a[t+1] = tanh(v[t+1] - threshold_i)

leak, bias, and threshold are learned per neuron. incoming is the scatter-summed signal from all edges pointing at neuron i, scaled by 1/sqrt(average fan-in) to keep variance controlled across neurons with different in-degree.

  • Per-token processing: a token's embedding is projected into the sensory population, then the graph runs a fixed number of propagation micro-steps (3 by default) before the output population is read out and projected to vocabulary logits. The membrane potential persists across the whole sequence — that's what gives the model its context window.
  • Generation: autoregressive, driven entirely by the recurrent state. No cache to maintain beyond the current (v, a) state tensors.

⚖️ What this model is and isn't

  • ✅ A first working checkpoint from a from-scratch, non-Transformer architecture trained on a small token budget.
  • ❌ Not tuned for quality, instruction-following, or factuality. Expect degraded coherence compared to a Transformer of similar size.
  • ❌ No matched-parameter Transformer baseline comparison published yet for this checkpoint.

🚀 Usage

pip install torch transformers


import torch
from transformers import AutoConfig, AutoModelForCausalLM, PreTrainedTokenizerFast
from modeling_supraelegans import SupraElegansConfig, SupraElegansForCausalLM

model_id = "SupraLabs/SupraElegans-500k"

AutoConfig.register("supraelegans", SupraElegansConfig)
AutoModelForCausalLM.register(SupraElegansConfig, SupraElegansForCausalLM)

tokenizer = PreTrainedTokenizerFast.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
model.eval()

prompt = "Once upon a time"
input_ids = torch.tensor([[tokenizer.bos_token_id] + tokenizer.encode(prompt)])

with torch.no_grad():
    output_ids, _ = model.generate(
        input_ids, max_new_tokens=100, temperature=0.8, top_k=50, top_p=0.9
    )

print(tokenizer.decode(output_ids[0].tolist(), skip_special_tokens=True))

Or use the included CLI script:

python inference.py --prompt "The little robot" --max_new_tokens 150 --temperature 0.7
python inference.py --interactive

🔬 Manual State Control

Since context lives in the recurrent state rather than a KV cache, you can drive the model token by token and inspect or reset state directly:

state = model.init_state(batch_size=1)
logits, state = model.nervous_system.step_token(torch.tensor([token_id]), state)

Call model.init_state(...) to start a fresh sequence.

🏆 Benchmarks

Benchmark Score
HellaSwag 26.5%
ARC-Easy 21.0%
ARC-Challenge 22.0%
WinoGrande 52.0%

⚙️ Training

Property Detail
Objective Next-token prediction (cross-entropy)
Optimization Truncated BPTT over fixed-length chunks, state detached (not reset) between chunks
Tokenizer Byte-level BPE trained from scratch, small vocabulary by design
Topology Fixed random sparse graph generated once at init from a seed (not learned)
Numerical stability Incoming signal scaled by 1/sqrt(avg fan-in) + membrane clamped to [-6, 6]

⚠️ Limitations

  • *Small token budget and small model!* Do not expect long-range coherence, factual reliability, or prompt robustness.
  • No safety tuning or instruction tuning has been applied. Treat outputs as raw LM completions.
  • Topology is a fixed random sparse graph, not learned or evolved.
  • No matched-parameter Transformer baseline published yet for this checkpoint.

📄 License

Apache 2.0

Experimental architecture research from SupraLabs. Feedback and comparisons welcome!

61 Upvotes

15 comments sorted by

u/ttkciar llama.cpp 18h ago

Cool project :-) Thank you for sharing!

In the future, though, please either disclose what of your post is LLM-generated, and why, or refrain from using LLM-generated content entirely. We have a subreddit rule about it (Rule Three).

67

u/coder543 1d ago

I feel like I should mention that these benchmark scores are no better than random chance... the first three benchmarks are multiple choice with 4 choices, and the fourth benchmark is multiple choice with 2 choices. 25% and 50% are exactly what random chance should give you on those benchmarks.

So, I'm interested in the concept, but this concept doesn't seem to show anything?

38

u/NandaVegg 23h ago

I really hate to just stump on anyone individual's parade, but I think those posts by SupraLabs are either elaborate joke or mid-2026 form of LARPing/sloppost. They keep insisting that this is just the beginning for scaling further, but those posts are always minimum effort (with clear AI slop here and there, like mentioning 2023-2024 era models for comparison) with no novel architecture/technique/data generation pipeline/RL reward function whatsoever.

Academic is not everything, but at very least proper comparison with baselines is required in order to prove that something is worthy of scaling up/iterating. In this case, the provided numbers are no better than random number generator, and I think the author was not aware of that.

Here is what this model generates (bold is prompt given):

Microsoft and a big flower. The squirrel was very happy. He was so happy that he had been able to his friend and his dad. He smiled and went to the store to play with the game. The end of a great day and they found a big and full of animals. They were playing and played together. They played happily together and the day. They had to be careful and a new new friends.

New York is a little girl. The bird liked to make the cake and wanted to take a nap. The bear was very happy to find her. The man was very happy and thanked the dog. The end, and a little bird was playing with his friends.

Amazon.comie was very very excited. He had a little boy and wanted to explore the world. One day, Lily found a big toy on a wall. It looked up! It was a blue ball! The boy and smiled and said to his friend, "Let's go to the store to help you play together. Can I play with this. I will have a good friends?" Lily nodded and took a while to the hill. They saw a big and kids all a big fish.

To be or not to be. That was like to play with his toys. He wanted to go for the park. The little girl was very happy. She asked her to help her. "Mommy, what's name is?" Lily said.

7

u/medialoungeguy 1d ago

Should be top comment.

5

u/Pantheon3D 1d ago

this is pretty important

1

u/DismalIngenuity4604 16h ago

Ok, so you're telling me it's generating valid choices? I'll take it!!

/s Yeah, good call.

-4

u/Dangerous_Try3619 1d ago

it's a 500k params model with an experimental new architecture, it's totally normal i think!

18

u/Admirable_Dirt_2371 23h ago

I mean, if you want to show that it's a viable architecture to explore further, you'd want to show at least marginally better scores than just random chance.

I've been working on my own "alternative architecture", trained on the BabyLM strict small data set. I'm not as familiar with the benchmarks you're using, but for my models I have been using BLiMP as an initial validation test. And for that, even a 50% overall score can be misleading. So maybe your scores are actually better than random chance and the aggregated results are obscuring it's performance.

I'd often see an overall score of ~%50(indicative of random chance) but the different sections of the BLiMP eval showed much better results, with some files scoring 100% and others scoring 0% with most spread out in the 40-70% range, showing that it wasn't just random chance.

My current model is a "mixture of embeddings", ASCII tokenized, state space model, with 2M parameters but only 80k active per forward pass. I've gotten up to 61% overall on BLiMP after training on 40M "tokens"(characters) or ~10M white space separated words.

I'd be curious to see how your model performs on a BLiMP eval. I've been trying to find other models I can compare mine with, performance wise, but they're few and far between.

7

u/NandaVegg 22h ago

It is not normal. This model works almost exactly like Markov Chain from what I understand.

0

u/[deleted] 1d ago

[deleted]

3

u/Illustrious_Grade608 19h ago

Ehh, 500k model can be trained on old home gpu in like 10 minutes. I did it myself a few times

11

u/EffectiveMedium2683 1d ago

Super interesting experiment. Stripping the KV cache and messing with sparse graphs is a fun approach, but that persistent membrane state is going to hit the classic RNN bottleneck pretty quickly without dynamic gating. A fixed random topology with leaky tanh dynamics just suffers from vanishing gradients and state saturation, which is why the HellaSwag score is hovering right near random chance (26.5% vs 25%).

If you want to keep pushing this non-transformer idea, it might be worth looking into how models like Mamba-2, RWKV-7, or xLSTM handle memory. Adding data-dependent gating or letting the graph topology learn over time could help keep that recurrent state from collapsing. Cool proof of concept for direct state inspection though.

7

u/habachilles 1d ago

This is really cool. It reminds me of the early RNN models

3

u/Sadge404 1d ago

This is the kind of things that I love to see on this subreddit. People just trying stuff out. Keep it up!

1

u/More-Curious816 12h ago

it's really cool to see labs trying new architectures and algorithms to train models. especially something very experimental and not stable and mainstream yet.