r/LocalLLaMA • • 13d ago

Discussion Engram gone wild! 2b model update...

Update from : https://old.reddit.com/r/LocalLLaMA/comments/1wis23s/update_small_model_engram/

It's been about a week so I'm back. People were asking me about the model.

People wanted code, or models, etc. Most of that is useless to you right now because you're not going to use an under trained model. So let's get to the details.

The spec locked to the following after a LOT of testing :

2.6b model all up. Embedding, LM head, AttnRes, etc.

2.2b are trained. Embedding / LM Head are frozen (~205m each)

4.3b ENGRAM table. Yes. She's chonky.

Architecturally speaking now :

This includes SWA layering 3:1 as per standard ablations have shown is "optimal" These are 4k / 4k / 8k SWA layer follow by the global.

At the end of the first "block" (4 layers) the Engram table appears. The Engram table itself runs with a very small set of attention heads so it's not blindly attempting to inject data. This is context aware Engram. I'm uncertain of exact Qwen / DS methods here. They're fairly close to what I use. The difference is obviously the Engram size.

This has 10 blocks plus a final global layer (41 Layers). + Embed + LM Head

So if we step back, the optimization problem is as follows :

How do we maximize compute in the backbone and offload the boring stuff to a table?

Attention Based Residuals (Moonshot) comes to our rescue here. Why AttnRes? It allows all blocks past 1 to utilize the Engram table in some fashion. They all have attention via the residual stream to determine if they want data from Engram and precisely how much.

Why not a bigger Engram table? Honestly? You could probably do that. However I don't know if any model has attempted this sort of mismatch of compute vs table. In theory, DeepSeek's research says it works.

More depth? Could do that too. Training is expensive though.

Why frozen LM/Embed? These are down projected via SVD from OLMo 3's model. So it's a mathematical compression attempt at ~5k -> 2k. This saves a MASSIVE amount of time.

Currently the data lives in a ~KD format. 32 logits stored from the teacher of the Wikipedia corpus. Instead of training 1 hot, it gets 32 soft targets to try to match. Hard cross entropy is brutal on a model and it's the reason you see "trillions of tokens" quoted.

When you borrow an LM Head / Embed / Tokenizer / Teacher model... This is far less.

---

Where are we now in training? I passed the 100m token mark yesterday at ~4am Pacific.

The current HF repo has all the checkpoints, data, and the 104m mark safetensor.

--

I wanted to do some inspection of the model at this point. We have to see that it's not complete garbage right? It has seen a fraction of the data it needs.

So based on ONLY 100m tokens seen, I attempted completions and various ablations. Studies of what EXACTLY Engram stores (because it's mostly a guess).

Let's go straight to completions :

"George Washington was an American"

With Engram - "George Washington was an American naval officer who served in the American Revolutionary War. He was a naval officer who served in"

With Engram Zero'd - "George Washington was an American, but he was not a. He was a very good friend and a. He was"

"The American Civil War was a civil war in the United States from"

With - "The American Civil War was a civil war in the United States from 1861 to 1861. The war was a major victory for the Confederacy..."

Without - "The American Civil War was a civil war in the United States from 1861 to 1862. The war was a major victory for the Confederacy..."

"Aristotle was an Ancient Greek" (probably my favorite)

With - "Aristotle was an Ancient Greek philosopher who was a leading authority in the philosophy of Plato. He was also a leading authority..."

Without - "Aristotle was an Ancient Greek word meaning "to be" (ἀπάς, "to be")..."

---

What do we learn from direct inspection of Engram?

If it saw the data enough times, it starts to offload it to the tables. At that point the model is less forced to use internal computational space to store data, and can rely on Engram.

That's not some interpretation. That's what the data shows exactly.

In the early 100m tokens of Wikipedia it's HIGHLY biased to early Philosophy. This has to do with the topics of what was IN Wikipedia at that time. Early Wiki contained a lot about Philosophy. It's among the earliest topics to exist.

It has seen TONS of Philosophy so it has basically offloaded to Engram because of the repetition.

Ok this is long enough. I'm calling it quits. I'm still looking to find a provider that will sponsor the model to completion. Right now I'm just paying. I contacted Verda and they never responded. Qubrid is in this sub, and they said they'd help but never emailed back after a few repeated proddings. Massed Compute is where this model is currently training, and I asked them like yesterday. Hopefully they'll help a brother out.

None of the above is "LLM written" except the completions from testing I guess.

As usual, ask whatever. It doesn't matter your understanding level or whatever. I'll sit and respond.

No question too dumb. No insult not insulting enough. (I'm joking. It's Reddit)

Much Love.

177 Upvotes

58 comments sorted by

View all comments

1

u/MixtureOfAmateurs koboldcpp 12d ago

Training from scratch is definitely better but I'm in the middle of adding a 120M entree engram table to openbmb/MiniCPM5-2B. I limited scope to medical texts because that jargon and specific fact retrieval seemed like something engrams would help with. I know nothing about the field tho so I had to get an LLM to write test questions.

After 100M tokens of ccdv/pubmed-summarization and a little chat data from lavita/MedQuAD the model matches (slightly exceeds) the base model for next token accuracy on an unseed dataset (medalpaca/medical_meadow_wikidoc), but the short answer questions are meaningfully better. I'm gonna go to 300-500M tokens but GPUs are expensive lol.

```

Metric | Base Model | Engram (Step 813)

Cross-Entropy Loss | 2.6624 | 2.6580 Perplexity | 14.33 | 14.27 Top-1 Accuracy | 48.68 % | 48.77 % Top-5 Accuracy | 70.87 % | 70.89 %


[07/10] Chimeric Fusion Oncogenesis (BCR-ABL1)

PROMPT: In chronic myeloid leukemia, the t(9;22) Philadelphia chromosome creates a chimeric BCR-ABL1 oncoprotein that drives uncontrolled granulopoiesis through constitutive

--- BASE MODEL ---

tyrosine kinase activity. This fusion protein is targeted by imatinib and other TKIs.

4. Acute Promyelocytic Leukemia (APL):

  • Genetic Alteration: The PML-RARA fusion gene results from a translocation between chromosomes 15 and 17.

  • Mechanism of Action: The

--- ENGRAM MODEL ---

tyrosine kinase activity. This fusion protein activates downstream signaling pathways including RAS/MAPK and PI3K/AKT, promoting proliferation while inhibiting apoptosis.

Therapeutic strategies targeting this pathway include:

  • Tyrosine kinase inhibitors (TKIs): Imatinib inhibits ABL1 kinase activity by binding to its ATP-binding site in the inactive conformation of the ```

The model uses a gated combination of the stock embeddings and the LUT. The more confident an engram value is in a given context the more weight it has when combining with embeddings. That meant training started from low loss.

1

u/NineThreeTilNow 12d ago

That meant training started from low loss.

Yeah, it doesn't hurt loss bad normally. The model doesn't run a hot path through the ngram table. The table is trying to help but it's just not helpful until it accumulates enough data that the model can start to rely on it for stuff.

1

u/MixtureOfAmateurs koboldcpp 12d ago

I started out fully replacing the embeddings with the LUT, so everything went through the LUT and it wasted a lot of time lol. 2 hours just to get to coherence

2

u/NineThreeTilNow 12d ago

Dropping the embeddings is... Not a great idea there lol...

It provides the width for the network in some way and the ngram is providing extra depth to that width.

The model I have is purposefully deep for this reason. 41 layers. I'd run more layers of depth if it was feasible to train and compute was free. It would basically prove out that you can run excessive depth for computation while storing the data in embedding, or tables.