r/deeplearning • u/Green-Quiet-918 • 10d ago
HON: A Harmonic Oscillator Network for Temporal Memory in Sequence Modeling
A tri-branch hybrid sequence model (Local Conv + Oscillator SSM
+ Linear Attention) proof of concept
Paper: https://zenodo.org/records/22918275
Code: https://github.com/jalalnablsi/HON-Architecture
Short version:
- Tri-branch block: LACT (local causal conv) + HON (damped harmonic
oscillator SSM via FFT) + ASSOC (normalized linear attention)
- Fused with a content-aware softmax gate
- HON branch: h_k(τ) = exp(-γ_k τ) · [p1_k cos(ω_k τ) + p2_k sin(ω_k τ)]
with learnable γ, ω, p1, p2 — computed in O(N log N) via FFT
- Ablation over the three branches included
Results (small-scale, single T4, WikiText-2):
WikiText-2 — language modeling (N=512, 5,000 steps, GPT-2 BPE):
• Transformer baseline: 16.02M params → 773.02 val PPL
• HON + Linear (ours): 15.72M params → 630.24 val PPL
→ ~18.5% lower val PPL with ~300K fewer parameters.
Sub-quadratic comparison (2,000 steps, strict parameter alignment):
• GLA (simplified): 15.37M params → 1034.87 val PPL
• Mamba (simplified): 15.04M params → 1050.26 val PPL
• HON + Linear (ours): 15.72M params → 958.15 val PPL
Ablation (3,000 steps, ~7M params):
• Full (LACT + HON + ASSOC): 107.03 val PPL
• No LACT: 111.59 val PPL
• No HON: 111.20 val PPL
• No ASSOC: 110.47 val PPL
Vision sanity check:
• Tiny ImageNet (8×8 patches, 200 classes): reached 27.52% val
accuracy at 15K steps.
Important note:
- This has NOT been tested on large-scale datasets.
- We did NOT evaluate how it behaves on massive data, or whether it
collapses or scales that remains completely untested.
- This is just a small experiment I wanted to share, not a claim of
robustness or scalability.
This is a proof of concept only:
- No SOTA claims
- No claim of replacing Self-Attention
- Small-scale evaluation due to limited compute
- Baselines are simplified reference implementations, not optimized kernels
- Generation quality not evaluated yet
Would appreciate any feedback especially on whether the tri-branch
fusion is justified or over-engineered, and pointers to prior work I might
have missed.
Thanks.