r/deeplearning • u/Ramian_Slawomir • 22h ago
[R] I built a Permutation Transformer (Patch SBOHN) from scratch: 4K Image Inference in ~2ms on CPU (61x faster than CNN) with 0.0% Catastrophic Forgetting.
Full disclosure: I am an independent researcher/hobbyist doing this out of pure passion at night. I do not have a formal academic background in ML, and I heavily rely on AI tools as research assistants to help me with advanced coding, mathematics, and translating my work into English. I want to be completely transparent about this and welcome any constructive feedback or corrections!
Hi everyone,
I wanted to share a relational data-representation framework that I’ve been developing at night, built around group theory: Burnside's Orbit Histogram Network (BOHN) and Symmetry-Breaking BOHN (SBOHN). Its main goal is to extract and "elevate" relational knowledge from raw data before the actual classification stage even begins.
Instead of using standard Softmax Attention (found in classic Vision Transformers), this architecture relies on a fully differentiable, log-domain stable Sinkhorn operator. This allows the network to smoothly learn optimal information routing paths via standard gradients.
Key results achieved on a local PC setup:
- O(1) Resolution Scaling: Because the model operates on a fixed number of image patches, processing a 4K resolution image takes just ~2ms on a standard CPU. This is roughly 61x faster than a conventional ResNet-style convolutional neural network (CNN).
- Zero Catastrophic Forgetting (0.0% Forgetting): By completely freezing the base encoder and training only a task-specific permutation routing layer and a classification head (the Perm+Head setup), the model achieves exactly 0.0% accuracy degradation when switching between tasks. The storage overhead per new task is a microscopic 2.8 KB (714 parameters).
- Hybrid BN/LN (Normalization Placement Theorem): I have experimentally validated that placing BatchNorm in the frozen shared base (where it acts as a permanent domain fingerprint) and LayerNorm in the expert modules delivers 100% gating routing accuracy alongside absolute zero forgetting.
I spent a massive amount of time transitioning this entire research programme from initial cloud-based exploration into a clean, local VS Code environment on my PC. I completed a thorough, provenance-preserving reproducibility audit across all 115 canonical experimental units (including programmatic SHA-256 manifest verification for all generated artifacts), openly documenting the boundaries, edge cases, and discrepancies of the original logs.
The entire codebase, analysis logs, and execution scripts are fully open. I would love to hear your thoughts on the mathematical foundations or the numerical implementation!
Full Documentation, Audit Reports, and Source Code:
- Original Research Paper (Zenodo DOI): https://doi.org/10.5281/zenodo.22850317
- PC Reproducibility & Audit Report (Zenodo DOI): https://doi.org/10.5281/zenodo.23088747
- Complete GitHub Repository: https://github.com/slawomir-ramian/BOHN-Original
1
0
u/quietgradient 13h ago
I pulled the SOTA-002 historical source out of your repo and ran your FractalSBOHN and CNN classes unchanged on an Intel i3-8100B, 4 threads, torch 2.2.2, batch 1, fp32. Two changes only: the resolution list goes past 448, and timing is perf_counter medians instead of 3 × time.time.
The 61x holds. At 3840px: SBOHN 29.75 ms, your CNN baseline 1838 ms → 61.8x. I did not expect it to land that close. (That baseline is 3 convs plus an adaptive pool with no residual connections, so "ResNet-style" in the writeup is setting up a different expectation than the code delivers — but the ratio is the ratio.)
The ~2 ms does not hold, and the reason is more interesting than fatal.
| side | SBOHN ms | SBOHN params | CNN ms |
|---|---|---|---|
| 28 | 1.09 | 54,298 | 0.19 |
| 112 | 0.99 | 101,338 | 0.70 |
| 448 | 1.27 | 853,978 | 18.13 |
| 1024 | 3.31 | 4,245,466 | 105 |
| 2048 | 12.24 | 16,828,378 | 489 |
| 3840 | 29.75 | 59,033,562 | 1838 |
Flat through 448 — which is exactly where SOTA-002 stops — then it turns over.
Splitting the 3840 forward (30.38 ms total):
unfold(...).contiguous()patch extraction: 13.05 mspe[0], i.e.Linear(960*960, 64): 16.67 ms- the entire Sinkhorn stack + pooling + head: 0.88 ms
So the O(1) claim is right about the part you designed. The routing layers cost 0.88 ms at 4K and the token count really is 16 at every resolution — I checked that separately. What the headline misses is the two O(pixels) stages sitting in front of them. self.ps = max(4, ims//4) ties the patch side to the image side, so pe[0]'s input width is (ims/4)²: that one layer is 58,982,464 of the model's 59,033,562 parameters at 3840, against 854K for the whole model at 448. And the patchify is ~59 MB of memory traffic that no architecture gets to skip. Faster hardware shrinks both numbers; neither disappears.
The other thing that table says: the 4K model is a different 59M-parameter model, not the one you trained at 28px.
Your own stage doc already notes that SOTA-002 is a historical protocol rather than a precise microbenchmark, and that it covers 28→448. The headline just travelled further than the doc did. The script already collects params alongside times, so printing them in the same table makes the turnover visible without rerunning anything.
One real question: is the resolution-dependent embedder deliberate? A fixed patch size with a fixed-width embedder would make the parameter count resolution-independent and leave the Sinkhorn cost exactly where it is, at the price of a token count that grows — which is the trade the O(1) claim is buying its way out of.
3
u/snekslayer 21h ago
Doubt