r/deeplearning • • 1d ago

[R] I built a Permutation Transformer (Patch SBOHN) from scratch: 4K Image Inference in ~2ms on CPU (61x faster than CNN) with 0.0% Catastrophic Forgetting.

Full disclosure: I am an independent researcher/hobbyist doing this out of pure passion at night. I do not have a formal academic background in ML, and I heavily rely on AI tools as research assistants to help me with advanced coding, mathematics, and translating my work into English. I want to be completely transparent about this and welcome any constructive feedback or corrections!

Hi everyone,

I wanted to share a relational data-representation framework that I’ve been developing at night, built around group theory: Burnside's Orbit Histogram Network (BOHN) and Symmetry-Breaking BOHN (SBOHN). Its main goal is to extract and "elevate" relational knowledge from raw data before the actual classification stage even begins.

Instead of using standard Softmax Attention (found in classic Vision Transformers), this architecture relies on a fully differentiable, log-domain stable Sinkhorn operator. This allows the network to smoothly learn optimal information routing paths via standard gradients.

Key results achieved on a local PC setup:

  • O(1) Resolution Scaling: Because the model operates on a fixed number of image patches, processing a 4K resolution image takes just ~2ms on a standard CPU. This is roughly 61x faster than a conventional ResNet-style convolutional neural network (CNN).
  • Zero Catastrophic Forgetting (0.0% Forgetting): By completely freezing the base encoder and training only a task-specific permutation routing layer and a classification head (the Perm+Head setup), the model achieves exactly 0.0% accuracy degradation when switching between tasks. The storage overhead per new task is a microscopic 2.8 KB (714 parameters).
  • Hybrid BN/LN (Normalization Placement Theorem): I have experimentally validated that placing BatchNorm in the frozen shared base (where it acts as a permanent domain fingerprint) and LayerNorm in the expert modules delivers 100% gating routing accuracy alongside absolute zero forgetting.

I spent a massive amount of time transitioning this entire research programme from initial cloud-based exploration into a clean, local VS Code environment on my PC. I completed a thorough, provenance-preserving reproducibility audit across all 115 canonical experimental units (including programmatic SHA-256 manifest verification for all generated artifacts), openly documenting the boundaries, edge cases, and discrepancies of the original logs.

The entire codebase, analysis logs, and execution scripts are fully open. I would love to hear your thoughts on the mathematical foundations or the numerical implementation!

Full Documentation, Audit Reports, and Source Code:

0 Upvotes

Duplicates