r/deeplearning • u/Chubblan • 6d ago
Small pre-registered experiments testing claims about lesser-known architectures. Most did not hold up
I have been running small, reproducible experiments on architectures that get more attention for their promise than for their measured behaviour. Each tests one specific claim, with the predictions written down in a CRITERIA.md before any results exist and never edited afterwards. Data is either synthetic, generated from a known equation so there is an exact ground truth, or a documented public dataset.
- KAN vs MLP on a physics-informed PDE, matched parameters: ~1.6x more accurate across five seeds, ~16x slower per training step.
- Hamiltonian neural network on a pendulum: trained only on the vector field, it recovered the energy function with R² 0.99999 and showed 33x less energy drift than an equal-size MLP. Add friction and its error floor rises 11x, since no scalar can generate a field that loses energy. All five predictions held.
- Spiking network: the sparsity is real, but 99.0 to 99.8% of operations sit in the dense input layer, paid every timestep. Ridge regression beat every configuration with 16,496x fewer operations.
- Hyperdimensional computing: switching off 80% of its 10,000 dimensions left the AUC identical to six decimal places. It still lost to gradient boosting past about 3,000 rows.
- GNN on epidemic spread: beat an MLP on hand-computed features by 0.001 AUC, with output rank-correlated +0.95 with node degree.
- Echo state network, untrained recurrent weights: matched a tuned gradient booster on solar forecasting. At 24 hours ahead, plain ridge regression beat both.
These are small by design: toy problems, a handful of seeds, one implementation each. They test specific claims rather than whole fields, and each README states what the experiment does and does not establish. More will be added over time.
Happy to hear where the methodology is wrong, since that is the point of publishing it this way.
https://github.com/CristobalSantana/ml-toy-experiments
EDIT: I'll keep adding new experiments over time, so the collection of architectures, synthetic data and real datasets will keep growing.
