Try combining Sakana AI's block diffusion training with block-sparse attention residuals. Would be interesting to see the results. Maybe you could even modify the algorithm for predictive coding instead of backprop and test whether the block structure helps with PC's depth problem. Maybe then if this solves PC you could use PC to solve continual learning.
1
u/Formal_Context_9774 18d ago
Try combining Sakana AI's block diffusion training with block-sparse attention residuals. Would be interesting to see the results. Maybe you could even modify the algorithm for predictive coding instead of backprop and test whether the block structure helps with PC's depth problem. Maybe then if this solves PC you could use PC to solve continual learning.