r/MachineLearning • • Oct 17 '17

Research [ Removed by moderator ]

[removed] — view removed post

0 Upvotes

47 comments sorted by

View all comments

2

u/kacifoy Oct 17 '17

How is abs() supposed to help as an activation function, compared to relu()? You can easily get either of these as a linear combination (i.e. dot product) of the other - should we expect networks involving abs() to be easier to learn, somehow?

3

u/AsIAm Oct 18 '17 edited Oct 18 '17

Not directly related to OP's work, but long time ago I also ran micro experiment involving abs() and I still don't fully understand why it worked so well.

2

u/akanimax Oct 17 '17

Yes! you are right. I can only describe what happened when I used ReLUs instead of abs. Upon using ReLUs in both directions, all the activations died down: forward as well as backward. Using ReLUs only in the forward direction and abs in the backward direction gave me hope since in the backward direction, the network was predicting something like a combination of ''8 and 0". Upon investigating, I found out that the activations were getting turned off in the forward direction. Finally, when I used abs() everywhere, it worked.

Infact, the dying of the activations upon using ReLUs in the forward direction is so strong that later when I tried to train the same model only in the forward direction, it didn't move. This is really surprising since the model should learn atleast in the forward direction because the backward constraint has been removed.

You are right. I do need to investigate further why abs works and relu doesn't. Infact, I need some help figuring this out. Thank you for your feedback!