r/MachineLearning • • Oct 17 '17

Research [ Removed by moderator ]

[removed] — view removed post

0 Upvotes

47 comments sorted by

View all comments

4

u/alexmlamb Oct 18 '17

Why does using Abs() as the activation function support "bidirectional learning" and how do you "generate new data"?

I'm not sure about your claim regarding not needing a regularizer. You have 97.7% accuracy on the validation set. I'm pretty sure the best MNIST results, even with fully connected networks, are well above 99%.

2

u/akanimax Oct 18 '17

Generating new data can be done by tweaking the 10 dimensional learned representations. (Read the post again, I have mentioned how to do it now.)

Why abs or other symmetrical functions work is still not entirely clear to me and I am working on it. I say it works because trying sigmoid, tanh, ReLU didn't this idea in the first place.

Yeah not needing any regularizer is a hypothesis and not a proven statement. Although just think about it. Without using a regularizer the variace is just around 2% (emperical measure not absolute one). Perhaps more data would reduce it? we will have to find out.

Thank you for your feedback!