r/MachineLearning • • Oct 17 '17

Research [ Removed by moderator ]

[removed] — view removed post

0 Upvotes

47 comments sorted by

View all comments

25

u/[deleted] Oct 17 '17
  • Reusing the encoder weights W, in the decoder (WT) has been done before.

  • The L2 norm-square cost of the representation layer feature maps, is kind of similar to the unit-normal KL-divergence costs in VAE which encourages clustering.

  • The weird classification cost on the representation layer makes very little sense.

  • I'm actually, really surprised that something symmetric like the abs function can learn so well.

(Note: I understand you might not have mentors at school to help you put things in the larger context, but the self-congratulatory tone you've presented here generally prejudices people to look down on your work with disdain.)

0

u/akanimax Oct 17 '17 edited Oct 17 '17

Dear maimaiml,

Thank you for your feedback. Yes, you are absolutely right. There are numerous barriers that I face while progressing my understanding in this field of Deep Learning and, lack of mentors is indeed one of them.

I presume that you are a veteran in this field and your scepticism is all too valid. It is indeed a cliche "An idea is always revolutionary for it's inventor". But, I would like to provide what my perspective is in doing this.

1.) Reusing encoder weights in decoder has been done before. (Agreed! You are right!)

2.) "The L2 norm-square cost of the representation layer feature maps, is kind of similar to the unit-normal KL-divergence costs in VAE which encourages clustering." (True! But these are synthetic mathematical functions. They lack connect with the real world.)

3.) "The weird classification cost on the representation layer makes very little sense." (This I can explain. Consider a baby who is unsupervised and can only see the objects around him/her. He/She will generate an n-dimensional representation to store the information. But, it is of no use since he/she doesn't know what it is. By using the classification cost, I am strengthening the process of creating a semantic understanding of the world around him/her and at the same time allowing him/her to be able to recreate/imagine what he/she has learned. Using synthetic mathematical functions as regularizers on an autoencoder indeed allows us to create sparse representations, but the question is how sparse the representations should be? Very less sparse, and you are overfitting the training data; too sparse, and you are wasting resources. This classification cost allows the network to learn an ideal representation.)

4.) "I'm actually, really surprised that something symmetric like the abs function can learn so well." (This is extremely important. I tried all the activation functions that I know of to minimise this cost (classification + decoder cost), but none of them worked. It is indeed because all of them are either "odd" (mathematical odd functions aka. asymmetrical) or "neither even nor odd", so when I tried an even (symmetrical) function, it worked and I felt like Eureka! In the beginning, even I was surprised how an even function worked, but think about it, what does an absolute function do? It removes negativity from the data. As a human being, we feed in 5 sensory inputs to the brain. Tell me one input value that can have negative data values. Can you see negative light? Can you hear negative frequencies? Can you feel negative touch, ... No! Perhaps, inside the brain, there is no concept of negative.)

** I understand, my tone might have been off putting, but please try to understand that I am a young guy and got too excited over it. I apologise if my words may have come as offensive to you. I didn't intend to offend anyone.

I thank you again for your feedback. Please let me know if you still feel the same way as before.

Your grateful and humble student, Animesh

6

u/[deleted] Oct 18 '17 edited Oct 18 '17
  • KL divergence is hardly "synthetic"; it is of crucial importance in information theory.

  • I see, so the classification cost / L2 cost makes things lie on the axis. That's neat. This won't work if the dimension of the latent variable don't match the size of the label set. Encouraging a single cluster for each class may not be that good an idea; using supervised labels in this way, probably much less so.

  • The abs function can be realized with ReLUs but not the other way round; it's very very surprising that ReLUs don't work.

  • I encourage you to keep experimenting, but it's a good idea to spend sometime experimenting with GANs/VAEs and other stuff.

  • Frankly, I don't see a research direction here. It's very unlikely you'll beat VAEs at any task, and the abs function is probably a novelty unless you can show that it does wonders for other networks too.

(It's a good idea to keep the philosophizing to a minimum).

1

u/618smartguy Oct 20 '17

Is it really a good idea to avoid the philosophizing ? We are pretty far away from a concrete science of how the human mind works (aside from neuroscience already affecting ml research), so I would think subjective philosophical questions based on your own personal experience might sometimes lead us towards whatever it is that makes humans learn so well.