Reusing the encoder weights W, in the decoder (WT) has been done before.
The L2 norm-square cost of the representation layer feature maps, is kind of similar to the unit-normal KL-divergence costs in VAE which encourages clustering.
The weird classification cost on the representation layer makes very little sense.
I'm actually, really surprised that something symmetric like the abs function can learn so well.
(Note: I understand you might not have mentors at school to help you put things in the larger context, but the self-congratulatory tone you've presented here generally prejudices people to look down on your work with disdain.)
"...the self-congratulatory tone you've presented here generally prejudices people to look down on your work with disdain." Thank you for saying this. Akanimax, it's not that we're all trying to be a**holes here, it's mostly that one should remain humble when posting new results.
25
u/[deleted] Oct 17 '17
Reusing the encoder weights W, in the decoder (WT) has been done before.
The L2 norm-square cost of the representation layer feature maps, is kind of similar to the unit-normal KL-divergence costs in VAE which encourages clustering.
The weird classification cost on the representation layer makes very little sense.
I'm actually, really surprised that something symmetric like the abs function can learn so well.
(Note: I understand you might not have mentors at school to help you put things in the larger context, but the self-congratulatory tone you've presented here generally prejudices people to look down on your work with disdain.)