Appendices ADetailsofOnlineFew-ShotLearningExperiments

Neural Information Processing Systems 

A.5 ArchitectureDetails The LSTM baseline and the controller in NTM have a hidden size of 400, followed by a 5-way classifier. Weuse SGD for adaptation instead of Adam in meta-testing, which resulted in significantly better results for our setting where trajectories are short, likely due to the lackofmomentum. We train the model with an outer loop learning rate of10 3 using Adam and an inner loop learning rate of10 2 using gradient descent. Finally, a 1 1 conv layer is used to generate logits at a pixellevel. This underlines that label injection can be used for avast variety of standardnetworks.

Similar Docs  Excel Report  more

TitleSimilaritySource
None found