Appendices ADetailsofOnlineFew-ShotLearningExperiments
–Neural Information Processing Systems
A.5 ArchitectureDetails The LSTM baseline and the controller in NTM have a hidden size of 400, followed by a 5-way classifier. Weuse SGD for adaptation instead of Adam in meta-testing, which resulted in significantly better results for our setting where trajectories are short, likely due to the lackofmomentum. We train the model with an outer loop learning rate of10 3 using Adam and an inner loop learning rate of10 2 using gradient descent. Finally, a 1 1 conv layer is used to generate logits at a pixellevel. This underlines that label injection can be used for avast variety of standardnetworks.
Neural Information Processing Systems
Feb-9-2026, 12:16:36 GMT
- Technology: