Goto

Collaborating Authors

 Technology



Supplementaryto"DSelect-k: Differentiable SelectionintheMixtureofExpertswithApplications toMulti-TaskLearning "

Neural Information Processing Systems

MTL: InMTL, deep learning-based architectures that perform soft-parameter sharing, i.e., share model parameters partially, are proving to be effective at exploiting both the commonalities and differences among tasks [6]. Ourwork is also related to [5] who introduced "routers" (similar to gates) that can choose which layers or components of layers to activate per-task. The routers in the latter work are not differentiable and requirereinforcementlearning. To construct ฮฑ, there are two cases to consider: (i)s = k and (ii) s < k. If s = k, then set ฮฑi = log(w ti) for i [k]. Our base case is fort = 1.



Meta-Query-Net: ResolvingPurity-InformativenessDilemmain Open-setActiveLearning (SupplementaryMaterial) ACompleteProofofTheorem4.1

Neural Information Processing Systems

Let g[1](zx) be g(zx) and W[1] be W for notation simplicity. Consider each dimension's scalar output ofg(zx), and it is denoted asg p (zx) where p is an index of the output dimension. For each AL round, a target modelฮ˜is trained via stochastic gradient descent(SGD) using IN examples in the labeled setSL (Lines 3-5). The initial learning rate of0.1 is decayed by a factor of 0.1 at 50% and 75% of the total training iterations. Owing to the ability to find the best balance between purity and informativeness, MQ-Net achieves the highest accuracy on every AL round.