Goto

Collaborating Authors

 Technology


Optimistic Rates for Multi-Task Representation Learning

Neural Information Processing Systems

We study the problem of transfer learning via Multi-Task Representation Learning (MTRL), wherein multiple source tasks are used to learn a good common representation, and a predictor is trained on top of it for the target task. Under standard regularity assumptions on the loss function and task diversity, we provide new statistical rates on the excess risk of the target task, which demonstrate the benefit of representation learning. Importantly, our rates are optimistic, i.e., they interpolate between the standard O(m 1/2)rate and the fast O(m 1)rate, depending on the difficulty of the learning task, where m is the number of samples for the target task. Besides the main result, we make several new contributions, including giving optimistic rates for excess risk of source tasks (Multi-Task Learning (MTL)), a local Rademacher complexity theorem for MTRL and MTL, as well as a chain rule for local Rademacher complexity for composite predictor classes.


Optimistic Rates for Multi-Task Representation Learning

Neural Information Processing Systems

We study the problem of transfer learning via Multi-Task Representation Learning (MTRL), wherein multiple source tasks are used to learn a good common representation, and a predictor is trained on top of it for the target task. Under standard regularity assumptions on the loss function and task diversity, we provide new statistical rates on the excess risk of the target task, which demonstrate the benefit of representation learning. Importantly, our rates are optimistic, i.e., they interpolate between the standard O(m 1/2)rate and the fast O(m 1)rate, depending on the difficulty of the learning task, where m is the number of samples for the target task. Besides the main result, we make several new contributions, including giving optimistic rates for excess risk of source tasks (Multi-Task Learning (MTL)), a local Rademacher complexity theorem for MTRL and MTL, as well as a chain rule for local Rademacher complexity for composite predictor classes.


Supplementary: Non-Local Latent Relation Distillation for Self-Adaptive 3DHuman Pose Estimation

Neural Information Processing Systems

The raw video frames are forwarded through a person-detector [15] to obtain the person-focused image sequences. Note that, the detector pruned video sequences may not have a smooth pixel transition. However, it retains the smooth pose transition at the view-variant root-relative system. In our work, the shared latent pose can be seen as a parametric form to represent plausible 3D poses. And, the image-to-latent model is trained to regress the latent pose parameters with latent being an intermediate 3D pose representation.



Zodiac Killer may be tied to Black Dahlia case after 'code cracked,' new suspect emerges

FOX News

This material may not be published, broadcast, rewritten, or redistributed. Quotes displayed in real-time or delayed by at least 15 minutes. Market data provided by Factset . Powered and implemented by FactSet Digital Solutions . Mutual Fund and ETF data provided by LSEG .


The New Masculinity of "DTF St. Louis"

The New Yorker

The show exists in a strange world where men repeatedly confess their love for each other. Does it make them better people? Much ink has been spilled, and countless TikToks recorded, in an effort to explain the female fervor unleashed by the series " Heated Rivalry ." I, a thirty-eight-year-old woman who owns a T-shirt that bears the logo of Shane Hollander's Montreal Metros and another that celebrates Ilya Rozanov's Boston Raiders (Valentine's Day gifts, it should be said, from my indulgent husband), don't find its appeal so mystifying. Two gorgeous young men, as elegantly muscled as Myron's discus thrower, have ecstatically unbridled, mutually satisfying sex to a soundtrack designed to tickle elder millennials' nostalgia-pleasure centers, all while falling in the kind of soul-sustaining love that most of us can only dream of.





Details

Neural Information Processing Systems

Here we derive Equation 8 for 0 and out = > 0. Since ESN(µ, 2,0) = NR(µ,), we can obtain Equation 4 for ID activation by specializing the result to =0 . We begin with a useful lemma. Let X ESN(0, 2,) and let a b 0, 0 c d. Then P(a X b)= (1+) h The result for P(c X d) follows analogously. For the reader's convenience, we summarize in detail a few common techniques for defining OOD scores that measure the degree of ID-ness on the given sample. All the methods derive the score post hoc on neural networks trained with in-distribution data only.