Goto

Collaborating Authors

 Country


Cycle-ContrastforSelf-SupervisedVideo RepresentationLearning

Neural Information Processing Systems

These methods giveeffectiverepresentations and decent results ofdownstream tasks, however we suggest that utilizing other nature characteristics of video can lead to different yet representative video representations.





YouNeverStopDancing: Non-freezingDance GenerationviaBank-constrainedManifoldProjection

Neural Information Processing Systems

One of the most overlooked challenges in dance generation is that the autoregressiveframeworks are prone tofreezing motions due tonoiseaccumulation. Inthispaper,wepresent twomodules thatcanbeplugged intotheexisting models to enable them to generate non-freezing and high fidelity dances.



40bb79c081828bebdc39d65a82367246-Supplemental-Conference.pdf

Neural Information Processing Systems

Table1: Linearnetwork Layer# Name Layer Inshape Outshape 1 Flatten() (3,32,32) 3072 2 fc1 nn.Linear(3072, 200) 3072 200 3 fc2 nn.Linear(200, 1) 200 1 Fully-connected Network We conduct further experiments on several different fully-connected networks with 4 hidden layers with various activation functions. Our subset is smaller because of the computation limitation when calculating the Gram matrix. Experiments show that the properties along GD trajectory(e.g. We consider simple linear networks, fully-connected networks, convolutional networks in this appendix. The following Figure 4 illustrates the positive correlation between thesharpness andtheA-norm, andtherelationship between theloss D(t) 2 and R(t) 2 alongthetrajectory.


40bb79c081828bebdc39d65a82367246-Paper-Conference.pdf

Neural Information Processing Systems

Recent findings demonstrate that modern neural networks trained by full-batch gradient descent typically enter a regime called Edge of Stability (EOS).


SUPER-ADAM: FasterandUniversalFrameworkof AdaptiveGradients

Neural Information Processing Systems

Although multiple adaptivegradient methods were recently studied, theymainly focus oneither empirical ortheoretical aspects and also only work for specific problems by using some specific adaptive learning rates.


SUPER-ADAM: FasterandUniversalFrameworkof AdaptiveGradients

Neural Information Processing Systems

Although multiple adaptivegradient methods were recently studied, theymainly focus oneither empirical ortheoretical aspects and also only work for specific problems by using some specific adaptive learning rates.