[D] thoughts on a few recent papers that could be useful in neural network tabular regression

#artificialintelligence 

Mish: A Self Regularized Non-Monotonic Neural Activation Function - Shows improvement over swish in some cases. Gradient Centralization: A New Optimization Technique for Deep Neural Networks - Basically just makes the mean of all weights excluding output layer to 0 E-Swish - [1801.07145] E-swish: Adjusting Activations to Different Network Depths - Adds a term to swish and shows that in some cases a parameterized swish provides better results. On the Variance of the Adaptive Learning Rate and Beyond - Uses smaller learning rates in first few epochs essentially just to get enough samples so that the variance of adaptive LR doesn't explode and causes detrimental learn rates in the first few steps. Lookahead Optimizer: k steps forward, 1 step back - maintains two optimizers/weights, one with a larger and more reactive learn rate, one with a smaller more stable learn rate.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found