Goto

Collaborating Authors

 Deep Learning




ThinkBig, TeachSmall: DoLanguageModelsDistilOccam'sRazor?

Neural Information Processing Systems

Large language models have recently shown a remarkable ability for few-shot learning, including patterns of algorithmic nature. However, it is still an open question to determine what kind of patterns these models can capture and how manyexamples theyneedintheirprompts.






0d9057d84a9fc37523bf826232ea6820-Paper-Conference.pdf

Neural Information Processing Systems

In the case of coupled skew tent maps, theproposedmethodconsistently outperforms afivelayerDeepNeuralNetwork (DNN) and Long Short Term Memory (LSTM) architecture for unidirectional coupling coefficient values ranging from0.1 to 0.7.