[D] Statistical language models are not good for NLU?
We are interviewing Walid Saba on *Friday* for Machine Learning Street Talk show (with Yannic Kilcher). He has just written an article, but written many before claiming that deep learning and memorisation / statistical approaches are completely flawed for NLU. He calls these approaches "BERTology" which I think it a funny name! He points out the "the missing text phenomenon" as the biggest issue i.e. "the corner table wants a beer" -- "the _person_ at the corner table wants a beer" ... and provides many other similar examples. He makes a "proof" for this by equating ML to "compressability" and NLU to "expansion" which is intuitive, although I would argue ML could just as easily be used to decompress, think a basic generative model to learn to decompress something.
Oct-22-2020, 07:00:52 GMT