r/MachineLearning - [D] Objective: Masked Language Model vs Autoencoding
Let's say we have a simple "autoencoding transformer" architecture: Now we ask about the properties of Z - the latent representation of the data, after the model is trained. Will Z differ between the two objectives? Will it capture different information? Which loss will preserve more information in Z? Does this have an obvious interpretation?
Dec-23-2019, 16:24:12 GMT
- Technology: