r/MachineLearning - [D] Objective: Masked Language Model vs Autoencoding

#artificialintelligence 

Let's say we have a simple "autoencoding transformer" architecture: Now we ask about the properties of Z - the latent representation of the data, after the model is trained. Will Z differ between the two objectives? Will it capture different information? Which loss will preserve more information in Z? Does this have an obvious interpretation?