Learning State Representations in Complex Systems with Multimodal Data

Solovev, Pavel, Aliev, Vladimir, Ostyakov, Pavel, Sterkin, Gleb, Logacheva, Elizaveta, Troeshestov, Stepan, Suvorov, Roman, Mashikhin, Anton, Khomenko, Oleg, Nikolenko, Sergey I.

arXiv.org Artificial Intelligence 

In order to be able to act in real world scenarios and control a complex system such as an airplane, a car, or an industrial facility, an automated agent needs to process very complex high-dimensional data coming from different domains: video feeds from different cameras, LIDAR sensors on a car, altitude and speed sensors on an airplane, various sensors related to the internal state of the system, and so on. An important problem in this regard would be to map this rich stream of multimodal information into a lower-dimensional space that would compress all modalities into a uniform latent representation (embedding); the agent could then use this embedding to learn or otherwise construct control algorithms. Thus, representation learning lies at the heart of optimal control for complex systems with multimodal unstructured features. Over the last decade, deep neural networks have surpassed other methods in processing nearly all modalities ofhigh-dimensional unstructured data, including images, natural language texts, sounds, and time series. One of the most important properties of neural networks that has made the deep learning revolution possible is their ability to extract meaningful low-dimensional representations of raw unstructured input data. Representation learningwith deep neural networks is a large and well-established area of research [15, 12, 22].

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found