Deep Learning
Renesas Accelerates Deep Learning Development for ADAS and Automated Driving Applications
Düsseldorf, September 21, 2021 ― Renesas Electronics Corporation (TSE:6723), a premier supplier of advanced semiconductor solutions, today unveiled the R-Car Software Development Kit (SDK), a complete software platform in a single package that enables quicker and easier software development and validation for smart camera and automated driving applications used in passenger, commercial, and off-road vehicles. "Software development and delivery has been a significant pain point for automotive system developers, involving resource-intensive customized packaging and full installations that typically take several days to complete," said Naoki Yoshida, Vice President, Automotive Digital Products Marketing Division at Renesas. "To alleviate these headaches when it comes to deep learning for automotive systems, Renesas is reinventing the developer experience, offering this new single package, multi-OS software platform that is easy for customers to access, learn, use, and install, enabling customers to quick start their deep learning development." Re-inventing SW development for Automotive Applications Automakers are increasingly turning to deep learning as they look for new ways to enable smart camera applications and automated driving systems for next-generation vehicles. However, most deep learning solutions available today are built on consumer or server applications, which do not operate under the same stringent constraints for functional safety, real-time responsiveness, and low power consumption.
78 AI Companies Around The World That Are Unicorns Today
AI is at the heart of digital disruption and on its way to becoming one of the biggest game-changers in the next few years. Early adopters of AI are reaping significant benefits and have differentiated themselves from the rest. As a result, the AI sector is garnering the attention of numerous investors globally, increasing the number of AI unicorns in just a few years. In India itself, as many as 11 startups earned unicorn tags during the black swan year 2020. This article lists all the AI companies that have reached a valuation of $1 billion or more. A technology platform company, Argo AI, is creating integrated self-driving systems. These are manufactured at scale for safe and reliable deployment in ride-sharing and goods delivery services. Along with Ford and Lyft, Argo AI is planning to launch a self-driving ride-hailing service in the US.
Multivariate Anomaly Detection based on Prediction Intervals Constructed using Deep Learning
It has been shown that deep learning models can under certain circumstances outperform traditional statistical methods at forecasting. Furthermore, various techniques have been developed for quantifying the forecast uncertainty (prediction intervals). In this paper, we utilize prediction intervals constructed with the aid of artificial neural networks to detect anomalies in the multivariate setting. Challenges with existing deep learning-based anomaly detection approaches include $(i)$ large sets of parameters that may be computationally intensive to tune, $(ii)$ returning too many false positives rendering the techniques impractical for use, $(iii)$ requiring labeled datasets for training which are often not prevalent in real life. Our approach overcomes these challenges. We benchmark our approach against the oft-preferred well-established statistical models. We focus on three deep learning architectures, namely, cascaded neural networks, reservoir computing and long short-term memory recurrent neural networks. Our finding is deep learning outperforms (or at the very least is competitive to) the latter.
Combining Image Features and Patient Metadata to Enhance Transfer Learning
In this work, we compare the performance of six state-of-the-art deep neural networks in classification tasks when using only image features, to when these are combined with patient metadata. We utilise transfer learning from networks pretrained on ImageNet to extract image features from the ISIC HAM10000 dataset prior to classification. Using several classification performance metrics, we evaluate the effects of including metadata with the image features. Furthermore, we repeat our experiments with data augmentation. Our results show an overall enhancement in performance of each network as assessed by all metrics, only noting degradation in a vgg16 architecture. Our results indicate that this performance enhancement may be a general property of deep networks and should be explored in other areas. Moreover, these improvements come at a negligible additional cost in computation time, and therefore are a practical method for other applications.
Social Recommendation with Self-Supervised Metagraph Informax Network
Long, Xiaoling, Huang, Chao, Xu, Yong, Xu, Huance, Dai, Peng, Xia, Lianghao, Bo, Liefeng
In recent years, researchers attempt to utilize online social information to alleviate data sparsity for collaborative filtering, based on the rationale that social networks offers the insights to understand the behavioral patterns. However, due to the overlook of inter-dependent knowledge across items (e.g., categories of products), existing social recommender systems are insufficient to distill the heterogeneous collaborative signals from both user and item sides. In this work, we propose a Self-Supervised Metagraph Infor-max Network (SMIN) which investigates the potential of jointly incorporating social- and knowledge-aware relational structures into the user preference representation for recommendation. To model relation heterogeneity, we design a metapath-guided heterogeneous graph neural network to aggregate feature embeddings from different types of meta-relations across users and items, em-powering SMIN to maintain dedicated representations for multi-faceted user- and item-wise dependencies. Additionally, to inject high-order collaborative signals, we generalize the mutual information learning paradigm under the self-supervised graph-based collaborative filtering. This endows the expressive modeling of user-item interactive patterns, by exploring global-level collaborative relations and underlying isomorphic transformation property of graph topology. Experimental results on several real-world datasets demonstrate the effectiveness of our SMIN model over various state-of-the-art recommendation methods. We release our source code at https://github.com/SocialRecsys/SMIN.
Field Extraction from Forms with Unlabeled Data
Gao, Mingfei, Chen, Zeyuan, Naik, Nikhil, Hashimoto, Kazuma, Xiong, Caiming, Xu, Ran
We propose a novel framework to conduct field extraction from forms with unlabeled data. To bootstrap the training process, we develop a rule-based method for mining noisy pseudo-labels from unlabeled forms. Using the supervisory signal from the pseudo-labels, we extract a discriminative token representation from a transformer-based model by modeling the interaction between text in the form. To prevent the model from overfitting to label noise, we introduce a refinement module based on a progressive pseudo-label ensemble. Experimental results demonstrate the effectiveness of our framework.
How Can AI Recognize Pain and Express Empathy
Cao, Siqi, Fu, Di, Yang, Xu, Barros, Pablo, Wermter, Stefan, Liu, Xun, Wu, Haiyan
Sensory and emotional experiences such as pain and empathy are relevant to mental and physical health. The current drive for automated pain recognition is motivated by a growing number of healthcare requirements and demands for social interaction make it increasingly essential. Despite being a trending area, they have not been explored in great detail. Over the past decades, behavioral science and neuroscience have uncovered mechanisms that explain the manifestations of pain. Recently, also artificial intelligence research has allowed empathic machine learning methods to be approachable. Generally, the purpose of this paper is to review the current developments for computational pain recognition and artificial empathy implementation. Our discussion covers the following topics: How can AI recognize pain from unimodality and multimodality? Is it necessary for AI to be empathic? How can we create an AI agent with proactive and reactive empathy? This article explores the challenges and opportunities of real-world multimodal pain recognition from a psychological, neuroscientific, and artificial intelligence perspective. Finally, we identify possible future implementations of artificial empathy and analyze how humans might benefit from an AI agent equipped with empathy.
Automated Feature-Specific Tree Species Identification from Natural Images using Deep Semi-Supervised Learning
Homan, Dewald, Preez, Johan A. du
Prior work on plant species classification predominantly focuses on building models from isolated plant attributes. Hence, there is a need for tools that can assist in species identification in the natural world. We present a novel and robust two-fold approach capable of identifying trees in a real-world natural setting. Further, we leverage unlabelled data through deep semi-supervised learning and demonstrate superior performance to supervised learning. Our single-GPU implementation for feature recognition uses minimal annotated data and achieves accuracies of 93.96% and 93.11% for leaves and bark, respectively. Further, we extract feature-specific datasets of 50 species by employing this technique. Finally, our semi-supervised species classification method attains 94.04% top-5 accuracy for leaves and 83.04% top-5 accuracy for bark.
Robustness Evaluation of Transformer-based Form Field Extractors via Form Attacks
Xue, Le, Gao, Mingfei, Chen, Zeyuan, Xiong, Caiming, Xu, Ran
We propose a novel framework to evaluate the robustness of transformer-based form field extraction methods via form attacks. We introduce 14 novel form transformations to evaluate the vulnerability of the state-of-the-art field extractors against form attacks from both OCR level and form level, including OCR location/order rearrangement, form background manipulation and form field-value augmentation. We conduct robustness evaluation using real invoices and receipts, and perform comprehensive research analysis. Experimental results suggest that the evaluated models are very susceptible to form perturbations such as the variation of field-values (~15% drop in F1 score), the disarrangement of input text order(~15% drop in F1 score) and the disruption of the neighboring words of field-values(~10% drop in F1 score). Guided by the analysis, we make recommendations to improve the design of field extractors and the process of data collection.
A guided journey through non-interactive automatic story generation
We present a literature survey on non-interactive computational story generation. The article starts with the presentation of requirements for creative systems, three types of models of creativity (computational, socio-cultural, and individual), and models of human creative writing. Then it reviews each class of story generation approach depending on the used technology: story-schemas, analogy, rules, planning, evolutionary algorithms, implicit knowledge learning, and explicit knowledge learning. Before the concluding section, the article analyses the contributions of the reviewed work to improve the quality of the generated stories. This analysis addresses the description of the story characters, the use of narrative knowledge including about character believability, and the possible lack of more comprehensive or more detailed knowledge or creativity models. Finally, the article presents concluding remarks in the form of suggestions of research topics that might have a significant impact on the advancement of the state of the art on autonomous non-interactive story generation systems. The article concludes that the autonomous generation and adoption of the main idea to be conveyed and the autonomous design of the creativity ensuring criteria are possibly two of most important topics for future research.