Goto

Collaborating Authors

 Education


Label Structure Preserving Contrastive Embedding for Multi-Label Learning with Missing Labels

arXiv.org Artificial Intelligence

Contrastive learning (CL) has shown impressive advances in image representation learning in whichever supervised multi-class classification or unsupervised learning. However, these CL methods fail to be directly adapted to multi-label image classification due to the difficulty in defining the positive and negative instances to contrast a given anchor image in multi-label scenario, let the label missing one alone, implying that borrowing a commonly-used way from contrastive multi-class learning to define them will incur a lot of false negative instances unfavorable for learning. In this paper, with the introduction of a label correction mechanism to identify missing labels, we first elegantly generate positives and negatives for individual semantic labels of an anchor image, then define a unique contrastive loss for multi-label image classification with missing labels (CLML), the loss is able to accurately bring images close to their true positive images and false negative images, far away from their true negative images. Different from existing multi-label CL losses, CLML also preserves low-rank global and local label dependencies in the latent representation space where such dependencies have been shown to be helpful in dealing with missing labels. To the best of our knowledge, this is the first general multi-label CL loss in the missing-label scenario and thus can seamlessly be paired with those losses of any existing multi-label learning methods just via a single hyperparameter. The proposed strategy has been shown to improve the classification performance of the Resnet101 model by margins of 1.2%, 1.6%, and 1.3% respectively on three standard datasets, MSCOCO, VOC, and NUS-WIDE. Code is available at https://github.com/chuangua/ContrastiveLossMLML.


Learn to Adapt to New Environment from Past Experience and Few Pilot

arXiv.org Artificial Intelligence

In recent years, deep learning has been widely applied in communications and achieved remarkable performance improvement. Most of the existing works are based on data-driven deep learning, which requires a significant amount of training data for the communication model to adapt to new environments and results in huge computing resources for collecting data and retraining the model. In this paper, we will significantly reduce the required amount of training data for new environments by leveraging the learning experience from the known environments. Therefore, we introduce few-shot learning to enable the communication model to generalize to new environments, which is realized by an attention-based method. With the attention network embedded into the deep learning-based communication model, environments with different power delay profiles can be learnt together in the training process, which is called the learning experience. By exploiting the learning experience, the communication model only requires few pilot blocks to perform well in the new environment. Through an example of deep-learning-based channel estimation, we demonstrate that this novel design method achieves better performance than the existing data-driven approach designed for few-shot learning.


A Novel Self-Knowledge Distillation Approach with Siamese Representation Learning for Action Recognition

arXiv.org Artificial Intelligence

Knowledge distillation is an effective transfer of knowledge from a heavy network (teacher) to a small network (student) to boost students' performance. Self-knowledge distillation, the special case of knowledge distillation, has been proposed to remove the large teacher network training process while preserving the student's performance. This paper introduces a novel Self-knowledge distillation approach via Siamese representation learning, which minimizes the difference between two representation vectors of the two different views from a given sample. Our proposed method, SKD-SRL, utilizes both soft label distillation and the similarity of representation vectors. Therefore, SKD-SRL can generate more consistent predictions and representations in various views of the same data point. Our benchmark has been evaluated on various standard datasets. The experimental results have shown that SKD-SRL significantly improves the accuracy compared to existing supervised learning and knowledge distillation methods regardless of the networks.


A PDE approach for regret bounds under partial monitoring

arXiv.org Artificial Intelligence

In this paper, we study a learning problem in which a forecaster only observes partial information. By properly rescaling the problem, we heuristically derive a limiting PDE on Wasserstein space which characterizes the asymptotic behavior of the regret of the forecaster. Using a verification type argument, we show that the problem of obtaining regret bounds and efficient algorithms can be tackled by finding appropriate smooth sub/supersolutions of this parabolic PDE.


University of Texas researchers develop brainlike transistors

#artificialintelligence

University of Texas researchers have developed new biocompatible transistors that mimic brain synapses, an advancement that could help scientists rebuild neural pathways or create brain implants. Though the transistors are not ready for use in humans, the team hopes to use them to create brainlike computers that can work alongside the human brain, said Jean Anne Incorvia, an assistant professor of electrical and computer engineering at UT. "There's a big goal in our field of building brain-inspired computers," Incorvia said. "We can imagine that it could seamlessly work alongside a human brain so the human brain is doing processing, and then at some point, it's connected to these (synaptic transistors), which then start doing neuromorphic computer processing with it. So we have this brain-machine combination for doing tasks." The transistors transfer signals in the same way synapses transfer impulses between neurons, co-author Dmitry Kireev said.


A Chinese game company has appointed the world's first humanoid robot as its CEO

#artificialintelligence

Tang Yu's duties also include providing a fair working environment for the employees. It would not make much sense to think of today's technology without artificial intelligence. Therefore, the founder of the company Dr. Dejian Liu, also emphasizes the importance of this spurt in the company. "We believe AI is the future of corporate management, and our appointment of Ms. Tang Yu represents our commitment to truly embrace the use of AI to transform the way we operate our business and ultimately drive our future strategic growth," said Dr. Dejian Liu. "Looking forward, we will continue to expand on our algorithms behind Tang Yu to build an open, interactive and highly transparent management model as we gradually transform to a metaverse-based working community, which will enable us to attract a much broader base of talents worldwide and put us in a position to achieve bigger goals," he also added. One of the most respected and well-known online game developers in China, NetDragon was founded in 1999 and has produced a number of popular games, including Eudemons Online, Heroes Evolved, Conquer Online, and Under Oath.


Researchers develop new strategies to teach computers to learn like humans do

#artificialintelligence

As demonstrated by breakthroughs in various fields of artificial intelligence (AI), such as image processing, smart health care, self-driving vehicles and smart cities, this is undoubtedly the golden period of deep learning. In the next decade or so, AI and computing systems will eventually be equipped with the ability to learn and think the way humans do--to process continuous flow of information and interact with the real world. However, current AI models suffer from a performance loss when they are trained consecutively on new information. This is because every time new data is generated, it is written on top of existing data, thus erasing previous information. This effect is known as "catastrophic forgetting."


Artificial Intelligence, Virtual Reality Or Gamification Lead The Educational Revolution

#artificialintelligence

The use of new technologies in the educational sector has increased since 2020 due to the advent of the pandemic. Tablets and computers have become essential devices, so much so that the latest data from INE indicates that 96% of Spanish households had internet access in 2021. During the same period, The training centers had to reinvent themselves to adapt to the new times and many of them used so called edtechBoth software and hardware tools that allow optimization and improvement of learning processes that yield better results. The application of technology to education has allowed greater access to knowledge as it has taken learning out of the classroom and, in turn, has provided more tools to create environments where the theoretical and practical work for technologies converge. Such as artificial intelligence or virtual reality, which are increasingly used in schools to generate better learning experiences.


Video-Guided Curriculum Learning for Spoken Video Grounding

arXiv.org Artificial Intelligence

In this paper, we introduce a new task, spoken video grounding (SVG), which aims to localize the desired video fragments from spoken language descriptions. Compared with using text, employing audio requires the model to directly exploit the useful phonemes and syllables related to the video from raw speech. Moreover, we randomly add environmental noises to this speech audio, further increasing the difficulty of this task and better simulating real applications. To rectify the discriminative phonemes and extract video-related information from noisy audio, we develop a novel video-guided curriculum learning (VGCL) during the audio pre-training process, which can make use of the vital visual perceptions to help understand the spoken language and suppress the external noise. Considering during inference the model can not obtain ground truth video segments, we design a curriculum strategy that gradually shifts the input video from the ground truth to the entire video content during pre-training. Finally, the model can learn how to extract critical visual information from the entire video clip to help understand the spoken language. In addition, we collect the first large-scale spoken video grounding dataset based on ActivityNet, which is named as ActivityNet Speech dataset. Extensive experiments demonstrate our proposed video-guided curriculum learning can facilitate the pre-training process to obtain a mutual audio encoder, significantly promoting the performance of spoken video grounding tasks. Moreover, we prove that in the case of noisy sound, our model outperforms the method that grounding video with ASR transcripts, further demonstrating the effectiveness of our curriculum strategy.


Learning with Differentiable Algorithms

arXiv.org Artificial Intelligence

Classic algorithms and machine learning systems like neural networks are both abundant in everyday life. While classic computer science algorithms are suitable for precise execution of exactly defined tasks such as finding the shortest path in a large graph, neural networks allow learning from data to predict the most likely answer in more complex tasks such as image classification, which cannot be reduced to an exact algorithm. To get the best of both worlds, this thesis explores combining both concepts leading to more robust, better performing, more interpretable, more computationally efficient, and more data efficient architectures. The thesis formalizes the idea of algorithmic supervision, which allows a neural network to learn from or in conjunction with an algorithm. When integrating an algorithm into a neural architecture, it is important that the algorithm is differentiable such that the architecture can be trained end-to-end and gradients can be propagated back through the algorithm in a meaningful way. To make algorithms differentiable, this thesis proposes a general method for continuously relaxing algorithms by perturbing variables and approximating the expectation value in closed form, i.e., without sampling. In addition, this thesis proposes differentiable algorithms, such as differentiable sorting networks, differentiable renderers, and differentiable logic gate networks. Finally, this thesis presents alternative training strategies for learning with algorithms.