Africa
Grasping, Part Identification, and Pose Refinement in One Shot with a Tactile Gripper
Lim, Joyce Xin-Yan, Pham, Quang-Cuong
The rise in additive manufacturing comes with unique opportunities and challenges. Rapid changes to part design and massive part customization distinctive to 3D-Print (3DP) can be easily achieved. Customized parts that are unique, yet exhibit similar features such as dental moulds, shoe insoles, or engine vanes could be industrially manufactured with 3DP. However, the opportunity for massive part customization comes with unique challenges for the existing production paradigm of robotics applications, as the current robotics paradigm for part identification and pose refinement is repetitive, where data-driven and object-dependent approaches are often used. Thus, a bottleneck exists in robotics applications for 3DP parts where massive customization is involved, as it is difficult for feature-based deep learning approaches to distinguish between similar parts such as shoe insoles belonging to different people. As such, we propose a method that augments patterns on 3DP parts so that grasping, part identification, and pose refinement can be executed in one shot with a tactile gripper. We also experimentally evaluate our approach from three perspectives, including real insertion tasks that mimic robotic sorting and packing, and achieved excellent classification results, a high insertion success rate of 95%, and a sub-millimeter pose refinement accuracy.
LLM Factoscope: Uncovering LLMs' Factual Discernment through Inner States Analysis
He, Jinwen, Gong, Yujia, Chen, Kai, Lin, Zijin, Wei, Chengan, Zhao, Yue
Large Language Models (LLMs) have revolutionized various domains with extensive knowledge and creative capabilities. However, a critical issue with LLMs is their tendency to produce outputs that diverge from factual reality. This phenomenon is particularly concerning in sensitive applications such as medical consultation and legal advice, where accuracy is paramount. In this paper, we introduce the LLM factoscope, a novel Siamese network-based model that leverages the inner states of LLMs for factual detection. Our investigation reveals distinguishable patterns in LLMs' inner states when generating factual versus non-factual content. We demonstrate the LLM factoscope's effectiveness across various architectures, achieving over 96% accuracy in factual detection. Our work opens a new avenue for utilizing LLMs' inner states for factual detection and encourages further exploration into LLMs' inner workings for enhanced reliability and transparency.
Beyond Isolation: Multi-Agent Synergy for Improving Knowledge Graph Construction
Ye, Hongbin, Gui, Honghao, Zhang, Aijia, Liu, Tong, Hua, Wei, Jia, Weiqiang
Knowledge graph construction (KGC) is a multifaceted undertaking involving the extraction of entities, relations, and events. Traditionally, large language models (LLMs) have been viewed as solitary task-solving agents in this complex landscape. However, this paper challenges this paradigm by introducing a novel framework, CooperKGC. Departing from the conventional approach, CooperKGC establishes a collaborative processing network, assembling a KGC collaboration team capable of concurrently addressing entity, relation, and event extraction tasks. Our experiments unequivocally demonstrate that fostering collaboration and information interaction among diverse agents within CooperKGC yields superior results compared to individual cognitive processes operating in isolation. Importantly, our findings reveal that the collaboration facilitated by CooperKGC enhances knowledge selection, correction, and aggregation capabilities across multiple rounds of interactions.
ImputeFormer: Low Rankness-Induced Transformers for Generalizable Spatiotemporal Imputation
Nie, Tong, Qin, Guoyang, Ma, Wei, Mei, Yuewen, Sun, Jian
Missing data is a pervasive issue in both scientific and engineering tasks, especially for the modeling of spatiotemporal data. This problem attracts many studies to contribute to machine learning solutions. Existing imputation solutions mainly include low-rank models and deep learning models. On the one hand, low-rank models assume general structural priors, but have limited model capacity. On the other hand, deep learning models possess salient features of expressivity, while lack prior knowledge of the spatiotemporal process. Leveraging the strengths of both two paradigms, we demonstrate a low rankness-induced Transformer model to achieve a balance between strong inductive bias and high model expressivity. The exploitation of the inherent structures of spatiotemporal data enables our model to learn balanced signal-noise representations, making it versatile for a variety of imputation problems. We demonstrate its superiority in terms of accuracy, efficiency, and generality in heterogeneous datasets, including traffic speed, traffic volume, solar energy, smart metering, and air quality. Comprehensive case studies are performed to further strengthen interpretability. Promising empirical results provide strong conviction that incorporating time series primitives, such as low-rank properties, can substantially facilitate the development of a generalizable model to approach a wide range of spatiotemporal imputation problems.
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
Zhang, Gengyuan, Zhang, Yurui, Zhang, Kerui, Tresp, Volker
Vision-Language Models (VLMs) are expected to be capable of reasoning with commonsense knowledge as human beings. One example is that humans can reason where and when an image is taken based on their knowledge. This makes us wonder if, based on visual cues, Vision-Language Models that are pre-trained with large-scale image-text resources can achieve and even outperform human's capability in reasoning times and location. To address this question, we propose a two-stage \recognition\space and \reasoning\space probing task, applied to discriminative and generative VLMs to uncover whether VLMs can recognize times and location-relevant features and further reason about it. To facilitate the investigation, we introduce WikiTiLo, a well-curated image dataset compromising images with rich socio-cultural cues. In the extensive experimental studies, we find that although VLMs can effectively retain relevant features in visual encoders, they still fail to make perfect reasoning. We will release our dataset and codes to facilitate future studies.
OpenAI became the nexus of the technology world in 2023
Let's take a look at how OpenAI and its chatbot have impacted consumer electronics in 2023 and where they might lead the industry in the new year. "Meteoric" doesn't do justice to OpenAI's rise this year. The company released ChatGPT on November 30, 2022. Within five days, the program had passed 1 million users; by January, 100 million people a month were logging on to use it. It took Facebook four and a half years to reach those sorts of engagement numbers.
McAfee Ultimate review: Comprehensive security that needs more polish
McAfee Ultimate offers strong antivirus protection and a vast array of online protections, but its apps, services, and tools could use more polish. Its scans also can tangibly decrease performance on mid-range and budget PCs. As attractive as this comprehensive all-in-one package is, it's currently a hard sell. Among the top-tier antivirus software plans, McAfee's version is an especially loaded offering--and less common in how it bundles together an extraordinary number of online protections. Many rivals have a premium antivirus suite, then offer services like a VPN, password manager, and identity protection and recovery as separate subscriptions. This simplifies how much you have to think about, of course, but there's just one problem--this security suite lacks the polish you'd expect of such a premium product. Further reading: See our roundup of the best antivirus software for Windows PCs to learn about competing products. The full list of features in McAfee's flagship subscription is exhaustive. Antivirus, link screening, and firewall protection are just the start.
Online Tensor Inference
Wen, Xin, Sun, Will Wei, Zhang, Yichen
Recent technological advances have led to contemporary applications that demand real-time processing and analysis of sequentially arriving tensor data. Traditional offline learning, involving the storage and utilization of all data in each computational iteration, becomes impractical for high-dimensional tensor data due to its voluminous size. Furthermore, existing low-rank tensor methods lack the capability for statistical inference in an online fashion, which is essential for real-time predictions and informed decision-making. This paper addresses these challenges by introducing a novel online inference framework for low-rank tensor learning. Our approach employs Stochastic Gradient Descent (SGD) to enable efficient real-time data processing without extensive memory requirements, thereby significantly reducing computational demands. We establish a non-asymptotic convergence result for the online low-rank SGD estimator, nearly matches the minimax optimal rate of estimation error in offline models that store all historical data. Building upon this foundation, we propose a simple yet powerful online debiasing approach for sequential statistical inference in low-rank tensor learning. The entire online procedure, covering both estimation and inference, eliminates the need for data splitting or storing historical data, making it suitable for on-the-fly hypothesis testing. Given the sequential nature of our data collection, traditional analyses relying on offline methods and sample splitting are inadequate. In our analysis, we control the sum of constructed super-martingales to ensure estimates along the entire solution path remain within the benign region. Additionally, a novel spectral representation tool is employed to address statistical dependencies among iterative estimates, establishing the desired asymptotic normality.
Multimodal Sentiment Analysis with Missing Modality: A Knowledge-Transfer Approach
Liu, Weide, Zhan, Huijing, Chen, Hao, Lv, Fengmao
Previous research studies [11, 12] have attempted to address the issue of missing modalities in multimodal sentiment Multimodal sentiment analysis aims to identify the emotions analysis. In particular, Tsai et al. [12] proposed a joint expressed by individuals through visual, language, and generative-discriminative objective to obtain a robust multimodal acoustic cues. However, most of the existing research efforts representation and a surrogate inference model for assume that all modalities are available during both missing modalities. Pham et al. [11] developed a multimodal training and testing, making their algorithms susceptible to translation network with a cyclic translation loss for forward the missing modality scenario. In this paper, we propose a adaptation between source and target modalities. However, novel knowledge-transfer network to translate between different the performances of their approaches degrade when complete modalities to reconstruct the missing audio modalities.
Beyond PID Controllers: PPO with Neuralized PID Policy for Proton Beam Intensity Control in Mu2e
Xu, Chenwei, Hu, Jerry Yao-Chieh, Narayanan, Aakaash, Thieme, Mattson, Nagaslaev, Vladimir, Austin, Mark, Arnold, Jeremy, Berlioz, Jose, Hanlet, Pierrick, Ibrahim, Aisha, Nicklaus, Dennis, Mitrevski, Jovan, John, Jason Michael St., Pradhan, Gauri, Saewert, Andrea, Seiya, Kiyomi, Schupbach, Brian, Thurman-Keup, Randy, Tran, Nhan, Shi, Rui, Ogrenci, Seda, Shuping, Alexis Maya-Isabelle, Hazelwood, Kyle, Liu, Han
We introduce a novel Proximal Policy Optimization (PPO) algorithm aimed at addressing the challenge of maintaining a uniform proton beam intensity delivery in the Muon to Electron Conversion Experiment (Mu2e) at Fermi National Accelerator Laboratory (Fermilab). Our primary objective is to regulate the spill process to ensure a consistent intensity profile, with the ultimate goal of creating an automated controller capable of providing real-time feedback and calibration of the Spill Regulation System (SRS) parameters on a millisecond timescale. We treat the Mu2e accelerator system as a Markov Decision Process suitable for Reinforcement Learning (RL), utilizing PPO to reduce bias and enhance training stability. A key innovation in our approach is the integration of a neuralized Proportional-Integral-Derivative (PID) controller into the policy function, resulting in a significant improvement in the Spill Duty Factor (SDF) by 13.6%, surpassing the performance of the current PID controller baseline by an additional 1.6%. This paper presents the preliminary offline results based on a differentiable simulator of the Mu2e accelerator. It paves the groundwork for real-time implementations and applications, representing a crucial step towards automated proton beam intensity control for the Mu2e experiment.