Information Fusion
Combining Incomplete Observational and Randomized Data for Heterogeneous Treatment Effects
Yao, Dong, Tang, Caizhi, Cui, Qing, Li, Longfei
Data from observational studies (OSs) is widely available and readily obtainable yet frequently contains confounding biases. On the other hand, data derived from randomized controlled trials (RCTs) helps to reduce these biases; however, it is expensive to gather, resulting in a tiny size of randomized data. For this reason, effectively fusing observational data and randomized data to better estimate heterogeneous treatment effects (HTEs) has gained increasing attention. However, existing methods for integrating observational data with randomized data must require \textit{complete} observational data, meaning that both treated subjects and untreated subjects must be included in OSs. This prerequisite confines the applicability of such methods to very specific situations, given that including all subjects, whether treated or untreated, in observational studies is not consistently achievable. In our paper, we propose a resilient approach to \textbf{C}ombine \textbf{I}ncomplete \textbf{O}bservational data and randomized data for HTE estimation, which we abbreviate as \textbf{CIO}. The CIO is capable of estimating HTEs efficiently regardless of the completeness of the observational data, be it full or partial. Concretely, a confounding bias function is first derived using the pseudo-experimental group from OSs, in conjunction with the pseudo-control group from RCTs, via an effect estimation procedure. This function is subsequently utilized as a corrective residual to rectify the observed outcomes of observational data during the HTE estimation by combining the available observational data and the all randomized data. To validate our approach, we have conducted experiments on a synthetic dataset and two semi-synthetic datasets.
Smart ETL and LLM-based contents classification: the European Smart Tourism Tools Observatory experience
Cosme, Diogo, Galvรฃo, Antรณnio, Abreu, Fernando Brito e
Purpose: Our research project focuses on improving the content update of the online European Smart Tourism Tools (STTs) Observatory by incorporating and categorizing STTs. The categorization is based on their taxonomy, and it facilitates the end user's search process. The use of a Smart ETL (Extract, Transform, and Load) process, where \emph{Smart} indicates the use of Artificial Intelligence (AI), is central to this endeavor. Methods: The contents describing STTs are derived from PDF catalogs, where PDF-scraping techniques extract QR codes, images, links, and text information. Duplicate STTs between the catalogs are removed, and the remaining ones are classified based on their text information using Large Language Models (LLMs). Finally, the data is transformed to comply with the Dublin Core metadata structure (the observatory's metadata structure), chosen for its wide acceptance and flexibility. Results: The Smart ETL process to import STTs to the observatory combines PDF-scraping techniques with LLMs for text content-based classification. Our preliminary results have demonstrated the potential of LLMs for text content-based classification. Conclusion: The proposed approach's feasibility is a step towards efficient content-based classification, not only in Smart Tourism but also adaptable to other fields. Future work will mainly focus on refining this classification process.
Precision Soil Quality Analysis Using Transformer-based Data Fusion Strategies: A Systematic Review
Saki, Mahdi, Keshavarz, Rasool, Franklin, Daniel, Abolhasan, Mehran, Lipman, Justin, Shariati, Negin
The transformer-based data fusion techniques in agricultural implementation of PA, also known as smart farming, relies remote sensing (RS), with a particular focus on soil on the ability to collect, process, and analyse spatial and analysis. Utilizing a systematic, data-driven approach, we temporal data to optimize field management practices demonstrate that transformers have significantly (Cisternas et al., 2020; Pyingkodi et al., 2022). Despite its outperformed conventional deep learning and machine enormous potential, the adoption of PA remains below learning methods since 2022, achieving prediction expectations due to factors such as high initial investment performance between 92% and 97%. The review is costs, the complexity of IT, and the need for specialized specifically focused on soil analysis, due to the importance knowledge (Cisternas et al., 2020). of soil condition in optimizing crop productivity and Remote sensing (RS) has seen rapid advancements and ensuring sustainable farming practices. Transformer-based widespread adoption in PA, offering high-resolution data models have shown remarkable capabilities in handling for applications ranging from crop monitoring to irrigation complex multivariate soil data, improving the accuracy of management (Sishodia et al., 2020). Remote sensing has soil moisture prediction, soil element analysis, and other proven to be an effective tool for capturing and monitoring soil-related applications. This systematic review primarily the spectral and temporal properties of the land surface focuses on 1) analysing research trends and patterns in the influenced by human activities at different spatial and literature, both chronologically and technically, and 2) temporal scales (Bรฉguรฉ et al., 2018).
Captions Speak Louder than Images (CASLIE): Generalizing Foundation Models for E-commerce from High-quality Multimodal Instruction Data
Ling, Xinyi, Peng, Bo, Du, Hanwen, Zhu, Zhihui, Ning, Xia
Leveraging multimodal data to drive breakthroughs in e-commerce applications through Multimodal Foundation Models (MFMs) is gaining increasing attention from the research community. However, there are significant challenges that hinder the optimal use of multimodal e-commerce data by foundation models: (1) the scarcity of large-scale, high-quality multimodal benchmark datasets; and (2) the lack of effective multimodal information integration methods. To address these challenges, in this paper, we introduce MMECInstruct, the first-ever, large-scale, and high-quality multimodal instruction dataset for e-commerce. We also develop CASLIE, a simple, lightweight, yet effective framework for integrating multimodal information for e-commerce. Leveraging MMECInstruct, we fine-tune a series of e-commerce MFMs within CASLIE, denoted as CASLIE models. Our comprehensive evaluation demonstrates that CASLIE models substantially outperform 5 categories of advanced baseline models in the in-domain evaluation. Moreover, CASLIE models show strong generalizability to out-of-domain settings. MMECInstruct and CASLIE models are publicly accessible through https://ninglab.github.io/CASLIE/.
Multi-Sensor Fusion for UAV Classification Based on Feature Maps of Image and Radar Data
Sakellariou, Nikos, Lalas, Antonios, Votis, Konstantinos, Tzovaras, Dimitrios
Continuous Wave (FMCW) radars represent the most Unmanned Aerial Vehicles (UAV) have successfully attractive and cost-efficient solutions [2]. While for the permeated modern society with various applications for civil verification and classification task various methods exist in and military purposes. Oil and gas, construction, metals and literature employing machine learning techniques such as mining already incorporate UAVs in their processes. SVM [3], Random Forests [4], Nearest Neighbor [5] and Furthermore, UAVs are employed for commercial purposes, Deep Neural Networks [6][7][8]. More recent DNN such as the monitoring of public places, cartography, survey approaches based on convolutional neural networks are wildlife, search and rescue (SAR), first aid and delivery of introduced in Samaras et al. [9]. The authors presented a deep goods. Big technological companies continuously challenge learning classification method based on data from an X-band the status quote by announcing breakthrough services. FMCW surveillance 2D radar that is able to reach a Moreover, progress in UAV regulation has driven classification accuracy of up to 95.0% utilizing a custom investments since 2019, to further increase the popularity and CNN based architecture. A similar approach is presented in use of UAVs in sectors that present significant potential but [10] where the authors proposed Res-Net-SP, a compressed still minimal use, such as agriculture, healthcare, architecture of ResNet-18 that is based in convolutional infrastructure, property management and insurance.
Kaninfradet3D:A Road-side Camera-LiDAR Fusion 3D Perception Model based on Nonlinear Feature Extraction and Intrinsic Correlation
Liu, Pei, Zheng, Nanfang, Li, Yiqun, Chen, Junlan, Pu, Ziyuan
With the development of AI-assisted driving, numerous methods have emerged for ego-vehicle 3D perception tasks, but there has been limited research on roadside perception. With its ability to provide a global view and a broader sensing range, the roadside perspective is worth developing. LiDAR provides precise three-dimensional spatial information, while cameras offer semantic information. These two modalities are complementary in 3D detection. However, adding camera data does not increase accuracy in some studies since the information extraction and fusion procedure is not sufficiently reliable. Recently, Kolmogorov-Arnold Networks (KANs) have been proposed as replacements for MLPs, which are better suited for high-dimensional, complex data. Both the camera and the LiDAR provide high-dimensional information, and employing KANs should enhance the extraction of valuable features to produce better fusion outcomes. This paper proposes Kaninfradet3D, which optimizes the feature extraction and fusion modules. To extract features from complex high-dimensional data, the model's encoder and fuser modules were improved using KAN Layers. Cross-attention was applied to enhance feature fusion, and visual comparisons verified that camera features were more evenly integrated. This addressed the issue of camera features being abnormally concentrated, negatively impacting fusion. Compared to the benchmark, our approach shows improvements of +9.87 mAP and +10.64 mAP in the two viewpoints of the TUMTraf Intersection Dataset and an improvement of +1.40 mAP in the roadside end of the TUMTraf V2X Cooperative Perception Dataset. The results indicate that Kaninfradet3D can effectively fuse features, demonstrating the potential of applying KANs in roadside perception tasks.
Random Token Fusion for Multi-View Medical Diagnosis
Guo, Jingyu, Matsoukas, Christos, Strand, Fredrik, Smith, Kevin
In multi-view medical diagnosis, deep learning-based models often fuse information from different imaging perspectives to improve diagnostic performance. However, existing approaches are prone to overfitting and rely heavily on view-specific features, which can lead to trivial solutions. In this work, we introduce Random Token Fusion (RTF), a novel technique designed to enhance multi-view medical image analysis using vision transformers. By integrating randomness into the feature fusion process during training, RTF addresses the issue of overfitting and enhances the robustness and accuracy of diagnostic models without incurring any additional cost at inference. We validate our approach on standard mammography and chest X-ray benchmark datasets. Through extensive experiments, we demonstrate that RTF consistently improves the performance of existing fusion methods, paving the way for a new generation of multi-view medical foundation models.
AMPLE: Emotion-Aware Multimodal Fusion Prompt Learning for Fake News Detection
Xu, Xiaoman, Li, Xiangrun, Wang, Taihang, Jiang, Ye
Detecting fake news in large datasets is challenging due to its diversity and complexity, with traditional approaches often focusing on textual features while underutilizing semantic and emotional elements. Current methods also rely heavily on large annotated datasets, limiting their effectiveness in more nuanced analysis. To address these challenges, this paper introduces Emotion-\textbf{A}ware \textbf{M}ultimodal Fusion \textbf{P}rompt \textbf{L}\textbf{E}arning (\textbf{AMPLE}) framework to address the above issue by combining text sentiment analysis with multimodal data and hybrid prompt templates. This framework extracts emotional elements from texts by leveraging sentiment analysis tools. It then employs Multi-Head Cross-Attention (MCA) mechanisms and similarity-aware fusion methods to integrate multimodal data. The proposed AMPLE framework demonstrates strong performance on two public datasets in both few-shot and data-rich settings, with results indicating the potential of emotional aspects in fake news detection. Furthermore, the study explores the impact of integrating large language models with this method for text sentiment extraction, revealing substantial room for further improvement. The code can be found at :\url{https://github.com/xxm1215/MMM2025_few-shot/
Collaborative State Fusion in Partially Known Multi-agent Environments
Zhou, Tianlong, Shang, Jun, Rao, Weixiong
In this paper, we study the collaborative state fusion problem in a multi-agent environment, where mobile agents collaborate to track movable targets. Due to the limited sensing range and potential errors of on-board sensors, it is necessary to aggregate individual observations to provide target state fusion for better target state estimation. Existing schemes do not perform well due to (1) impractical assumption of the fully known prior target state-space model and (2) observation outliers from individual sensors. To address the issues, we propose a two-stage collaborative fusion framework, namely \underline{L}earnable Weighted R\underline{o}bust \underline{F}usion (\textsf{LoF}). \textsf{LoF} combines a local state estimator (e.g., Kalman Filter) with a learnable weight generator to address the mismatch between the prior state-space model and underlying patterns of moving targets. Moreover, given observation outliers, we develop a time-series soft medoid(TSM) scheme to perform robust fusion. We evaluate \textsf{LoF} in a collaborative detection simulation environment with promising results. In an example setting with 4 agents and 2 targets, \textsf{LoF} leads to a 9.1\% higher fusion gain compared to the state-of-the-art.
Cocoon: Robust Multi-Modal Perception with Uncertainty-Aware Sensor Fusion
Cho, Minkyoung, Cao, Yulong, Sun, Jiachen, Zhang, Qingzhao, Pavone, Marco, Park, Jeong Joon, Yang, Heng, Mao, Z. Morley
An important paradigm in 3D object detection is the use of multiple modalities to enhance accuracy in both normal and challenging conditions, particularly for long-tail scenarios. To address this, recent studies have explored two directions of adaptive approaches: MoE-based adaptive fusion, which struggles with uncertainties arising from distinct object configurations, and late fusion for output-level adaptive fusion, which relies on separate detection pipelines and limits comprehensive understanding. In this work, we introduce Cocoon, an object- and feature-level uncertainty-aware fusion framework. The key innovation lies in uncertainty quantification for heterogeneous representations, enabling fair comparison across modalities through the introduction of a feature aligner and a learnable surrogate ground truth, termed feature impression. We also define a training objective to ensure that their relationship provides a valid metric for uncertainty quantification. Cocoon consistently outperforms existing static and adaptive methods in both normal and challenging conditions, including those with natural and artificial corruptions. Furthermore, we show the validity and efficacy of our uncertainty metric across diverse datasets.