Africa
Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report Generation
Sun, Liwen, Zhao, James, Han, Megan, Xiong, Chenyan
Multimodal foundation models hold significant potential for automating radiology report generation, thereby assisting clinicians in diagnosing cardiac diseases. However, generated reports often suffer from serious factual inaccuracy. In this paper, we introduce a fact-aware multimodal retrieval-augmented pipeline in generating accurate radiology reports (FactMM-RAG). We first leverage RadGraph to mine factual report pairs, then integrate factual knowledge to train a universal multimodal retriever. Given a radiology image, our retriever can identify high-quality reference reports to augment multimodal foundation models, thus enhancing the factual completeness and correctness of report generation. Experiments on two benchmark datasets show that our multimodal retriever outperforms state-of-the-art retrievers on both language generation and radiology-specific metrics, up to 6.5% and 2% score in F1CheXbert and F1RadGraph. Further analysis indicates that employing our factually-informed training strategy imposes an effective supervision signal, without relying on explicit diagnostic label guidance, and successfully propagates fact-aware capabilities from the multimodal retriever to the multimodal foundation model in radiology report generation.
Improving Minimum Bayes Risk Decoding with Multi-Prompt
Heineman, David, Dou, Yao, Xu, Wei
While instruction fine-tuned LLMs are effective text generators, sensitivity to prompt construction makes performance unstable and sub-optimal in practice. Relying on a single "best" prompt cannot capture all differing approaches to a generation problem. Using this observation, we propose multi-prompt decoding, where many candidate generations are decoded from a prompt bank at inference-time. To ensemble candidates, we use Minimum Bayes Risk (MBR) decoding, which selects a final output using a trained value metric. We show multi-prompt improves MBR across a comprehensive set of conditional generation tasks, and show this is a result of estimating a more diverse and higher quality candidate space than that of a single prompt. Further experiments confirm multi-prompt improves generation across tasks, models and metrics.
Houthi Drone Strike Highlights Dilemmas for Israel
One immediate, short-term response, some analysts said, might be a cease-fire deal between Hamas and Israel, a move that could halt attacks from Hamas's allies, like the Houthis and Hezbollah in Lebanon. While the Houthis' opposition to Israel long preceded the war in Gaza, the group had rarely attacked Israeli interests before it began. A truce in Gaza could "prompt some kind of a lull for a while" in Yemen and Lebanon, said Relik Shafir, a former general in the Israeli Air Force. But while mediators say they are edging closer to sealing a Gaza cease-fire, key gaps between Israel and Hamas remain, and parts of Prime Minister Benjamin Netanyahu's right-wing coalition oppose compromising on Hamas's main demands. In the long term, the Houthis also remain committed to Israel's total destruction and would most likely not be placated for long by a temporary truce in Gaza or an end to Israel's occupation of the West Bank. The Houthis are a Yemeni Shiite militia that over the past decade seized control of large parts of western Yemen, including its capital, Sana, and Red Sea coastline.
Diffusion Models as Data Mining Tools
Siglidis, Ioannis, Holynski, Aleksander, Efros, Alexei A., Aubry, Mathieu, Ginosar, Shiry
This paper demonstrates how to use generative models trained for image synthesis as tools for visual data mining. Our insight is that since contemporary generative models learn an accurate representation of their training data, we can use them to summarize the data by mining for visual patterns. Concretely, we show that after finetuning conditional diffusion models to synthesize images from a specific dataset, we can use these models to define a typicality measure on that dataset. This measure assesses how typical visual elements are for different data labels, such as geographic location, time stamps, semantic labels, or even the presence of a disease. This analysis-by-synthesis approach to data mining has two key advantages. First, it scales much better than traditional correspondence-based approaches since it does not require explicitly comparing all pairs of visual elements. Second, while most previous works on visual data mining focus on a single dataset, our approach works on diverse datasets in terms of content and scale, including a historical car dataset, a historical face dataset, a large worldwide street-view dataset, and an even larger scene dataset. Furthermore, our approach allows for translating visual elements across class labels and analyzing consistent changes.
Diff4VS: HIV-inhibiting Molecules Generation with Classifier Guidance Diffusion for Virtual Screening
Lyu, Jiaqing, Chen, Changjie, Liang, Bing, Zhang, Yijia
The AIDS epidemic has killed 40 million people and caused serious global problems. The identification of new HIV-inhibiting molecules is of great importance for combating the AIDS epidemic. Here, the Classifier Guidance Diffusion model and ligand-based virtual screening strategy are combined to discover potential HIV-inhibiting molecules for the first time. We call it Diff4VS. An extra classifier is trained using the HIV molecule dataset, and the gradient of the classifier is used to guide the Diffusion to generate HIV-inhibiting molecules. Experiments show that Diff4VS can generate more candidate HIV-inhibiting molecules than other methods. Inspired by ligand-based virtual screening, a new metric DrugIndex is proposed. The DrugIndex is the ratio of the proportion of candidate drug molecules in the generated molecule to the proportion of candidate drug molecules in the training set. DrugIndex provides a new evaluation method for evolving molecular generative models from a pharmaceutical perspective. Besides, we report a new phenomenon observed when using molecule generation models for virtual screening. Compared to real molecules, the generated molecules have a lower proportion that is highly similar to known drug molecules. We call it Degradation in molecule generation. Based on the data analysis, the Degradation may result from the difficulty of generating molecules with a specific structure in the generative model. Our research contributes to the application of generative models in drug design from method, metric, and phenomenon analysis.
Enhancing Microgrid Performance Prediction with Attention-based Deep Learning Models
Maddineni, Vinod Kumar, Koganti, Naga Babu, Damacharla, Praveen
In this research, an effort is made to address microgrid systems' operational challenges, characterized by power oscillations that eventually contribute to grid instability. An integrated strategy is proposed, leveraging the strengths of convolutional and Gated Recurrent Unit (GRU) layers. This approach is aimed at effectively extracting temporal data from energy datasets to improve the precision of microgrid behavior forecasts. Additionally, an attention layer is employed to underscore significant features within the time-series data, optimizing the forecasting process. The framework is anchored by a Multi-Layer Perceptron (MLP) model, which is tasked with comprehensive load forecasting and the identification of abnormal grid behaviors. Our methodology underwent rigorous evaluation using the Micro-grid Tariff Assessment Tool dataset, with Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and the coefficient of determination (r2-score) serving as the primary metrics. The approach demonstrated exemplary performance, evidenced by a MAE of 0.39, RMSE of 0.28, and an r2-score of 98.89\% in load forecasting, along with near-perfect zero state prediction accuracy (approximately 99.9\%). Significantly outperforming conventional machine learning models such as support vector regression and random forest regression, our model's streamlined architecture is particularly suitable for real-time applications, thereby facilitating more effective and reliable microgrid management.
Visual Geo-Localization from images
Algorithms process this data to pinpoint exact coordinates[11][12]. Geo-localization is important for organizing and analyzing large volumes of imagery data, as demonstrated by systems like the US Geological Survey (USGS), which classify and locate satellite and drone images to streamline data collection and analysis. Social media platforms like Instagram use geo-localization to tag photos with specific locations, enabling users to explore location-based content[11]. Despite its significance, many images and videos lack geo-localization data, particularly those collected in the past or by devices without GPS capabilities[12].
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models
Hossain, Md Zarif, Imteaj, Ahmed
Vision-language models (VLMs) have achieved significant strides in recent times specially in multimodal tasks, yet they remain susceptible to adversarial attacks on their vision components. To address this, we propose Sim-CLIP, an unsupervised adversarial fine-tuning method that enhances the robustness of the widely-used CLIP vision encoder against such attacks while maintaining semantic richness and specificity. By employing a Siamese architecture with cosine similarity loss, Sim-CLIP learns semantically meaningful and attack-resilient visual representations without requiring large batch sizes or momentum encoders. Our results demonstrate that VLMs enhanced with Sim-CLIP's fine-tuned CLIP encoder exhibit significantly enhanced robustness against adversarial attacks, while preserving semantic meaning of the perturbed images. Notably, Sim-CLIP does not require additional training or fine-tuning of the VLM itself; replacing the original vision encoder with our fine-tuned Sim-CLIP suffices to provide robustness. This work underscores the significance of reinforcing foundational models like CLIP to safeguard the reliability of downstream VLM applications, paving the way for more secure and effective multimodal systems.
Israel defense minister says country will 'settle the score' after Houthi drone attack on Tel Aviv
Israel's defense minister struck an ominous tone Friday after an Iranian-made drone fired by Houthi rebels in Yemen struck Tel Aviv, telling Israeli media that Jerusalem would "settle the score." "I held an operational situation assessment this morning to review the steps required to strengthen our defense arrays in light of events overnight, as well as the intelligence and operational activities required against those responsible for the attack," Israeli Minister of Defense Yoav Gallant said in a statement. "The year 2024 is marked by war. We must be prepared for every scenario and every arena." Israeli Minister of Defense Yoav Gallant sits with defense officials after a Yemen-based Houthi drone strike on Tel Aviv July 19, 2024.
Houthi drone strikes Tel Aviv: How significant is the attack?
Yemen's Houthi group has claimed responsibility for the drone that struck overnight in Tel Aviv, Israel, killing one person and injuring eight. Israeli media identified the dead man as 50-year-old Yevgeny Ferder, who had moved to Israel from Belarus at the beginning of the Russia-Ukraine war. Last night's strike is unique -- it's the first time the group is known to have hit Tel Aviv, though the Houthi have waged a continued campaign against targets they claim are linked to Israel since the ongoing devastating war on Gaza broke out in October. The drone struck in central Tel Aviv in the early hours of Friday morning. The site itself is thought to be close to a number of hotels, many hosting those displaced from Israel's northern border with Lebanon. A US embassy office is also close to the site of the attack.