Government
Evidence-Enhanced Triplet Generation Framework for Hallucination Alleviation in Generative Question Answering
Du, Haowei, Zhang, Huishuai, Zhao, Dongyan
To address the hallucination in generative question answering (GQA) where the answer can not be derived from the document, we propose a novel evidence-enhanced triplet generation framework, EATQA, encouraging the model to predict all the combinations of (Question, Evidence, Answer) triplet by flipping the source pair and the target label to understand their logical relationships, i.e., predict Answer(A), Question(Q), and Evidence(E) given a QE, EA, and QA pairs, respectively. Furthermore, we bridge the distribution gap to distill the knowledge from evidence in inference stage. Our framework ensures the model to learn the logical relation between query, evidence and answer, which simultaneously improves the evidence generation and query answering. In this paper, we apply EATQA to LLama and it outperforms other LLMs-based methods and hallucination mitigation approaches on two challenging GQA benchmarks. Further analysis shows that our method not only keeps prior knowledge within LLM, but also mitigates hallucination and generates faithful answers.
PolicyLR: A Logic Representation For Privacy Policies
Hooda, Ashish, Khandelwal, Rishabh, Chalasani, Prasad, Fawaz, Kassem, Jha, Somesh
Privacy policies are crucial in the online ecosystem, defining how services handle user data and adhere to regulations such as GDPR and CCPA. However, their complexity and frequent updates often make them difficult for stakeholders to understand and analyze. Current automated analysis methods, which utilize natural language processing, have limitations. They typically focus on individual tasks and fail to capture the full context of the policies. We propose PolicyLR, a new paradigm that offers a comprehensive machine-readable representation of privacy policies, serving as an all-in-one solution for multiple downstream tasks. PolicyLR converts privacy policies into a machine-readable format using valuations of atomic formulae, allowing for formal definitions of tasks like compliance and consistency. We have developed a compiler that transforms unstructured policy text into this format using off-the-shelf Large Language Models (LLMs). This compiler breaks down the transformation task into a two-stage translation and entailment procedure. This procedure considers the full context of the privacy policy to infer a complex formula, where each formula consists of simpler atomic formulae. The advantage of this model is that PolicyLR is interpretable by design and grounded in segments of the privacy policy. We evaluated the compiler using ToS;DR, a community-annotated privacy policy entailment dataset. Utilizing open-source LLMs, our compiler achieves precision and recall values of 0.91 and 0.88, respectively. Finally, we demonstrate the utility of PolicyLR in three privacy tasks: Policy Compliance, Inconsistency Detection, and Privacy Comparison Shopping.
Intertwined Biases Across Social Media Spheres: Unpacking Correlations in Media Bias Dimensions
Liu, Yifan, Li, Yike, Wang, Dong
Media bias significantly shapes public perception by reinforcing stereotypes and exacerbating societal divisions. Prior research has often focused on isolated media bias dimensions such as \textit{political bias} or \textit{racial bias}, neglecting the complex interrelationships among various bias dimensions across different topic domains. Moreover, we observe that models trained on existing media bias benchmarks fail to generalize effectively on recent social media posts, particularly in certain bias identification tasks. This shortfall primarily arises because these benchmarks do not adequately reflect the rapidly evolving nature of social media content, which is characterized by shifting user behaviors and emerging trends. In response to these limitations, our research introduces a novel dataset collected from YouTube and Reddit over the past five years. Our dataset includes automated annotations for YouTube content across a broad spectrum of bias dimensions, such as gender, racial, and political biases, as well as hate speech, among others. It spans diverse domains including politics, sports, healthcare, education, and entertainment, reflecting the complex interplay of biases across different societal sectors. Through comprehensive statistical analysis, we identify significant differences in bias expression patterns and intra-domain bias correlations across these domains. By utilizing our understanding of the correlations among various bias dimensions, we lay the groundwork for creating advanced systems capable of detecting multiple biases simultaneously. Overall, our dataset advances the field of media bias identification, contributing to the development of tools that promote fairer media consumption. The comprehensive awareness of existing media bias fosters more ethical journalism, promotes cultural sensitivity, and supports a more informed and equitable public discourse.
Enhancing Robustness of Human Detection Algorithms in Maritime SAR through Augmented Aerial Images to Simulate Weather Conditions
Tjia, Miguel, Kim, Artem, Wijaya, Elaine Wynette, Tefara, Hanna, Zhu, Kevin
Through the utilizations of YOLO, we were able to run different weather conditions and lighting from our augmented dataset for training. YOLO then utilizes CNNs to apply a series of convolutions and pooling layers to the input image, where the convolution layers are able to extract the main features of the image [2]. Through this, our YOLO model is able to learn to differentiate different objects which may considerably improve its accuracy, possibly enhancing the efficiency of SAR operations through enhanced detection accuracy. This paper aims to improve the model's accuracy of human detection in maritime SAR by evaluating a robust datasets containing various elevations and geological locations, as well as through data augmentation which simulates different weather and lighting. We observed that models trained on augmented datasets outperformed their non-augmented counterparts in which the human recall scores ranged from 0.891 to 0.911 with an improvement rate of 3.4% on the YOLOv5l model. Results showed that these models demonstrate greater robustness to real-world conditions in varying of weather, brightness, tint, and contrast.
Low-Budget Simulation-Based Inference with Bayesian Neural Networks
Delaunoy, Arnaud, Bonardeaux, Maxence de la Brassinne, Mishra-Sharma, Siddharth, Louppe, Gilles
Simulation-based inference methods have been shown to be inaccurate in the data-poor regime, when training simulations are limited or expensive. Under these circumstances, the inference network is particularly prone to overfitting, and using it without accounting for the computational uncertainty arising from the lack of identifiability of the network weights can lead to unreliable results. To address this issue, we propose using Bayesian neural networks in low-budget simulation-based inference, thereby explicitly accounting for the computational uncertainty of the posterior approximation. We design a family of Bayesian neural network priors that are tailored for inference and show that they lead to well-calibrated posteriors on tested benchmarks, even when as few as $O(10)$ simulations are available. This opens up the possibility of performing reliable simulation-based inference using very expensive simulators, as we demonstrate on a problem from the field of cosmology where single simulations are computationally expensive. We show that Bayesian neural networks produce informative and well-calibrated posterior estimates with only a few hundred simulations.
Russia, Ukraine trade drone attacks in renewed escalation
Russia has launched several strikes across Ukraine, killing at least five people and wounding several, in an attack that appeared to target energy infrastructure. Ukraine also launched a drone attack on Russia's central region of Saratov, injuring four. The exchange began around midnight on Sunday and continued beyond daybreak on Monday. Ukraine's air force reported multiple groups of Russian drones moving towards its eastern, northern, southern, and central regions, followed by numerous cruise and ballistic missiles. Authorities in at least six Ukrainian regions said blasts had been heard.
Twenty-one civilians killed in Mali drone strikes: Separatist group
At least 21 people, including 11 children, have been killed in drone attacks in the town of Tinzaouaten in northern Mali. A spokesperson for the coalition of Tuareg-majority groups fighting for independence in northern Mali said on Monday that the drones hit a pharmacy and a group of people, leaving dozens wounded. Mali's army confirmed the drone attacks on national television, saying the "precision strikes targeted terrorists". Tinzaouaten has witnessed air attacks before and as recently as July when the Tuareg-led groups claimed to have killed a large number of Malian soldiers and Russian Wagner Group mercenaries. The separatists said they killed at least 47 soldiers and 84 Wagner mercenaries in the July attacks, but the army did not confirm that death toll.
Islamic State supporters turn to AI to bolster online support
Days after a deadly Islamic State attack on a Russian concert hall in March, a man clad in military fatigues and a helmet appeared in an online video, celebrating the assault in which more than 140 people were killed. "The Islamic State delivered a strong blow to Russia with a bloody attack, the fiercest that hit it in years," the man said in Arabic, according to the SITE Intelligence Group, an organisation that tracks and analyses such online content. But the man in the video, which the Thomson Reuters Foundation was not able to view independently, was not real -- he was created using artificial intelligence, according to SITE and other online researchers.
North Korean leader Kim Jong Un oversees suicide drone tests
North Korean leader Kim Jong Un has supervised a test of domestically-developed attack drones, state media KCNA reported. Photos published by North Korean media on Monday showed a white drone with X-shaped tails and wings crashing into and destroying a target resembling South Korea's K-2 main battle tank. Kim, who was pictured at a desk surrounded by advisers, has been modernising his country's military and developing its weapons capabilities amid rising tensions with Washington and Seoul. The North Korean leader supervised the test on a visit to the Drone Institute of North Korea's Academy of Defense Science, KCNA said. Kim said that global trends in military technologies and modern combat showed the importance of drones in war and that Pyongyang's military should be equipped with them "as early as possible".
Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation
Wang, Haoyu, Wu, Bingzhe, Bian, Yatao, Chang, Yongzhe, Wang, Xueqian, Zhao, Peilin
Large Language Models (LLMs) are implicit troublemakers. While they provide valuable insights and assist in problem-solving, they can also potentially serve as a resource for malicious activities. Implementing safety alignment could mitigate the risk of LLMs generating harmful responses. We argue that: even when an LLM appears to successfully block harmful queries, there may still be hidden vulnerabilities that could act as ticking time bombs. To identify these underlying weaknesses, we propose to use a cost value model as both a detector and an attacker. Trained on external or self-generated harmful datasets, the cost value model could successfully influence the original safe LLM to output toxic content in decoding process. For instance, LLaMA-2-chat 7B outputs 39.18% concrete toxic content, along with only 22.16% refusals without any harmful suffixes. These potential weaknesses can then be exploited via prompt optimization such as soft prompts on images. We name this decoding strategy: Jailbreak Value Decoding (JVD), emphasizing that seemingly secure LLMs may not be as safe as we initially believe. They could be used to gather harmful data or launch covert attacks.