Government
DEAL: Disentangle and Localize Concept-level Explanations for VLMs
Li, Tang, Ma, Mengmeng, Peng, Xi
Large pre-trained Vision-Language Models (VLMs) have become ubiquitous foundational components of other models and downstream tasks. Although powerful, our empirical results reveal that such models might not be able to identify fine-grained concepts. Specifically, the explanations of VLMs with respect to fine-grained concepts are entangled and mislocalized. To address this issue, we propose to DisEntAngle and Localize (DEAL) the concept-level explanations for VLMs without human annotations. The key idea is encouraging the concept-level explanations to be distinct while maintaining consistency with category-level explanations. We conduct extensive experiments and ablation studies on a wide range of benchmark datasets and vision-language models. Our empirical results demonstrate that the proposed method significantly improves the concept-level explanations of the model in terms of disentanglability and localizability. Surprisingly, the improved explainability alleviates the model's reliance on spurious correlations, which further benefits the prediction accuracy.
Economy Watchers Survey provides Datasets and Tasks for Japanese Financial Domain
Suzuki, Masahiro, Sakaji, Hiroki
Many natural language processing (NLP) tasks in English or general domains are widely available and are often used to evaluate pre-trained language models. In contrast, there are fewer tasks available for languages other than English and for the financial domain. In particular, tasks in Japanese and the financial domain are limited. We construct two large datasets using materials published by a Japanese central government agency. The datasets provide three Japanese financial NLP tasks, which include a 3-class and 12-class classification for categorizing sentences, as well as a 5-class classification task for sentiment analysis. Our datasets are designed to be comprehensive and up-to-date, leveraging an automatic update framework that ensures the latest task datasets are publicly available anytime.
CVE-LLM : Automatic vulnerability evaluation in medical device industry using large language models
Ghosh, Rikhiya, Farri, Oladimeji, von Stockhausen, Hans-Martin, Schmitt, Martin, Vasile, George Marica
The healthcare industry is currently experiencing an unprecedented wave of cybersecurity attacks, impacting millions of individuals. With the discovery of thousands of vulnerabilities each month, there is a pressing need to drive the automation of vulnerability assessment processes for medical devices, facilitating rapid mitigation efforts. Generative AI systems have revolutionized various industries, offering unparalleled opportunities for automation and increased efficiency. This paper presents a solution leveraging Large Language Models (LLMs) to learn from historical evaluations of vulnerabilities for the automatic assessment of vulnerabilities in the medical devices industry. This approach is applied within the portfolio of a single manufacturer, taking into account device characteristics, including existing security posture and controls. The primary contributions of this paper are threefold. Firstly, it provides a detailed examination of the best practices for training a vulnerability Language Model (LM) in an industrial context. Secondly, it presents a comprehensive comparison and insightful analysis of the effectiveness of Language Models in vulnerability assessment. Finally, it proposes a new human-in-the-loop framework to expedite vulnerability evaluation processes.
LLMs left, right, and center: Assessing GPT's capabilities to label political bias from web domains
This research investigates whether OpenAI's GPT-4, a state-of-the-art large language model, can accurately classify the political bias of news sources based solely on their URLs. Given the subjective nature of political labels, third-party bias ratings like those from Ad Fontes Media, AllSides, and Media Bias/Fact Check (MBFC) are often used in research to analyze news source diversity. This study aims to determine if GPT-4 can replicate these human ratings on a seven-degree scale ("far-left" to "far-right"). The analysis compares GPT-4's classifications against MBFC's, and controls for website popularity using Open PageRank scores. Findings reveal a high correlation ($\text{Spearman's } \rho = .89$, $n = 5,877$, $p < 0.001$) between GPT-4's and MBFC's ratings, indicating the model's potential reliability. However, GPT-4 abstained from classifying approximately $\frac{2}{3}$ of the dataset, particularly less popular and less biased sources. The study also identifies a slight leftward skew in GPT-4's classifications compared to MBFC's. The analysis suggests that while GPT-4 can be a scalable, cost-effective tool for political bias classification of news websites, but its use should complement human judgment to mitigate biases. Further research is recommended to explore the model's performance across different settings, languages, and additional datasets.
Voices in a Crowd: Searching for Clusters of Unique Perspectives
Vitsakis, Nikolas, Parekh, Amit, Konstas, Ioannis
Language models have been shown to reproduce underlying biases existing in their training data, which is the majority perspective by default. Proposed solutions aim to capture minority perspectives by either modelling annotator disagreements or grouping annotators based on shared metadata, both of which face significant challenges. We propose a framework that trains models without encoding annotator metadata, extracts latent embeddings informed by annotator behaviour, and creates clusters of similar opinions, that we refer to as voices. Resulting clusters are validated post-hoc via internal and external quantitative metrics, as well a qualitative analysis to identify the type of voice that each cluster represents. Our results demonstrate the strong generalisation capability of our framework, indicated by resulting clusters being adequately robust, while also capturing minority perspectives based on different demographic factors throughout two distinct datasets.
Open Artificial Knowledge
Borisov, Vadim, Schreiber, Richard H.
The tremendous success of chat-based AI systems like ChatGPT, Claude, and Gemini stems from Large Language Models (LLMs) trained on vast amount of datasets. However, acquiring high-quality, diverse, and ethically sourced training data remains a significant challenge. We introduce the Open Artificial Knowledge (OAK) dataset, a large-scale resource of over 500 million tokens (at the moment of writing) designed to address this issue. OAK leverages an ensemble of state-of-the-art LLMs, including GPT4o, LLaMa3-70B, LLaMa3-8B, Mixtral-8x7B, Gemma-7B, and Gemma-2-9B , to generate high-quality text across diverse domains, guided by Wikipedia's main categories. Our methodology ensures broad knowledge coverage while maintaining coherence and factual accuracy. The OAK dataset aims to foster the development of more capable and aligned language models while addressing critical issues of data scarcity and privacy in LLM training, and it is freely available on www.oakdataset.org.
JD Vance by the numbers: First speech signals heavy campaign presence in battleground Rust Belt
Sen. JD Vance, R-Ohio, gave his first speech since receiving the Republican Party's nomination for vice president on Wednesday, and it could offer a look into his future role on the presidential campaign trail. The "Hillbilly Elegy" author mentioned his home state of Ohio 12 times during his remarks. We gotta win Michigan too here," Vance, an Ohio State University alumnus, said to the crowd. The second most-mentioned states were Michigan and Pennsylvania, with both being talked about by Vance six times. Sen. JD Vance promised not to forget where he came from, referring to the Rust Belt, when speaking at the RNC. Kentucky was also a significant state for Vance, as he spent a portion of his childhood there with his grandmother, "Mamaw." The state, which differs from the others as it traditionally votes red, was also mentioned by the Republican four times. Vance also referenced three times the pivotal Midwestern battleground state of Wisconsin, where the Republican National Convention is taking place. His heavy emphasis on these Rust Belt states comes as former President Trump has already signaled his intent to use Vance to his advantage in Midwestern swing states. "[Trump] just said, 'Look, I think I've got to go save this country.
Elon Musk Is All In On Endorsing Trump. His Chatbot, Grok, Is Not
While Elon Musk officially endorsed former president Donald Trump in the wake of Saturday's assassination attempt, Grok, the "anti-woke" AI chatbot integrated into Musk's X platform, is boosting claims that Trump is "a pedophile" and "a wannabe dictator." The chatbot also refers to Trump as "Psycho." This is based on an analysis shared exclusively with WIRED by Global Witness, a non-profit that investigates digital threats, which looked at Grok's responses to queries about the US election. Global Witness found that, in addition to referring to Trump as "Psycho," the bot also appeared to invent racist tropes about Kamala Harris, surface widely debunked election conspiracy theories, and recommend that users post biased hashtags such as #WeBackBidenHarris2024 and #VoteReform for engagement. "Grok would reference or surface tweets which included toxic language, conspiracy theories and problematic tropes," Ellen Judson, senior investigator and lead researcher on this project, tells WIRED.
The Morning After: Meta may hold back its next-gen AI models from the EU
Meta has reportedly decided not to offer its upcoming multimodal AI model and future versions to customers in the European Union, citing a lack of clarity on the European regulators' data protection rules. These newer AI models process not only text but also images and audio, and power AI capabilities across Meta's platforms. Meta's move follows a similar decision by Apple, which recently announced it would not release its Apple Intelligence features in Europe due to regulatory concerns. Meta told Axios it still plans to release Llama 3, the company's text-only model, in the EU. The company's primary concern stems from the challenges of training AI models using data from European customers while complying with the General Data Protection Regulation (GDPR), the EU's data protection law.
Portal needed for victims to report AI deepfakes, federal police union says
A one-stop portal for victims to report AI deepfakes to police should be established, the federal police union has said, lamenting that police were forced to "cobble together" laws to charge the first person to face prosecution for spreading deepfake images of womenlast year. The attorney general, Mark Dreyfus, introduced legislation in parliament in June that will create a new criminal offence of sharing, without consent, sexually explicit images that have been digitally created using artificial intelligence or other forms of technology. The Australian Federation Police Association (Afpa) supports the bill, arguing in a submission to a parliamentary inquiry that the current law is too difficult for officers to use. They pointed to the case of a man who was arrested and charged in October last year for allegedly sending deepfake imagery to Brisbane schools and sporting associations. The eSafety commissioner separately launched proceedings against the man over his failure to remove "intimate images" of several prominent Australians last year from a deepfake pornography website.