Goto

Collaborating Authors

 Government


Context versus Prior Knowledge in Language Models

arXiv.org Artificial Intelligence

To answer a question, language models often need to integrate prior knowledge learned during pretraining and new information presented in context. We hypothesize that models perform this integration in a predictable way across different questions and contexts: models will rely more on prior knowledge for questions about entities (e.g., persons, places, etc.) that they are more familiar with due to higher exposure in the training corpus, and be more easily persuaded by some contexts than others. To formalize this problem, we propose two mutual information-based metrics to measure a model's dependency on a context and on its prior about an entity: first, the persuasion score of a given context represents how much a model depends on the context in its decision, and second, the susceptibility score of a given entity represents how much the model can be swayed away from its original answer distribution about an entity. We empirically test our metrics for their validity and reliability. Finally, we explore and find a relationship between the scores and the model's expected familiarity with an entity, and provide two use cases to illustrate their benefits.


Evaluating Supply Chain Resilience During Pandemic Using Agent-based Simulation

arXiv.org Artificial Intelligence

Recent pandemics have highlighted vulnerabilities in our global economic systems, especially supply chains. Possible future pandemic raises a dilemma for businesses owners between short-term profitability and long-term supply chain resilience planning. In this study, we propose a novel agent-based simulation model integrating extended Susceptible-Infected-Recovered (SIR) epidemiological model and supply and demand economic model to evaluate supply chain resilience strategies during pandemics. Using this model, we explore a range of supply chain resilience strategies under pandemic scenarios using in silico experiments. We find that a balanced approach to supply chain resilience performs better in both pandemic and non-pandemic times compared to extreme strategies, highlighting the importance of preparedness in the form of a better supply chain resilience. However, our analysis shows that the exact supply chain resilience strategy is hard to obtain for each firm and is relatively sensitive to the exact profile of the pandemic and economic state at the beginning of the pandemic. As such, we used a machine learning model that uses the agent-based simulation to estimate a near-optimal supply chain resilience strategy for a firm. The proposed model offers insights for policymakers and businesses to enhance supply chain resilience in the face of future pandemics, contributing to understanding the trade-offs between short-term gains and long-term sustainability in supply chain management before and during pandemics.


Mitigating Biases of Large Language Models in Stance Detection with Calibration

arXiv.org Artificial Intelligence

Large language models (LLMs) have achieved remarkable progress in many natural language processing tasks. However, our experiment reveals that, in stance detection tasks, LLMs may generate biased stances due to sentiment-stance spurious correlations and preference towards certain individuals and topics, thus harming their performance. Therefore, in this paper, we propose to Mitigate Biases of LLMs in stance detection with Calibration (MB-Cal). To be specific, a novel calibration network is devised to calibrate potential bias in the stance prediction of LLMs. Further, to address the challenge of effectively learning bias representations and the difficulty in the generalizability of debiasing, we construct counterfactual augmented data. This approach enhances the calibration network, facilitating the debiasing and out-of-domain generalization. Experimental results on in-target and zero-shot stance detection tasks show that the proposed MB-Cal can effectively mitigate biases of LLMs, achieving state-of-the-art results.


Schr\"{o}dinger Bridge with Quadratic State Cost is Exactly Solvable

arXiv.org Machine Learning

Schr\"odinger bridge is a diffusion process that steers a given distribution to another in a prescribed time while minimizing the effort to do so. It can be seen as the stochastic dynamical version of the optimal mass transport, and has growing applications in generative diffusion models and stochastic optimal control. In this work, we propose a regularized variant of the Schr\"odinger bridge with a quadratic state cost-to-go that incentivizes the optimal sample paths to stay close to a nominal level. Unlike the conventional Schr\"odinger bridge, the regularization induces a state-dependent rate of killing and creation of probability mass, and its solution requires determining the Markov kernel of a reaction-diffusion partial differential equation. We derive this Markov kernel in closed form. Our solution recovers the heat kernel in the vanishing regularization (i.e., diffusion without reaction) limit, thereby recovering the solution of the conventional Schr\"odinger bridge. Our results enable the use of dynamic Sinkhorn recursion for computing the Schr\"odinger bridge with a quadratic state cost-to-go, which would otherwise be challenging to use in this setting. We deduce properties of the new kernel and explain its connections with certain exactly solvable models in quantum mechanics.


A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery

arXiv.org Artificial Intelligence

In many scientific fields, large language models (LLMs) have revolutionized the way with which text and other modalities of data (e.g., molecules and proteins) are dealt, achieving superior performance in various applications and augmenting the scientific discovery process. Nevertheless, previous surveys on scientific LLMs often concentrate on one to two fields or a single modality. In this paper, we aim to provide a more holistic view of the research landscape by unveiling cross-field and cross-modal connections between scientific LLMs regarding their architectures and pre-training techniques. To this end, we comprehensively survey over 250 scientific LLMs, discuss their commonalities and differences, as well as summarize pre-training datasets and evaluation tasks for each field and modality. Moreover, we investigate how LLMs have been deployed to benefit scientific discovery. Resources related to this survey are available at https://github.com/yuzhimanhua/Awesome-Scientific-Language-Models.


How's this for a bombshell – the US must make AI its next Manhattan Project John Naughton

The Guardian

Ten years ago, the Oxford philosopher Nick Bostrom published Superintelligence, a book exploring how superintelligent machines could be created and what the implications of such technology might be. One was that such a machine, if it were created, would be difficult to control and might even take over the world in order to achieve its goals (which in Bostrom's celebrated thought experiment was to make paperclips). The book was a big seller, triggering lively debates but also attracting a good deal of disagreement. Critics complained that it was based on a simplistic view of "intelligence", that it overestimated the likelihood of superintelligent machines emerging any time soon and that it failed to suggest credible solutions for the problems that it had raised. But it had the great merit of making people think about a possibility that had hitherto been confined to the remoter fringes of academia and sci-fi. Now, 10 years later, comes another shot at the same target.


Ransomware Attacks Are Getting Worse

WIRED

Despite years worth of efforts to eliminate the scourge of ransomware targeting schools, hospitals, and critical infrastructure worldwide, experts are warning that the crisis is only heating up, with criminal gangs growing ever more aggressive in their tactics. The threat of real-world violence now looms, some experts warn, as the data stolen grows increasingly sensitive and millions in potential profits hang in the balance. "We know where your CEO lives," read a message reportedly received by one victim. Attacks targeting the medical sector are blooming in response to the 44 million payout by Change Healthcare this March. United States lawmakers and intelligence officials are circling their wagons following the revelation of Israel's involvement in a malign influence campaign that targeted US voters--an attempt by America's Middle East ally to artificially boost support for an increasingly unpopular war that was kicked off by Hamas' unprecedented Oct. 7th attack.


How AI cops are ALREADY patrolling Britain's streets: From 'the eye in the sky' to facial recognition surveillance in supermarkets - the Orwellian technologies being used to tackle crime

Daily Mail - Science & tech

In his classic novel, 1984, George Orwell imagined how Britain might one day become a totalitarian surveillance state. Yet as Orwell's novel celebrates its 75th anniversary this month, British police are already deploying technologies that would put Big Brother to shame. From the facial recognition cameras watching you shop to the algorithms predicting crimes before they happen, these tools feel as if they've been ripped from the pages of science fiction. But there is nothing fictional about the AI cops already patrolling Britain's streets - and experts say there is only more to come. Jake Hufurt, head of research and investigations at Big Brother Watch, warned MailOnline: 'We're sleepwalking into a high-tech police state.'


Israeli-deployed AI in Gaza likely helps IDF reduce civilian casualties, expert says

FOX News

After loudly touting the use of artificial intelligence (AI) during their 11-day conflict against Hamas in 2021, the Israel Defense Forces (IDF) have been fairly tight-lipped about the AI systems they've employed in the post-Oct. Numerous media outlets have speculated that Israel's AI platforms are being used recklessly, but Blaise Misztal, Vice President for Policy at the Jewish Institute for National Security of America (JINSA), told Fox News Digital that he believes Israel is using AI-powered drone swarms, mapping drones and targeting systems as a means to minimize civilian casualties as they seek out Hamas terrorists hiding among the populace or holed up in tunnel systems laced beneath civilian architecture. Misztal says that available evidence implies drones are a "near constant companion for ground troops as they're maneuvering through Gaza," with the IDF telling JINSA researchers that "each unit has its own mini-Air Force" supporting troop movements. A number of AI-powered drones may be mapping the underground tunnels built below Gaza, or protecting those who are traversing them as they seek out terrorists or hostages. Iris, a ground-based, throwable unit manufactured by Elbit Systems "can enter small and confined spaces, above or underground, to explore hazardous areas while relaying intelligence and reconnaissance information in real-time."


Applications of Generative AI in Healthcare: algorithmic, ethical, legal and societal considerations

arXiv.org Artificial Intelligence

Generative AI is rapidly transforming medical imaging and text analysis, offering immense potential for enhanced diagnosis and personalized care. However, this transformative technology raises crucial ethical, societal, and legal questions. This paper delves into these complexities, examining issues of accuracy, informed consent, data privacy, and algorithmic limitations in the context of generative AI's application to medical imaging and text. We explore the legal landscape surrounding liability and accountability, emphasizing the need for robust regulatory frameworks. Furthermore, we dissect the algorithmic challenges, including data biases, model limitations, and workflow integration. By critically analyzing these challenges and proposing responsible solutions, we aim to foster a roadmap for ethical and responsible implementation of generative AI in healthcare, ensuring its transformative potential serves humanity with utmost care and precision.