Goto

Collaborating Authors

 Atlantic Ocean


Machine-learning prediction of tipping and collapse of the Atlantic Meridional Overturning Circulation

arXiv.org Artificial Intelligence

Department of Physics, Arizona State University, Tempe, Arizona 85287, USA (Dated: February 26, 2024) Recent research on the Atlantic Meridional Overturning Circulation (AMOC) raised concern about its potential collapse through a tipping point due to the climate-change caused increase in the freshwater input into the North Atlantic. The predicted time window of collapse is centered about the middle of the century and the earliest possible start is approximately two years from now. More generally, anticipating a tipping point at which the system transitions from one stable steady state to another is relevant to a broad range of fields. We develop a machine-learning approach to predicting tipping in noisy dynamical systems with a time-varying parameter and test it on a number of systems including the AMOC, ecological networks, an electrical power system, and a climate model. For the AMOC, our prediction based on simulated fingerprint data and real data of the sea surface temperature places the time window of a potential collapse between the years 2040 and 2065.


Coercing LLMs to do and reveal (almost) anything

arXiv.org Artificial Intelligence

It has recently been shown that adversarial attacks on large language models (LLMs) can'jailbreak' the model into making harmful statements. In this work, we argue that the spectrum of adversarial attacks on LLMs is much larger than merely jailbreaking. We provide a broad overview of possible attack surfaces and attack goals. Based on a series of concrete examples, we discuss, categorize and systematize attacks that coerce varied unintended behaviors, such as misdirection, model control, denial-of-service, or data extraction. We analyze these attacks in controlled experiments, and find that many of them stem from the practice of pre-training LLMs with coding capabilities, as well as the continued existence of strange'glitch' tokens in common LLM vocabularies that should be removed for security reasons. We conclude that the spectrum of adversarial attacks on LLMs is much broader than previously thought, and that the security of these models must be addressed through a comprehensive understanding of their capabilities and limitations.")] Some figures and tables below contain profanity or offensive text.


The Importance of Architecture Choice in Deep Learning for Climate Applications

arXiv.org Artificial Intelligence

Machine Learning has become a pervasive tool in climate science applications. However, current models fail to address nonstationarity induced by anthropogenic alterations in greenhouse emissions and do not routinely quantify the uncertainty of proposed projections. In this paper, we model the Atlantic Meridional Overturning Circulation (AMOC) which is of major importance to climate in Europe and the US East Coast by transporting warm water to these regions, and has the potential for abrupt collapse. We can generate arbitrarily extreme climate scenarios through arbitrary time scales which we then predict using neural networks. Our analysis shows that the AMOC is predictable using neural networks under a diverse set of climate scenarios. Further experiments reveal that MLPs and Deep Ensembles can learn the physics of the AMOC instead of imitating its progression through autocorrelation. With quantified uncertainty, an intriguing pattern of "spikes" before critical points of collapse in the AMOC casts doubt on previous analyses that predicted an AMOC collapse within this century. Our results show that Bayesian Neural Networks perform poorly compared to more dense architectures and care should be taken when applying neural networks to nonstationary scenarios such as climate projections. Further, our results highlight that big NN models might have difficulty in modeling global Earth System dynamics accurately and be successfully applied in nonstationary climate scenarios due to the physics being challenging for neural networks to capture.


Diffusion Visual Counterfactual Explanations

Neural Information Processing Systems

Visual Counterfactual Explanations (VCEs) are an important tool to understand the decisions of an image classifier. They are "small" but "realistic" semantic changes of the image changing the classifier decision. Current approaches for the generation of VCEs are restricted to adversarially robust models and often contain non-realistic artefacts, or are limited to image classification problems with few classes. In this paper, we overcome this by generating Diffusion Visual Counterfactual Explanations (DVCEs) for arbitrary ImageNet classifiers via a diffusion process. Two modifications to the diffusion process are key for our DVCEs: first, an adaptive parameterization, whose hyperparameters generalize across images and models, together with distance regularization and late start of the diffusion process, allow us to generate images with minimal semantic changes to the original ones but different classification. Second, our cone regularization via an adversarially robust model ensures that the diffusion process does not converge to trivial non-semantic changes, but instead produces realistic images of the target class which achieve high confidence by the classifier.


Can We Verify Step by Step for Incorrect Answer Detection?

arXiv.org Artificial Intelligence

Chain-of-Thought (CoT) prompting has marked a significant advancement in enhancing the reasoning capabilities of large language models (LLMs). Previous studies have developed various extensions of CoT, which focus primarily on enhancing end-task performance. In addition, there has been research on assessing the quality of reasoning chains in CoT. This raises an intriguing question: Is it possible to predict the accuracy of LLM outputs by scrutinizing the reasoning chains they generate? To answer this research question, we introduce a benchmark, R2PE, designed specifically to explore the relationship between reasoning chains and performance in various reasoning tasks spanning five different domains. This benchmark aims to measure the falsehood of the final output of LLMs based on the reasoning steps. To make full use of information in multiple reasoning chains, we propose the process discernibility score (PDS) framework that beats the answer-checking baseline by a large margin. Concretely, this resulted in an average of 5.1% increase in the F1 score across all 45 subsets within R2PE. We further demonstrate our PDS's efficacy in advancing open-domain QA accuracy. Data and code are available at https://github.com/XinXU-USTC/R2PE.


Missile strike on Belgorod, Russia, kills 6, injures 18

FOX News

Seven people, including three children, were killed in a Russian drone attack on a gas station in the Ukrainian city of Kharkiv on Saturday. A missile strike on the Russian city of Belgorod near the Ukraine border on Thursday killed six people, including a child, and injured 18 others, a Russian official said. It was the latest in exchanges of long-range missile and rocket fire in Russia's war on Ukraine. Hours earlier, Russia fired two dozen cruise and ballistic missiles at a broad area of Ukraine, hitting multiple regions after a midnight strike in Ukraine's northeast killed five people in an apartment building, authorities said. Five of the 18 people injured in Belgorod, a city of around 340,000 people, were children, regional Gov. Vyacheslav Gladkov said on Telegram.


Kyiv aims to use more Ukrainian drones; Trump, Biden clash on NATO

Al Jazeera

Ukraine changed its military leadership and announced a change of tactics in the past week, as a vote in the US Senate brought renewed hope of US aid for the embattled country. Ukrainian President Volodymyr Zelenskyy appointed ground forces commander Oleksandr Syrskii as commander-in-chief of the armed forces on February 8. Zelenskyy reportedly asked the outgoing Valery Zaluzhny to "continue to be part of the team", without specifying what that meant. "We stood against a vile and powerful enemy. Endured together," wrote Zaluzhny, an immensely popular general who stopped Russia's invasion in February 2022 and ordered a counterattack in August that year, which claimed more than 1,500sq km (580sq miles) Since then, Ukrainian forces have become bogged down in positional warfare. A counteroffensive last summer failed to achieve its goal of cutting the Russian front in two.


Russia-Ukraine war: List of key events, day 723

Al Jazeera

Ukraine said it critically damaged the Caesar Kunikov, a Russian landing warship, off occupied Crimea, in a drone attack, the latest blow to the Russian navy's Black Sea Fleet. Ukraine said the ship, one of Russia's newest vessels, had a crew of 87 and had taken part in wars in Georgia and Syria as well as Ukraine. There was no official comment from Russia on the attack. Newly-appointed Ukrainian armed forces chief Oleksandr Syrskyii visited troops fighting around the key flashpoint of Avdiivka on the eastern front line, and described the situation as "extremely complex and stressful". Syrskyii, who was accompanied by Defence Minister Rustem Umerov, said Russian forces had "a numerical advantage in personnel".


Generative Representational Instruction Tuning

arXiv.org Artificial Intelligence

All text-based language problems can be reduced to either generation or embedding. Current models only perform well at one or the other. We introduce generative representational instruction tuning (GRIT) whereby a large language model is trained to handle both generative and embedding tasks by distinguishing between them through instructions. Compared to other open models, our resulting GritLM 7B sets a new state of the art on the Massive Text Embedding Benchmark (MTEB) and outperforms all models up to its size on a range of generative tasks. By scaling up further, GritLM 8x7B outperforms all open generative language models that we tried while still being among the best embedding models. Notably, we find that GRIT matches training on only generative or embedding data, thus we can unify both at no performance loss. Among other benefits, the unification via GRIT speeds up Retrieval-Augmented Generation (RAG) by > 60% for long documents, by no longer requiring separate retrieval and generation models. Models, code, etc. are freely available at https://github.com/ContextualAI/gritlm.


Russian landing ship Caesar Kunikov sunk off Crimea, says Ukraine

BBC News

There was no confirmation from Russia's navy that the Caesar Kunikov had been sunk in the Black Sea, merely that six Ukrainian drones had been destroyed. Video appearing to show the aftermath of the Ukrainian attack was uploaded only recently, BBC Verify confirmed.