Goto

Collaborating Authors

 Atlantic Ocean


Russia-Ukraine war: List of key events, day 586

Al Jazeera

Ukraine said its air defence systems shot down 16 of about 30 drones launched by Russia on Sunday. Authorities said civilian infrastructure and grain storage warehouses were damaged in the Cherkasy region as well as the southern Mykolaiv and eastern Dnipropetrovsk regions. Russia's defence ministry said its forces' air defences in eastern Ukraine had intercepted five United States-made HIMARS shells, an air-launched JDAM bomb and 37 Ukrainian drones. Kyiv began a counteroffensive in June to retake Ukrainian land occupied by Russia since it launched its full-scale invasion of the country in February 2022. Russia's defence ministry said it shot down six Ukrainian drones over Russian regions and two Ukrainian missiles over Crimea, which Moscow annexed from Ukraine in 2014.


Making Retrieval-Augmented Language Models Robust to Irrelevant Context

arXiv.org Artificial Intelligence

Retrieval-augmented language models (RALMs) hold promise to produce language understanding systems that are are factual, efficient, and up-to-date. An important desideratum of RALMs, is that retrieved information helps model performance when it is relevant, and does not harm performance when it is not. This is particularly important in multi-hop reasoning scenarios, where misuse of irrelevant evidence can lead to cascading errors. However, recent work has shown that retrieval augmentation can sometimes have a negative effect on performance. In this work, we present a thorough analysis on five open-domain question answering benchmarks, characterizing cases when retrieval reduces accuracy. We then propose two methods to mitigate this issue. First, a simple baseline that filters out retrieved passages that do not entail question-answer pairs according to a natural language inference (NLI) model. This is effective in preventing performance reduction, but at a cost of also discarding relevant passages. Thus, we propose a method for automatically generating data to fine-tune the language model to properly leverage retrieved passages, using a mix of relevant and irrelevant contexts at training time. We empirically show that even 1,000 examples suffice to train the model to be robust to irrelevant contexts while maintaining high performance on examples with relevant ones.


Language Model Decoding as Direct Metrics Optimization

arXiv.org Artificial Intelligence

Despite the remarkable advances in language modeling, current mainstream decoding methods still struggle to generate texts that align with human texts across different aspects. In particular, sampling-based methods produce less-repetitive texts which are often disjunctive in discourse, while search-based methods maintain topic coherence at the cost of increased repetition. Overall, these methods fall short in achieving holistic alignment across a broad range of aspects. In this work, we frame decoding from a language model as an optimization problem with the goal of strictly matching the expected performance with human texts measured by multiple metrics of desired aspects simultaneously. The resulting decoding distribution enjoys an analytical solution that scales the input language model distribution via a sequence-level energy function defined by these metrics. And most importantly, we prove that this induced distribution is guaranteed to improve the perplexity on human texts, which suggests a better approximation to the underlying distribution of human texts. To facilitate tractable sampling from this globally normalized distribution, we adopt the Sampling-Importance-Resampling technique. Experiments on various domains and model scales demonstrate the superiority of our method in metrics alignment with human texts and human evaluation over strong baselines.


Knowledge Engineering for Wind Energy

arXiv.org Artificial Intelligence

To this end, vast amounts of data generated by various sources, including sensors and other monitoring systems, need to be effectively structured and represented in a way that can be easily understood and processed by both Artificial Intelligence (AI) systems and humans. The digitalisation of the wind energy sector is one of the key drivers for reducing costs and risks over the whole wind energy project life cycle [2]. The digitalisation process encompasses solutions such as digital twins, decision support systems and AI systems, some of which need to still be developed, in order to contribute to reducing operation and maintenance costs, for increasing the amount of energy delivered, as well as for maximising the efficiency of wind energy systems. In this context, the term Knowledge-Based Systems (KBS) refers to AI systems that formalize knowledge as rules, logical expressions, and conceptualisations [3, 4]. Such systems can be realised as AI-enabled digital twins or decision support systems that rely on databases of knowledge (also referred to as knowledge bases or knowledge graphs), which contain machine-readable facts, rules, and logics about a domain of interest, to assist with problem-solving and decision-making [5].


OceanNet: A principled neural operator-based digital twin for regional oceans

arXiv.org Artificial Intelligence

While data-driven approaches demonstrate great potential in atmospheric modeling and weather forecasting, ocean modeling poses distinct challenges due to complex bathymetry, land, vertical structure, and flow non-linearity. This study introduces OceanNet, a principled neural operator-based digital twin for ocean circulation. OceanNet uses a Fourier neural operator and predictor-evaluate-corrector integration scheme to mitigate autoregressive error growth and enhance stability over extended time scales. A spectral regularizer counteracts spectral bias at smaller scales. OceanNet is applied to the northwest Atlantic Ocean western boundary current (the Gulf Stream), focusing on the task of seasonal prediction for Loop Current eddies and the Gulf Stream meander. Trained using historical sea surface height (SSH) data, OceanNet demonstrates competitive forecast skill by outperforming SSH predictions by an uncoupled, state-of-the-art dynamical ocean model forecast, reducing computation by 500,000 times. These accomplishments demonstrate the potential of physics-inspired deep neural operators as cost-effective alternatives to high-resolution numerical ocean models.


LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

arXiv.org Artificial Intelligence

Studying how people interact with large language models (LLMs) in real-world scenarios is increasingly important due to their widespread use in various applications. In this paper, we introduce LMSYS-Chat-1M, a large-scale dataset containing one million real-world conversations with 25 state-of-the-art LLMs. This dataset is collected from 210K unique IP addresses in the wild on our Vicuna demo and Chatbot Arena website. We offer an overview of the dataset's content, including its curation process, basic statistics, and topic distribution, highlighting its diversity, originality, and scale. We demonstrate its versatility through four use cases: developing content moderation models that perform similarly to GPT-4, building a safety benchmark, training instruction-following models that perform similarly to Vicuna, and creating challenging benchmark questions. We believe that this dataset will serve as a valuable resource for understanding and advancing LLM capabilities. The dataset is publicly available at https://huggingface.co/datasets/lmsys/lmsys-chat-1m.


Towards Mitigating Spurious Correlations in the Wild: A Benchmark and a more Realistic Dataset

arXiv.org Artificial Intelligence

Deep neural networks often exploit non-predictive features that are spuriously correlated with class labels, leading to poor performance on groups of examples without such features. Despite the growing body of recent works on remedying spurious correlations, the lack of a standardized benchmark hinders reproducible evaluation and comparison of the proposed solutions. To address this, we present SpuCo, a python package with modular implementations of state-of-the-art solutions enabling easy and reproducible evaluation of current methods. Using SpuCo, we demonstrate the limitations of existing datasets and evaluation schemes in validating the learning of predictive features over spurious ones. To overcome these limitations, we propose two new vision datasets: (1) SpuCoMNIST, a synthetic dataset that enables simulating the effect of real world data properties e.g. difficulty of learning spurious feature, as well as noise in the labels and features; (2) SpuCoAnimals, a large-scale dataset curated from ImageNet that captures spurious correlations in the wild much more closely than existing datasets. These contributions highlight the shortcomings of current methods and provide a direction for future research in tackling spurious correlations. SpuCo, containing the benchmark and datasets, can be found at https://github.com/BigML-CS-UCLA/SpuCo, with detailed documentation available at https://spuco.readthedocs.io/en/latest/.


Ukraine's drone warfare strategy has brought war home to 'Mother Russia'

FOX News

Former U.S. Defense intel officer Rebekah Koffler discusses additional aid pledged to Ukraine and the U.S.'s decision to launch an unarmed ICBM in California. Last Friday, responding to questions about recent strikes on Crimea, Vice Prime Minister and Minister of Digital Transformation of Ukraine Mykhailo Fedorov, acknowledged, albeit indirectly, that Ukraine was behind them. He also warned that there will be more drone attacks on Russian warships. Drone warfare is a critical component to Ukrainian President Volodymyr Zelenskyy's new asymmetric strategy, likely intended to ensure that Ukrainian armed forces are able to stay in the fight, over the long run, even if they are unable to secure a clear military victory over their highly entrenched opponent. Zelenskyy probably calculates that by systematically employing small scale drone attacks, Ukraine may be able to frustrate, demoralize and exhaust the Russian forces and psychologically dislodge Russian civilians.


Russia-Ukraine war: List of key events, day 581

Al Jazeera

Russia released a video reportedly showing Viktor Sokolov, commander of Russia's Black Sea Fleet in Crimea, at a meeting with Defence Minister Sergei Shoigu and other military top brass a day after Ukrainian special forces claimed he was among dozens of officers killed in an attack on the fleet's Sevastopol naval base. Ukraine said it was clarifying information regarding Sokolov. The United Kingdom's defence ministry said "a dynamic, deep strike battle" was under way in the Black Sea after the Russian Black Sea Fleet suffered a series of major attacks. Kyiv said its air defences destroyed 26 of 38 Russian drones fired overnight but that some of the drones hit the Danube River port of Izmail, damaging more than 30 vehicles and injuring two drivers during a two-hour attack. The drone barrage also prompted the temporary suspension of ferry services to Romania.


Measurement Models For Sailboats Price vs. Features And Regional Areas

arXiv.org Artificial Intelligence

In this study, we investigated the relationship between sailboat technical specifications and their prices, as well as regional pricing influences. Utilizing a dataset encompassing characteristics like length, beam, draft, displacement, sail area, and waterline, we applied multiple machine learning models to predict sailboat prices. The gradient descent model demonstrated superior performance, producing the lowest MSE and MAE. Our analysis revealed that monohulled boats are generally more affordable than catamarans, and that certain specifications such as length, beam, displacement, and sail area directly correlate with higher prices. Interestingly, lower draft was associated with higher listing prices. We also explored regional price determinants and found that the United States tops the list in average sailboat prices, followed by Europe, Hong Kong, and the Caribbean. Contrary to our initial hypothesis, a country's GDP showed no direct correlation with sailboat prices. Utilizing a 50% cross-validation method, our models yielded consistent results across test groups. Our research offers a machine learning-enhanced perspective on sailboat pricing, aiding prospective buyers in making informed decisions.