Goto

Collaborating Authors

 Government


REASONS: A benchmark for REtrieval and Automated citationS Of scieNtific Sentences using Public and Proprietary LLMs

arXiv.org Artificial Intelligence

Automatic citation generation for sentences in a document or report is paramount for intelligence analysts, cybersecurity, news agencies, and education personnel. In this research, we investigate whether large language models (LLMs) are capable of generating references based on two forms of sentence queries: (a) Direct Queries, LLMs are asked to provide author names of the given research article, and (b) Indirect Queries, LLMs are asked to provide the title of a mentioned article when given a sentence from a different article. To demonstrate where LLM stands in this task, we introduce a large dataset called REASONS comprising abstracts of the 12 most popular domains of scientific research on arXiv. From around 20K research articles, we make the following deductions on public and proprietary LLMs: (a) State-of-the-art, often called anthropomorphic GPT-4 and GPT-3.5, suffers from high pass percentage (PP) to minimize the hallucination rate (HR). When tested with Perplexity.ai (7B), they unexpectedly made more errors; (b) Augmenting relevant metadata lowered the PP and gave the lowest HR; (c) Advance retrieval-augmented generation (RAG) using Mistral demonstrates consistent and robust citation support on indirect queries and matched performance to GPT-3.5 and GPT-4. The HR across all domains and models decreased by an average of 41.93%, and the PP was reduced to 0% in most cases. In terms of generation quality, the average F1 Score and BLEU were 68.09% and 57.51%, respectively; (d) Testing with adversarial samples showed that LLMs, including the Advance RAG Mistral, struggle to understand context, but the extent of this issue was small in Mistral and GPT-4-Preview. Our study contributes valuable insights into the reliability of RAG for automated citation generation tasks.


Interpretability Needs a New Paradigm

arXiv.org Machine Learning

Interpretability is the study of explaining models in understandable terms to humans. At present, interpretability is divided into two paradigms: the intrinsic paradigm, which believes that only models designed to be explained can be explained, and the post-hoc paradigm, which believes that black-box models can be explained. At the core of this debate is how each paradigm ensures its explanations are faithful, i.e., true to the model's behavior. This is important, as false but convincing explanations lead to unsupported confidence in artificial intelligence (AI), which can be dangerous. This paper's position is that we should think about new paradigms while staying vigilant regarding faithfulness. First, by examining the history of paradigms in science, we see that paradigms are constantly evolving. Then, by examining the current paradigms, we can understand their underlying beliefs, the value they bring, and their limitations. Finally, this paper presents 3 emerging paradigms for interpretability. The first paradigm designs models such that faithfulness can be easily measured. Another optimizes models such that explanations become faithful. The last paradigm proposes to develop models that produce both a prediction and an explanation.


Kathy Hochul Really Outdid Herself With This Gaffe

Slate

This is Totally Normal Quote of the Day, a feature highlighting a statement from the news that exemplifies just how extremely normal everything has become. "Right now, we have young Black kids growing up in the Bronx who don't even know what the word computer is. They don't know, they don't know these things." If recent polling is any indication, it seems pretty clear to everyone that New York Gov. Kathy Hochul could be doing a better job of running her state. If I may offer a little advice, maybe she could start by understanding New York City a little better--and by being just a biiiiiiit less racist.


TikTok and ByteDance sue US to block law forcing sale of the app

The Guardian

TikTok and its parent company ByteDance have sued to block a law signed by Joe Biden just weeks ago that would force the sale of the short video app or ban it from the US. The companies filed a lawsuit on Tuesday against the US government in the court of appeals for the District of Columbia, arguing the law is unconstitutional and violates free speech protections. Signed by the president on 24 April as part of a broader foreign aid package, the law gives China's ByteDance until 19 January 2025 to sell TikTok to an approved buyer. If it does not, the US would prohibit app stores from offering TikTok and bar internet hosting services from supporting TikTok. The companies argue in the suit that the divestiture required by the bill "is simply not commercially, legally, or technically possible. "There is no question: the Act (law) will force a shutdown of TikTok by January 19, 2025, silencing the 170 million Americans who use the platform to communicate in ways that cannot be replicated elsewhere," the suit said. The suit confirmed previous reports that ByteDance would not sell TikTok without the powerful recommendation algorithm that has fueled the platform's success. The Chinese government "has made clear that it would not permit a divestment of the recommendation engine that is a key to the success of TikTok in the United States", the suit said. The potential for a ban of TikTok has been escalating since Donald Trump first unsuccessfully attempted to block it in 2020. Critics of TikTok have expressed worry that the platform's China-based parent company could collect sensitive user data and censor content that goes against the Chinese government โ€“ claims TikTok denies. Our US morning briefing breaks down the key stories of the day, telling you what's happening and why it matters Amid the political fallout, TikTok spent more than 2bn to implement measures to protect the data of US users, according to the suit. The suit also highlighted additional commitments the company made in a 90-page draft National Security Agreement developed through negotiations with the Committee on Foreign Investment in the United States (CFIUS), an interagency committee, chaired by the US Treasury Department, that reviews foreign investments in American businesses that implicate national security concerns. CFIUS had been in talks with TikTok to find solutions, though the agreement included TikTok agreeing to a "shut-down option" that would give the US government the authority to suspend TikTok in the US if it violated some obligations", according to the suit.


Using high-tech drones, Russia is pressing aerial advantage against beleaguered Ukrainian artillery

FOX News

Polish Foreign Minister Radosล‚aw Sikorski provides his analysis of the Russia-Ukraine war as it marks its third Easter, the passing of the foreign aid package and his expectations for the upcoming NATO summit. Rumbling out of its forest hideout, the hulking German-supplied howitzer has only a few minutes to fire before slipping back under cover to evade Russian surveillance in the skies above. Across the hills and valleys of the east, Ukrainian artillery units play a cat-and-mouse game with Russian drones hunting high-value artillery weapons such as this self-propelled Panzerhaubitze 2000. Moscow's troops have stepped up ground attacks along the 621-mile front in the south and east of Ukraine, threatening some of the industrialized Donetsk region's last big cities held by Kyiv more than two years after Russia's full-scale invasion. Counterbattery efforts are crucial to suppressing enemy fire that rains on Ukrainian lines and artillery units, and paves the way for Russian advances.


AIhub coffee corner: Responsible and trustworthy AI

AIHub

This month, our trustees tackle the topic of trustworthy AI. Joining the conversation this time are: Tom Dietterich (Oregon State University), Sabine Hauert (University of Bristol), and Sarit Kraus (Bar-Ilan University). Sabine Hauert: There was a big trustworthy autonomous systems conference a few weeks back in London, and on the back of that they've launched a big responsible AI portfolio. I know Europe has been focusing on trustworthiness and how responsible these algorithms are. Deploying these systems in a responsible way is something that people are thinking about more and more. It was interesting at that conference because, while a lot of it had to do with ethics, interfacing with humans and thinking holistically about these algorithms, there was also a strong military track discussing how you make military tools trustworthy. I always find it quite interesting that trustworthiness and responsible AI mean completely different things to different communities.


Gov. Hochul says she 'misspoke' when she said some 'black kids' don't know the word 'computer'

FOX News

New York Gov. Kathy Hochul tells an audience at the Milken Institute that there are "young black kids in the Bronx" who "don't even know what the word'computer' is." (Credit: Governor Kathy Hochul) New York Gov. Kathy Hochul apologized this week after saying there are black kids in the Bronx who don't know what the word "computer" means. Hochil made the remarks during an address at the Milken Institute Global Conference in Los Angeles, California. "Now what we have is the money to build a phenomenal super computer that is gonna be accessible to the researchers in New York, college students, will attract more federal grants, and this is how we lay down the mark," Hochul said. "No state has done this. In fact, I talk to a lot of other people who say, 'I wish my governor had thought of that first.' I say, 'No no, this is New York. We like to be first,' with all due respect to you from other states."


From gun gear to prosthetic leg covers, volunteers boost Ukraine's army

Al Jazeera

Chernihiv, Ukraine โ€“ In combat, the speed of loading your assault gun's magazine is a matter of life and death. Sometimes, a soldier has to load the rounds in sub-zero temperatures, with wet or wounded hands. An improperly loaded magazine could jam the rifle and get its owner killed. A simple and inexpensive accessory โ€“ magazine speed loaders known among gun enthusiasts as "magloaders" or "thumb savers" โ€“ pushes the magazine's top so that the rounds are inserted with little or no pressure. Widely available in the United States, the speed loaders were virtually unknown in Ukraine until Take Back Our History, a volunteer group in the northern city of Chernihiv, began manufacturing and supplying them to the military, free of charge.


Defense think tank MITRE to build AI supercomputer with Nvidia

Washington Post - Technology News

A key supplier to the Pentagon and U.S. intelligence agencies is building a 20 million supercomputer with buzzy chipmaker Nvidia to speed deployment of artificial-intelligence capabilities across the U.S. federal government, the MITRE think tank said Tuesday.


New allometric models for the USA create a step-change in forest carbon estimation, modeling, and mapping

arXiv.org Artificial Intelligence

The United States national forest inventory (NFI) serves as the foundation for forest aboveground biomass (AGB) and carbon accounting across the nation. These data enable design-based estimates of forest carbon stocks and stock-changes at state and regional levels, but also serve as inputs to model-based approaches for characterizing forest carbon stocks and stock-changes at finer resolutions. Although NFI tree and plot-level data are often treated as truth in these models, they are in fact estimates based on regional species-group models known collectively as the Component Ratio Method (CRM). In late 2023 the Forest Inventory and Analysis (FIA) program introduced a new National Scale Volume and Biomass Estimators (NSVB) system to replace CRM nationwide and offer more precise and accurate representations of forest AGB and carbon. Given the prevalence of model-based AGB studies relying on FIA, there is concern about the transferability of methods from CRM to NSVB models, as well as the comparability of existing CRM AGB products (e.g. maps) to new and forthcoming NSVB AGB products. To begin addressing these concerns we compared previously published CRM AGB maps to new maps produced using identical methods with NSVB AGB reference data. Our results suggest that models relying on passive satellite imagery (e.g. Landsat) provide acceptable estimates of point-in-time NSVB AGB and carbon stocks, but fail to accurately quantify growth in mature closed-canopy forests. We highlight that existing estimates, models, and maps based on FIA reference data are no longer compatible with NSVB, and recommend new methods as well as updated models and maps for accommodating this step-change. Our collective ability to adopt NSVB in our modeling and mapping workflows will help us provide the most accurate spatial forest carbon data possible in order to better inform local management and decision making.