Goto

Collaborating Authors

 Government


ConQRet: Benchmarking Fine-Grained Evaluation of Retrieval Augmented Argumentation with LLM Judges

arXiv.org Artificial Intelligence

Computational argumentation, which involves generating answers or summaries for controversial topics like abortion bans and vaccination, has become increasingly important in today's polarized environment. Sophisticated LLM capabilities offer the potential to provide nuanced, evidence-based answers to such questions through Retrieval-Augmented Argumentation (RAArg), leveraging real-world evidence for high-quality, grounded arguments. However, evaluating RAArg remains challenging, as human evaluation is costly and difficult for complex, lengthy answers on complicated topics. At the same time, re-using existing argumentation datasets is no longer sufficient, as they lack long, complex arguments and realistic evidence from potentially misleading sources, limiting holistic evaluation of retrieval effectiveness and argument quality. To address these gaps, we investigate automated evaluation methods using multiple fine-grained LLM judges, providing better and more interpretable assessments than traditional single-score metrics and even previously reported human crowdsourcing. To validate the proposed techniques, we introduce ConQRet, a new benchmark featuring long and complex human-authored arguments on debated topics, grounded in real-world websites, allowing an exhaustive evaluation across retrieval effectiveness, argument quality, and groundedness. We validate our LLM Judges on a prior dataset and the new ConQRet benchmark. Our proposed LLM Judges and the ConQRet benchmark can enable rapid progress in computational argumentation and can be naturally extended to other complex retrieval-augmented generation tasks.


From Defects to Demands: A Unified, Iterative, and Heuristically Guided LLM-Based Framework for Automated Software Repair and Requirement Realization

arXiv.org Artificial Intelligence

As software systems evolve, developers face a dual challenge: maintaining correctness by fixing bugs and continuously adapting functionality to meet new user demands. Traditional software engineering processes rely heavily on human developers to interpret requirements, fix errors, and ensure correctness against specifications. With the advancement of Large Language Models (LLMs) adept at code generation, the opportunity arises to shift portions of these responsibilities onto machine-driven processes. However, simply prompting an LLM to solve a complex programming task--be it eliminating a subtle bug or implementing a new feature--often falls short. Complex codebases exceed the model's context window, specification details are not always fully captured in a single prompt, and correctness requires iterative refinement guided by tests, analysis, and verification. This paper proposes a holistic, iterative framework that enables an LLM to evolve a codebase from an initial, potentially buggy state to one that satisfies not only pre-existing correctness criteria but also newly introduced feature demands. Key contributions include: 1. Unified Framework for Bug-to-Demand Resolution: We present a method by which the LLM iteratively refines code, starting from an imperfect state (with known or unknown bugs) and incrementally adjusting the codebase to meet a set of evolving functional and nonfunctional requirements introduced over time.


OCEAN: Open-World Contrastive Authorship Identification

arXiv.org Artificial Intelligence

In an era where cyberattacks increasingly target the software supply chain, the ability to accurately attribute code authorship in binary files is critical to improving cybersecurity measures. We propose OCEAN, a contrastive learning-based system for function-level authorship attribution. OCEAN is the first framework to explore code authorship attribution on compiled binaries in an open-world and extreme scenario, where two code samples from unknown authors are compared to determine if they are developed by the same author. To evaluate OCEAN, we introduce new realistic datasets: CONAN, to improve the performance of authorship attribution systems in real-world use cases, and SNOOPY, to increase the robustness of the evaluation of such systems. We use CONAN to train our model and evaluate on SNOOPY, a fully unseen dataset, resulting in an AUROC score of 0.86 even when using high compiler optimizations. We further show that CONAN improves performance by 7% compared to the previously used Google Code Jam dataset. Additionally, OCEAN outperforms previous methods in their settings, achieving a 10% improvement over state-of-the-art SCS-Gan in scenarios analyzing source code. Furthermore, OCEAN can detect code injections from an unknown author in a software update, underscoring its value for securing software supply chains.


Project Report: Requirements for a Social Robot as an Information Provider in the Public Sector

arXiv.org Artificial Intelligence

Is it possible to integrate a humanoid social robot into the work processes or customer care in an official environment, e.g. in municipal offices? If so, what could such an application scenario look like and what skills would the robot need to have when interacting with human customers? What are requirements for this kind of interactions? We have devised an application scenario for such a case, determined the necessary or desirable capabilities of the robot, developed a corresponding robot application and carried out initial tests and evaluations in a project together with the Kiel City Council. One of the most important insights gained in the project was that a humanoid robot with natural language processing capabilities based on large language models as well as human-like gestures and posture changes (animations) proved to be much more preferred by users compared to standard browser-based solutions on tablets for an information system in the City Council. Furthermore, we propose a connection of the ACT-R cognitive architecture with the robot, where an ACT-R model is used in interaction with the robot application to cognitively process and enhance a dialogue between human and robot.


CompCap: Improving Multimodal Large Language Models with Composite Captions

arXiv.org Artificial Intelligence

How well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as charts, posters, or screenshots, rather than being captured directly by a camera. While CIs are prevalent in real-world applications, recent MLLM developments have primarily focused on interpreting natural images (NIs). Our research reveals that current MLLMs face significant challenges in accurately understanding CIs, often struggling to extract information or perform complex reasoning based on these images. We find that existing training data for CIs are mostly formatted for question-answer tasks (e.g., in datasets like ChartQA and ScienceQA), while high-quality image-caption datasets, critical for robust vision-language alignment, are only available for NIs. To bridge this gap, we introduce Composite Captions (CompCap), a flexible framework that leverages Large Language Models (LLMs) and automation tools to synthesize CIs with accurate and detailed captions. Using CompCap, we curate CompCap-118K, a dataset containing 118K image-caption pairs across six CI types. We validate the effectiveness of CompCap-118K by supervised fine-tuning MLLMs of three sizes: xGen-MM-inst.-4B and LLaVA-NeXT-Vicuna-7B/13B. Empirical results show that CompCap-118K significantly enhances MLLMs' understanding of CIs, yielding average gains of 1.7%, 2.0%, and 2.9% across eleven benchmarks, respectively.


'A needle in a haystack:' How AI is helping uncover abandoned oil wells

Popular Science

The continental United States is jam-packed with reminders of our ravenous oil appetite. Since the 1850s, there have been an estimated 3.5 million oil and gas wells drilled across the country. Many of those were abandoned after the companies running them ran out of business or otherwise ceased operating. These forgotten fossil fuel artifacts, referred to officially as "undocumented orphan wells" (UOWs) are often left behind without meaningful efforts taken to safely seal them. Unplugged orphan wells can leak out dangerous methane, oil, and other chemicals for years which can pollute the air and potentially contaminate nearby water sources.


Authorities stress 'no known threat to public safety' following unusual drones near Trump Bedminster club

FOX News

Officials are still investigating unusual drone activity that has been reported in recent weeks in New Jersey. The FAA set temporary restrictions above Trump National Golf Club in Bedminster in response. Authorities investigating the unusual drone activity observed several times in northern New Jersey in recent days, including the vicinity of President-elect Trump's Bedminster golf club, continue to stress that there is no threat to public safety. Multiple videos show drones flying in Somerset and Morris counties over the past few weeks, including Dec. 1 and Dec. 3. In a video from Nov. 25, a Morris County resident named Mike Walsh spotted drones flying over Black River Middle School in Chester.


The US Department of Defense is investing in deepfake detection

MIT Technology Review

"This work represents a significant step forward in strengthening our information advantage as we combat sophisticated disinformation campaigns and synthetic-media threats," says Bustamante. Hive was chosen out of a pool of 36 companies to test its deepfake detection and attribution technology with the DOD. The contract could enable the department to detect and counter AI deception at scale. "This is the evolution of cyberwarfare." Hive's technology has been trained on a large amount of content, some AI-generated and some not.


China's Tencent seems to have AI chips banned by US export controls

New Scientist

Tencent is one of China's largest technology companies Chinese tech giant Tencent doesn't seem to be affected by US export bans of computer chips that are crucial to the development of artificial intelligence systems – but even if such bans were more stringent, they may not be able to slow the country's AI advancement. Ritwik Gupta and his colleagues at the University of California, Berkeley, have analysed publications released by researchers at Tencent about the firm's latest models, including its Hunyuan AI models. The team's findings suggest that, in recent months, Tencent has publicly described using…


Tencent seems unaffected by US AI chip export ban, research shows

New Scientist

Tencent is one of China's largest technology companies Chinese tech giant Tencent doesn't seem to be affected by US export bans of computer chips that are crucial to the development of artificial intelligence systems – but even if such bans were more stringent, they may not be able to slow the country's AI advancement, US research shows. Ritwik Gupta and his colleagues at the University of California, Berkeley, have analysed publications released by researchers at Tencent about the firm's latest models, including its Hunyuan AI models. The team's findings suggest that, in recent months, Tencent has…