Goto

Collaborating Authors

 great gatsby


Decomposition-Enhanced Training for Post-Hoc Attributions In Language Models

arXiv.org Artificial Intelligence

Large language models (LLMs) are increasingly used for long-document question answering, where reliable attribution to sources is critical for trust. Existing post-hoc attribution methods work well for extractive QA but struggle in multi-hop, abstractive, and semi-extractive settings, where answers synthesize information across passages. To address these challenges, we argue that post-hoc attribution can be reframed as a reasoning problem, where answers are decomposed into constituent units, each tied to specific context. We first show that prompting models to generate such decompositions alongside attributions improves performance. Building on this, we introduce DecompTune, a post-training method that teaches models to produce answer decompositions as intermediate reasoning steps. We curate a diverse dataset of complex QA tasks, annotated with decompositions by a strong LLM, and post-train Qwen-2.5 (7B and 14B) using a two-stage SFT + GRPO pipeline with task-specific curated rewards. Across extensive experiments and ablations, DecompTune substantially improves attribution quality, outperforming prior methods and matching or exceeding state-of-the-art frontier models.


Gatsby Without the 'E': Crafting Lipograms with LLMs

arXiv.org Artificial Intelligence

Lipograms are a unique form of constrained writing where all occurrences of a particular letter are excluded from the text, typified by the novel Gadsby, which daringly avoids all usage of the letter 'e'. In this study, we explore the power of modern large language models (LLMs) by transforming the novel F. Scott Fitzgerald's The Great Gatsby into a fully 'e'-less text. We experimented with a range of techniques, from baseline methods like synonym replacement to sophisticated generative models enhanced with beam search and named entity analysis. We show that excluding up to 3.6% of the most common letters (up to the letter 'u') had minimal impact on the text's meaning, although translation fidelity rapidly and predictably decays with stronger lipogram constraints. Our work highlights the surprising flexibility of English under strict constraints, revealing just how adaptable and creative language can be.


Meta's AI memorised books verbatim – that could cost it billions

New Scientist

Authors and publishers have filed multiple lawsuits over this issue, and in a new twist, researchers have shown that at least one AI model has not only used popular books in its training data, but also memorised their contents verbatim. But now, researchers have tested multiple models to see how much of that training data they can spit back out verbatim. They found that many models do not retain the exact text of the books in their training data – but one of Meta's models has memorised almost the entirety of certain books. If judges rule against the company, the researchers estimate that this could make Meta liable for at least 1 billion in damages. "That means, on the one hand, that AI models are not just'plagiarism machines', as some have alleged, but it also means that they do more than just learn general relationships between words," says Mark Lemley at Stanford University in California.


IAO Prompting: Making Knowledge Flow Explicit in LLMs through Structured Reasoning Templates

arXiv.org Artificial Intelligence

While Large Language Models (LLMs) demonstrate impressive reasoning capabilities, understanding and validating their knowledge utilization remains challenging. Chain-of-thought (CoT) prompting partially addresses this by revealing intermediate reasoning steps, but the knowledge flow and application remain implicit. We introduce IAO (Input-Action-Output) prompting, a structured template-based method that explicitly models how LLMs access and apply their knowledge during complex reasoning tasks. IAO decomposes problems into sequential steps, each clearly identifying the input knowledge being used, the action being performed, and the resulting output. This structured decomposition enables us to trace knowledge flow, verify factual consistency, and identify potential knowledge gaps or misapplications. Through experiments across diverse reasoning tasks, we demonstrate that IAO not only improves zero-shot performance but also provides transparency in how LLMs leverage their stored knowledge. Human evaluation confirms that this structured approach enhances our ability to verify knowledge utilization and detect potential hallucinations or reasoning errors. Our findings provide insights into both knowledge representation within LLMs and methods for more reliable knowledge application.


Fennec: Fine-grained Language Model Evaluation and Correction Extended through Branching and Bridging

arXiv.org Artificial Intelligence

The rapid advancement of large language models has given rise to a plethora of applications across a myriad of real-world tasks, mainly centered on aligning with human intent. However, the complexities inherent in human intent necessitate a dependence on labor-intensive and time-consuming human evaluation. To alleviate this constraint, we delve into the paradigm of employing open-source large language models as evaluators, aligning with the prevailing trend of utilizing GPT-4. Particularly, we present a step-by-step evaluation framework: \textbf{Fennec}, capable of \textbf{F}ine-grained \textbf{E}valuatio\textbf{N} and correctio\textbf{N} \textbf{E}xtended through bran\textbf{C}hing and bridging. Specifically, the branching operation dissects the evaluation task into various dimensions and granularities, thereby alleviating the challenges associated with evaluation. Concurrently, the bridging operation amalgamates diverse training datasets, augmenting the variety of evaluation tasks. In experimental trials, our 7B model consistently outperforms open-source larger-scale evaluation models across various widely adopted benchmarks in terms of both \textit{Agreement} and \textit{Consistency}, closely approaching the capabilities of GPT-4. We employ the fine-grained correction capabilities induced by the evaluation model to refine multiple model responses, and the results show that the refinement elevates the quality of responses, leading to an improvement of 1-2 points on the MT-Bench. Our code is available at Github\footnote{\url{https://github.com/dropreg/Fennec}}.


Can YOU guess the book? AI reimagines famous houses from literature to celebrate World Book Day

Daily Mail - Science & tech

While your body is lying in bed, your mind may be strolling around the manicured gardens of a manor house or the gritty streets of Victorian London. But now you can see some of the most iconic homes in literature with your own eyes, thanks an artificial intelligence (AI). These include Pemberley House, Mr Darcy's lavish estate in'Pride and Prejudice', and the residence of the world's most famous detective, Sherlock Holmes. Book lovers at Hammonds Furniture used the text-to-image software Midjourney to bring fictional homes to life in celebration of World Book Day 2023 - but how many of them can you guess? Jay Gatsby's mansion in'The Great Gatsby' (pictured) is described as a'colossal affair by any standard' and an'imitation of some Hôtel de Ville in Normandy' Daisy Buchanan's estate in'The Great Gatsby' (pictured) is described as a'cheerful red-and-white Georgian Colonial mansion', as well as'elaborate', 'bright' and'rosy-coloured' The above two houses are depictions of those from'The Great Gatsby', a novel set in 1922 that follows the life of mysterious millionaire Jay Gatsby.


Watson the word lover - Watson

#artificialintelligence

Or, more precisely, I love words and stories. I often find myself poring over specific passages and savoring them, searching for nuance and pattern. In many ways, Watson is a word lover too, finding connections and insights within unstructured data. Usually, Watson reads medical journals or legal regulations, but I have often wondered what Watson might find in some of my favorite novels. So, I decided to see how Watson would'read' a scene from The Great Gatsby by F. Scott Fitzgerald.


Watson the word lover - Watson

#artificialintelligence

Or, more precisely, I love words and stories. I often find myself poring over specific passages and savoring them, searching for nuance and pattern. In many ways, Watson is a word lover too, finding connections and insights within unstructured data. Usually, Watson reads medical journals or legal regulations, but I have often wondered what Watson might find in some of my favorite novels. So, I decided to see how Watson would'read' a scene from The Great Gatsby by F. Scott Fitzgerald.