Government
AutoHall: Automated Hallucination Dataset Generation for Large Language Models
Cao, Zouying, Yang, Yifei, Zhao, Hai
While Large language models (LLMs) have garnered widespread applications across various domains due to their powerful language understanding and generation capabilities, the detection of non-factual or hallucinatory content generated by LLMs remains scarce. Currently, one significant challenge in hallucination detection is the laborious task of time-consuming and expensive manual annotation of the hallucinatory generation. To address this issue, this paper first introduces a method for automatically constructing model-specific hallucination datasets based on existing fact-checking datasets called AutoHall. Furthermore, we propose a zero-resource and black-box hallucination detection method based on self-contradiction. We conduct experiments towards prevalent open-/closed-source LLMs, achieving superior hallucination detection performance compared to extant baselines. Moreover, our experiments reveal variations in hallucination proportions and types among different models.
Exploring the Impact of Training Data Distribution and Subword Tokenization on Gender Bias in Machine Translation
Iluz, Bar, Limisiewicz, Tomasz, Stanovsky, Gabriel, Mareฤek, David
We study the effect of tokenization on gender bias in machine translation, an aspect that has been largely overlooked in previous works. Specifically, we focus on the interactions between the frequency of gendered profession names in training data, their representation in the subword tokenizer's vocabulary, and gender bias. We observe that female and non-stereotypical gender inflections of profession names (e.g., Spanish "doctora" for "female doctor") tend to be split into multiple subword tokens. Our results indicate that the imbalance of gender forms in the model's training corpus is a major factor contributing to gender bias and has a greater impact than subword splitting. We show that analyzing subword splits provides good estimates of gender-form imbalance in the training data and can be used even when the corpus is not publicly available. We also demonstrate that fine-tuning just the token embedding layer can decrease the gap in gender prediction accuracy between female and male forms without impairing the translation quality.
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
Gou, Zhibin, Shao, Zhihong, Gong, Yeyun, Shen, Yelong, Yang, Yujiu, Duan, Nan, Chen, Weizhu
Recent developments in large language models (LLMs) have been impressive. However, these models sometimes show inconsistencies and problematic behavior, such as hallucinating facts, generating flawed code, or creating offensive and toxic content. Unlike these models, humans typically utilize external tools to cross-check and refine their initial content, like using a search engine for fact-checking, or a code interpreter for debugging. Inspired by this observation, we introduce a framework called CRITIC that allows LLMs, which are essentially "black boxes" to validate and progressively amend their own outputs in a manner similar to human interaction with tools. More specifically, starting with an initial output, CRITIC interacts with appropriate tools to evaluate certain aspects of the text, and then revises the output based on the feedback obtained during this validation process. Comprehensive evaluations involving free-form question answering, mathematical program synthesis, and toxicity reduction demonstrate that CRITIC consistently enhances the performance of LLMs. Meanwhile, our research highlights the crucial importance of external feedback in promoting the ongoing self-improvement of LLMs.
Comfetch: Federated Learning of Large Networks on Constrained Clients via Sketching
Rabbani, Tahseen, Feng, Brandon, Bornstein, Marco, Sang, Kyle Rui, Yang, Yifan, Rajkumar, Arjun, Varshney, Amitabh, Huang, Furong
Federated learning (FL) is a popular paradigm for private and collaborative model training on the edge. In centralized FL, the parameters of a global architecture (such as a deep neural network) are maintained and distributed by a central server/controller to clients who transmit model updates (gradients) back to the server based on local optimization. While many efforts have focused on reducing the communication complexity of gradient transmission, the vast majority of compression-based algorithms assume that each participating client is able to download and train the current and full set of parameters, which may not be a practical assumption depending on the resource constraints of smaller clients such as mobile devices. In this work, we propose a simple yet effective novel algorithm, Comfetch, which allows clients to train large networks using reduced representations of the global architecture via the count sketch, which reduces local computational and memory costs along with bi-directional communication complexity. We provide a nonconvex convergence guarantee and experimentally demonstrate that it is possible to learn large models, such as a deep convolutional network, through federated training on their sketched counterparts. The resulting global models exhibit competitive test accuracy over CIFAR10/100 classification when compared against un-compressed model training.
T-Stochastic Graphs
Previous statistical approaches to hierarchical clustering for social network analysis all construct an "ultrametric" hierarchy. While the assumption of ultrametricity has been discussed and studied in the phylogenetics literature, it has not yet been acknowledged in the social network literature. We show that "non-ultrametric structure" in the network introduces significant instabilities in the existing top-down recovery algorithms. To address this issue, we introduce an instability diagnostic plot and use it to examine a collection of empirical networks. These networks appear to violate the "ultrametric" assumption. We propose a deceptively simple and yet general class of probabilistic models called $\mathbb{T}$-Stochastic Graphs which impose no topological restrictions on the latent hierarchy. To illustrate this model, we propose six alternative forms of hierarchical network models and then show that all six are equivalent to the $\mathbb{T}$-Stochastic Graph model. These alternative models motivate a novel approach to hierarchical clustering that combines spectral techniques with the well-known Neighbor-Joining algorithm from phylogenetic reconstruction. We prove this spectral approach is statistically consistent.
The Slatest for Sept. 29: The Questions Dianne Feinstein Leaves Behind
Dianne Feinstein's office announced Friday morning that she has died at the age of 90, after more than 30 years representing California in the Senate. As her colleagues share memories of her, some huge, high-stakes questions are looming--namely, who will take her seat, and what will become of her spot on the powerful Judiciary Committee. Jim Newell walks us through what seems likely to happen, and what still remains unknown. Plus: The Waves reflects on the senator's legacy of fighting gun violence and conflict with her left-wing constituents. Unless Congress passes a bill to fund the government by Oct. 1, we're cruising for a government shutdown.
Assassin's Creed Mirage preview: Finally, a return to stealth roots
Assassin's Creed Mirage is a dream for stealth kings. People who loved Sam Fisher in Splinter Cell or simply the old Assassin's Creeds will have a tremendous fun in beautiful 9th century Baghdad, our recent hands-on with the game revealed. We throw coins, briefly distract a guard, dart around corners. In that game, we are a bear of a man, with arms like tree trunks as we swing the axe and make the English army tremble. Valhalla also had its moments, but in Mirage there is much more of a hand-built feel.
Pennsylvania Gov. Shapiro noncommittal on future of carbon pricing plan
Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. Gov. Josh Shapiro on Friday remained noncommittal on a strategy to reduce planet-warming greenhouse gases after a task force the Democrat appointed came to an uncertain conclusion over how to make Pennsylvania the first major fossil fuel state to adopt carbon pricing over power plant emissions. The task force sprang from Shapiro questioning his predecessor's use of regulatory authority to join the Regional Greenhouse Gas Initiative, a consortium of 12 eastern states that imposes a price and declining cap on carbon dioxide emissions from power plants. However, the 17-member task force -- comprised of supporters and opponents of former Democratic Gov. Tom Wolf's plan -- could come to no consensus on it. Wolf's regulation allowing Pennsylvania to join the consortium remains hung up in the courts, and Shapiro gave no sign Friday whether he would carry out the consortium's carbon pricing policy should it survive the legal challenge.
The NSA has a new security center specifically for guarding against AI
The National Security Agency (NSA) is starting a dedicated artificial intelligence security center, as reported by AP. This move comes after the government has begun to increasingly rely on AI, integrating multiple algorithms into defense and intelligence systems. The security center will work to protect these systems from theft and sabotage, in addition to safeguarding the country from external AI-based threats. The NSA's recent move toward AI security was announced Thursday by outgoing director General Paul Nakasone. He says that the division will operate underneath the umbrella of the pre-existing Cybersecurity Collaboration Center.
The Creator review: A visually stunning, yet deeply shallow, AI epic
Equal parts Terminator, The Golden Child and The Matrix prequel, The Creator is yet another sci-fi epic about a war between humans and AI, one told by someone who just can't shut up about their time backpacking across Asia. Director Gareth Edwards clearly understands the power of scale and spectacle, something he demonstrated with his indie knockout Monsters, as well as his big-budget efforts, Godzilla and Rogue One. But The Creator, like those films, also suffers from a disjointed narrative, weak characters and a surprisingly shallow exploration of its (potentially interesting!) themes. It's a shame -- at times, the film also proves he can be a genuine visual poet. The Creator stars John David Washington, fresh off of Christopher Nolan's Tenet, as Joshua, an American soldier embedded among a group of AI rebels as a double-agent. When an operation goes wrong early on, he loses his rebel wife Maya (Gemma Chan) and the will to keep fighting the war between the anti-AI West and the AI-loving country of New Asia.