Goto

Collaborating Authors

 Personal


Contrastive Decoding: Open-ended Text Generation as Optimization

arXiv.org Artificial Intelligence

Given a language model (LM), maximum probability is a poor decoding objective for open-ended generation, because it produces short and repetitive text. On the other hand, sampling can often produce incoherent text that drifts from the original topics. We propose contrastive decoding (CD), a reliable decoding approach that optimizes a contrastive objective subject to a plausibility constraint. The contrastive objective returns the difference between the likelihood under a large LM (called the expert, e.g. OPT-13B) and a small LM (called the amateur, e.g. OPT-125M), and the constraint ensures that the outputs are plausible. CD is inspired by the fact that the failures of larger LMs (e.g., repetition, incoherence) are even more prevalent in smaller LMs, and that this difference signals which texts should be preferred. CD requires zero additional training, and produces higher quality text than decoding from the larger LM alone. It also works across model scales (OPT-13B and GPT2-1.5B) and significantly outperforms four strong decoding algorithms (e.g., nucleus, top-k) in automatic and human evaluations across wikipedia, news and story domains.


Most Women Ignore Their "Reply Guys." Then There Are These People.

Slate

In May, Sydney Leathers confessed to her tens of thousands of Twitter followers that she was smitten. Where'd she meet the guy? Not on a dating app, or through friends, but in the last place she ever expected to find a real connection: her mentions. "Still can't believe I fell in love with one of my reply guys. Apparently, things had progressed since December, when she last posted about him: "I had sex with someone who started as my reply guy and I hope this doesn't inspire confidence in the rest of you because frankly your replies are not that good," she wrote. Leathers is a writer, adult performer, and startup employee whose name you may recognize from her part in the Anthony Weiner sexting scandal--this wasn't exactly her first brush with online flirtation. But it was her first time falling for a reply guy, or someone who was, effectively, a fan. The term "reply guy" emerged on Twitter about five years ago to describe the behavior of a certain subset of people, usually with very few social media followers of their own, who stake out space in the mentions of prominent users. They can be counted on to reply promptly and frequently to the tweets of whomever they've chosen as their object of devotion, and they often seek attention by nitpicking, mansplaining, joke one-upping, and harassing them. Because of this, reply guys--who can also be girls, or people of any gender--are generally understood to be pathetic creatures, without a chance in hell of getting said person to like their replies, much less return their affections. So the revelation that this gambit actually worked for someone is โ€ฆ pretty noteworthy. Reply guy success stories may be happening more than we realize. Abby, a 25-year-old in Brooklyn who runs a meme page on Instagram with several thousand followers, told me that she got frisky with one of her reply guys last year. "I'm not the only person that I know that has hooked up with reply guys," she said. "It's not as uncommon as you might think." Now, Leathers' Twitter feed is a monument to her relationship, by turns adorable and lewd. "This definitely caught me by surprise," she told me. "But it's been the best, happiest relationship I've had." To attain this goal, a reply guy's first challenge is to stand out from the crowd. The meme account Abby is the admin for is about politics, so she likes when a guy can show not just that he's hot, but that they share a political sensibility. "I have to be attracted to them," she said. "And they have to have some sort of compelling thing to say." "I feel like I've never more than mildly acknowledged a reply guy before now," she said. "I generally don't even follow them back." But when her now-boyfriend started responding to her tweets last year after discovering her through a winding path that involved the singer of the band Eve 6, she took notice. "I'd seen him reply to my stuff a few times.


Designing a Direct Feedback Loop between Humans and Convolutional Neural Networks through Local Explanations

arXiv.org Artificial Intelligence

The local explanation provides heatmaps on images to explain how Convolutional Neural Networks (CNNs) derive their output. Due to its visual straightforwardness, the method has been one of the most popular explainable AI (XAI) methods for diagnosing CNNs. Through our formative study (S1), however, we captured ML engineers' ambivalent perspective about the local explanation as a valuable and indispensable envision in building CNNs versus the process that exhausts them due to the heuristic nature of detecting vulnerability. Moreover, steering the CNNs based on the vulnerability learned from the diagnosis seemed highly challenging. To mitigate the gap, we designed DeepFuse, the first interactive design that realizes the direct feedback loop between a user and CNNs in diagnosing and revising CNN's vulnerability using local explanations. DeepFuse helps CNN engineers to systemically search "unreasonable" local explanations and annotate the new boundaries for those identified as unreasonable in a labor-efficient manner. Next, it steers the model based on the given annotation such that the model doesn't introduce similar mistakes. We conducted a two-day study (S2) with 12 experienced CNN engineers. Using DeepFuse, participants made a more accurate and "reasonable" model than the current state-of-the-art. Also, participants found the way DeepFuse guides case-based reasoning can practically improve their current practice. We provide implications for design that explain how future HCI-driven design can move our practice forward to make XAI-driven insights more actionable.


AIhub coffee corner: AI risks, pause letters and the ensuing discourse

AIHub

This month, in light of the recent prominent discussions relating to perceived AI risks, we consider the pause letters and risk statements, the debate around existential threats, and how this discourse could impact the field and public perceptions. Joining the discussion this time are: Sanmay Das (George Mason University), Tom Dietterich (Oregon State University), Sabine Hauert (University of Bristol), Sarit Kraus (Bar-Ilan University), Anna Tahovskรก (Czech Technical University), and Oskar von Stryk (Technische Universitรคt Darmstadt). Sabine Hauert: In today's discussion we're going to talk about potential AI risks and the recent discourse around existential threats. Does anyone have any hot reactions? How do you feel about the discourse of existential threat? Tom Dietterich: I agree with Emily Bender and a lot of the critics that it's a distraction and a diversion from thinking about the more immediate threats.


VisKoP: Visual Knowledge oriented Programming for Interactive Knowledge Base Question Answering

arXiv.org Artificial Intelligence

We present Visual Knowledge oriented Programming platform (VisKoP), a knowledge base question answering (KBQA) system that integrates human into the loop to edit and debug the knowledge base (KB) queries. VisKoP not only provides a neural program induction module, which converts natural language questions into knowledge oriented program language (KoPL), but also maps KoPL programs into graphical elements. KoPL programs can be edited with simple graphical operators, such as dragging to add knowledge operators and slot filling to designate operator arguments. Moreover, VisKoP provides auto-completion for its knowledge base schema and users can easily debug the KoPL program by checking its intermediate results. To facilitate the practical KBQA on a million-entity-level KB, we design a highly efficient KoPL execution engine for the back-end. Experiment results show that VisKoP is highly efficient and user interaction can fix a large portion of wrong KoPL programs to acquire the correct answer. The VisKoP online demo https://demoviskop.xlore.cn (Stable release of this paper) and https://viskop.xlore.cn (Beta release with new features), highly efficient KoPL engine https://pypi.org/project/kopl-engine, and screencast video https://youtu.be/zAbJtxFPTXo are now publicly available.


Convergence of Communications, Control, and Machine Learning for Secure and Autonomous Vehicle Navigation

arXiv.org Artificial Intelligence

Connected and autonomous vehicles (CAVs) can reduce human errors in traffic accidents, increase road efficiency, and execute various tasks ranging from delivery to smart city surveillance. Reaping these benefits requires CAVs to autonomously navigate to target destinations. To this end, each CAV's navigation controller must leverage the information collected by sensors and wireless systems for decision-making on longitudinal and lateral movements. However, enabling autonomous navigation for CAVs requires a convergent integration of communication, control, and learning systems. The goal of this article is to explicitly expose the challenges related to this convergence and propose solutions to address them in two major use cases: Uncoordinated and coordinated CAVs. In particular, challenges related to the navigation of uncoordinated CAVs include stable path tracking, robust control against cyber-physical attacks, and adaptive navigation controller design. Meanwhile, when multiple CAVs coordinate their movements during navigation, fundamental problems such as stable formation, fast collaborative learning, and distributed intrusion detection are analyzed. For both cases, solutions using the convergence of communication theory, control theory, and machine learning are proposed to enable effective and secure CAV navigation. Preliminary simulation results are provided to show the merits of proposed solutions.


A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

arXiv.org Artificial Intelligence

The development of large language models (LLMs) such as ChatGPT has brought a lot of attention recently. However, their evaluation in the benchmark academic datasets remains under-explored due to the difficulty of evaluating the generative outputs produced by this model against the ground truth. In this paper, we aim to present a thorough evaluation of ChatGPT's performance on diverse academic datasets, covering tasks like question-answering, text summarization, code generation, commonsense reasoning, mathematical problem-solving, machine translation, bias detection, and ethical considerations. Specifically, we evaluate ChatGPT across 140 tasks and analyze 255K responses it generates in these datasets. This makes our work the largest evaluation of ChatGPT in NLP benchmarks. In short, our study aims to validate the strengths and weaknesses of ChatGPT in various tasks and provide insights for future research using LLMs. We also report a new emergent ability to follow multi-query instructions that we mostly found in ChatGPT and other instruction-tuned models. Our extensive evaluation shows that even though ChatGPT is capable of performing a wide variety of tasks, and may obtain impressive performance in several benchmark datasets, it is still far from achieving the ability to reliably solve many challenging tasks. By providing a thorough assessment of ChatGPT's performance across diverse NLP tasks, this paper sets the stage for a targeted deployment of ChatGPT-like LLMs in real-world applications.


Mitigating the Learning Bias towards Repetition by Self-Contrastive Training for Open-Ended Generation

arXiv.org Artificial Intelligence

Despite the huge progress in myriad generation tasks, pretrained language models (LMs) such as GPT2 still tend to generate repetitive texts with maximization-based decoding algorithms for open-ended generation. We attribute their overestimation of token-level repetition probabilities to the learning bias: LMs capture simple repetitive patterns faster with the MLE loss. We propose self-contrastive training to penalize the output of a premature checkpoint of the same model when it incorrectly predicts repetition, which is shown to mitigate repetition effectively while maintaining fluency on two datasets. Furthermore, we find that LMs use longer-range dependencies to predict repetitive tokens than non-repetitive ones, which may be the cause of sentence-level repetition loops.


Scenario-Based Motion Planning with Bounded Probability of Collision

arXiv.org Artificial Intelligence

Robots will increasingly operate near humans that introduce uncertainties in the motion planning problem due to their complex nature. Typically, chance constraints are introduced in the planner to optimize performance while guaranteeing probabilistic safety. However, existing methods do not consider the actual probability of collision for the planned trajectory, but rather its marginalization, that is, the independent collision probabilities for each planning step and/or dynamic obstacle, resulting in conservative trajectories. To address this issue, we introduce a novel real-time capable method termed Safe Horizon MPC, that explicitly constrains the joint probability of collision with all obstacles over the duration of the motion plan. This is achieved by reformulating the chance-constrained planning problem using scenario optimization and predictive control. Our method is less conservative than state-of-the-art approaches, applicable to arbitrary probability distributions of the obstacles' trajectories, computationally tractable and scalable. We demonstrate our proposed approach using a mobile robot and an autonomous vehicle in an environment shared with humans.


A Comprehensive Survey of Artificial Intelligence Techniques for Talent Analytics

arXiv.org Artificial Intelligence

In today's competitive and fast-evolving business environment, it is a critical time for organizations to rethink how to make talent-related decisions in a quantitative manner. Indeed, the recent development of Big Data and Artificial Intelligence (AI) techniques have revolutionized human resource management. The availability of large-scale talent and management-related data provides unparalleled opportunities for business leaders to comprehend organizational behaviors and gain tangible knowledge from a data science perspective, which in turn delivers intelligence for real-time decision-making and effective talent management at work for their organizations. In the last decade, talent analytics has emerged as a promising field in applied data science for human resource management, garnering significant attention from AI communities and inspiring numerous research efforts. To this end, we present an up-to-date and comprehensive survey on AI technologies used for talent analytics in the field of human resource management. Specifically, we first provide the background knowledge of talent analytics and categorize various pertinent data. Subsequently, we offer a comprehensive taxonomy of relevant research efforts, categorized based on three distinct application-driven scenarios: talent management, organization management, and labor market analysis. In conclusion, we summarize the open challenges and potential prospects for future research directions in the domain of AI-driven talent analytics.