Media
A Neural-Symbolic Approach Towards Identifying Grammatically Correct Sentences
Textual content around us is growing on a daily basis. Numerous articles are being written as we speak on online newspapers, blogs, or social media. Similarly, recent advances in the AI field, like language models or traditional classic AI approaches, are utilizing all the above to improve their learned representation to tackle NLP challenges with human-like accuracy. It is commonly accepted that it is crucial to have access to well-written text from valid sources to tackle challenges like text summarization, question-answering, machine translation, or even pronoun resolution. For instance, to summarize well, one needs to select the most important sentences in order to concatenate them to form the summary. However, what happens if we do not have access to well-formed English sentences or even non-valid sentences? Despite the importance of having access to well-written sentences, figuring out ways to validate them is still an open area of research. To address this problem, we present a simplified way to validate English sentences through a novel neural-symbolic approach. Lately, neural-symbolic approaches have triggered an increasing interest towards tackling various NLP challenges, as they are demonstrating their effectiveness as a central component in various AI systems. Through combining Classic with Modern AI, which involves the blending of grammatical and syntactical rules with language models, we effectively tackle the Corpus of Linguistic Acceptability (COLA), a task that shows whether or not a sequence of words is an English grammatical sentence. Among others, undertaken experiments effectively show that blending symbolic and non-symbolic systems helps the former provide insights about the latter's accuracy results.
Understanding and Mitigating Spurious Correlations in Text Classification with Neighborhood Analysis
Chew, Oscar, Lin, Hsuan-Tien, Chang, Kai-Wei, Huang, Kuan-Hao
Recent research has revealed that deep learning models have a tendency to leverage spurious correlations that exist in the training set but may not hold true in general circumstances. For instance, a sentiment classifier may erroneously learn that the token performances is commonly associated with positive movie reviews. Relying on these spurious correlations degrades the classifiers performance when it deploys on out-of-distribution data. In this paper, we examine the implications of spurious correlations through a novel perspective called neighborhood analysis. The analysis uncovers how spurious correlations lead unrelated words to erroneously cluster together in the embedding space. Driven by the analysis, we design a metric to detect spurious tokens and also propose a family of regularization methods, NFL (doN't Forget your Language) to mitigate spurious correlations in text classification. Experiments show that NFL can effectively prevent erroneous clusters and significantly improve the robustness of classifiers.
Moviegoers weigh in on demands made by striking actors, writers: 'I'm all for it''
Hollywood shuts down as actors and writers hit the picket line in the first industry-wide strike in over 60 years. Fox News spoke to Americans from New York, Texas, Tennessee and Wisconsin to get their thoughts on Tinseltown going dark. Hollywood actors joined screenwriters in their months long strike against studios, streaming services and production companies represented by the Alliance of Motion Picture and Television Producers (AMPTP) on Thursday, marking the first time in over six decades that the two unions have been on strike at the same time. Many moviegoers and TV fanatics who Fox News Digital spoke to realize the impact the strike will have on their favorite shows and movies as production grinds to a halt. Since May, writers, represented by the Writers Guild of America (WGA) have been on strike, asking for a guaranteed number of writers per room, increased pay, and regulated use of artificial intelligence (AI) in the writing process.
The New em Mission: Impossible /em Reveals That the Franchise Has Always Had an Unlikely Big Bad
Over the course of the six movies and nearly 30 years leading up to this one, Mission: Impossible's Ethan Hunt has fought double agents and shadowy terrorist networks, scaled skyscrapers, and thrown himself out of planes. But he's never fought an adversary like the Entity, the rogue artificial intelligence he takes on in Dead Reckoning Part One. For one thing, it has no physical form, which means Tom Cruise can't catch it no matter how fast he runs. And for another, the Entity doesn't just want to defeat Ethan: It wants to replace him. In a briefing of the U.S. top intelligence officials in which exposition is passed from one actor to the next like a red-hot baton, one alphabet-agency higher-up describes the Entity as a "godless, stateless, amoral" being that can infiltrate any system in the world--not unlike the Impossible Mission Force itself, which, though nominally a branch of the U.S. government, doesn't take orders from the military-industrial complex so much as consider its suggestions.
Latin America looks to use AI to narrow the technology gap, but fear of 'risks' could accelerate the divide
GOP Rep. Nancy Mace spoke exclusively with Fox News Digital about her thoughts on the rapidly advancing AI sector, as Congress races to get ahead of the burgeoning technology. The incredible potential of artificial intelligence (AI) threatens to accelerate the technological divide that runs deep throughout Latin America, an expert told Fox News Digital. "The use of AI is going to increase the quality of life of all those countries for sure, but what the gap could be, and who is generating those other AI, and then who's controlling the data that's feeding that AI?," Jordi Albo-Canals, a Chilean native and CSO and co-founder of Lighthouse Disruptive Innovation Group, told Fox News Digital. Albo-Canals suggested that with the right regulation, the tech could help to actually close the gap, but for some parts of the region, the access to the technology remains limited. Different countries in Latin America have approached the burgeoning AI technology in different ways, but each formed by their own experience with technology so far: Mexican media company Radio Formula introduced an AI news anchor called NAT in March, who presented short news capsules -- the first of its kind, according to the company, Mexico Business News reported.
Cruz shoots down Schumer effort to regulate AI: 'More harm than good'
Fox News anchor Julie Banderas reacts to the vice president's gaffe and CNN calling Dylan Mulvaney a man on'Jesse Watters Primetime.' EXCLUSIVE: Sen. Ted Cruz is criticizing Democrats for what he believes is a rush to regulate the artificial intelligence sector, and says new rules for AI would be a drag on the U.S. in the critical tech race against China. "I am concerned China is investing heavily in AI. I'm also concerned that Democrats want to impose such stringent regulations on the development of AI that it stifles innovation in the United States, and allows China to take the lead," Cruz told Fox News Digital after a classified briefing on AI and national security this week. "That would be a generational mistake," he said of the Democrats' effort.
A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image Models
Allingham, James Urquhart, Ren, Jie, Dusenberry, Michael W, Gu, Xiuye, Cui, Yin, Tran, Dustin, Liu, Jeremiah Zhe, Lakshminarayanan, Balaji
Contrastively trained text-image models have the remarkable ability to perform zero-shot classification, that is, classifying previously unseen images into categories that the model has never been explicitly trained to identify. However, these zero-shot classifiers need prompt engineering to achieve high accuracy. Prompt engineering typically requires hand-crafting a set of prompts for individual downstream tasks. In this work, we aim to automate this prompt engineering and improve zero-shot accuracy through prompt ensembling. In particular, we ask "Given a large pool of prompts, can we automatically score the prompts and ensemble those that are most suitable for a particular downstream dataset, without needing access to labeled validation data?". We demonstrate that this is possible. In doing so, we identify several pathologies in a naive prompt scoring method where the score can be easily overconfident due to biases in pre-training and test data, and we propose a novel prompt scoring method that corrects for the biases. Using our proposed scoring method to create a weighted average prompt ensemble, our method outperforms equal average ensemble, as well as hand-crafted prompts, on ImageNet, 4 of its variants, and 11 fine-grained classification benchmarks, all while being fully automatic, optimization-free, and not requiring access to labeled validation data.
Unifying Structure Reasoning and Language Model Pre-training for Complex Reasoning
Wang, Siyuan, Wei, Zhongyu, Xu, Jiarong, Li, Taishan, Fan, Zhihao
Recent pre-trained language models (PLMs) equipped with foundation reasoning skills have shown remarkable performance on downstream complex tasks. However, the significant structure reasoning skill has been rarely studied, which involves modeling implicit structure information within the text and performing explicit logical reasoning over them to deduce the conclusion. This paper proposes a unified learning framework that combines explicit structure reasoning and language pre-training to endow PLMs with the structure reasoning skill. It first identifies several elementary structures within contexts to construct structured queries and performs step-by-step reasoning along the queries to identify the answer entity. The fusion of textual semantics and structure reasoning is achieved by using contextual representations learned by PLMs to initialize the representation space of structures, and performing stepwise reasoning on this semantic representation space. Experimental results on four datasets demonstrate that the proposed model achieves significant improvements in complex reasoning tasks involving diverse structures, and shows transferability to downstream tasks with limited training data and effectiveness for complex reasoning of KGs modality.
Self-Supervised Beat Tracking in Musical Signals with Polyphonic Contrastive Learning
Annotating musical beats is a very long and tedious process. In order to combat this problem, we present a new self-supervised learning pretext task for beat tracking and downbeat estimation. This task makes use of Spleeter, an audio source separation model, to separate a song's drums from the rest of its signal. The first set of signals are used as positives, and by extension negatives, for contrastive learning pre-training. The drum-less signals, on the other hand, are used as anchors. When pre-training a fully-convolutional and recurrent model using this pretext task, an onset function is learned. In some cases, this function is found to be mapped to periodic elements in a song. We find that pre-trained models outperform randomly initialized models when a beat tracking training set is extremely small (less than 10 examples). When this is not the case, pre-training leads to a learning speed-up that causes the model to overfit to the training set. More generally, this work defines new perspectives in the realm of musical self-supervised learning. It is notably one of the first works to use audio source separation as a fundamental component of self-supervision.
The Unlikely Stars of the Actors Strike (So Far)
The idea that studios want actors to relinquish their digital doubles forever in exchange for a few tanks of gas (and that's before taxes!) was just too deliciously infuriating not to retweet or Thread. The scenario transformed the studios into creative labor–devouring supervillains. Even if you hadn't previously cared much about the looming strike, suddenly you were angry--the studios wanted to get away with replacing human actors for free. And who knew how they'd use those digital doubles! If you've watched the recent episode of Black Mirror in which a streaming site rationalizes all kinds of twisted uses of Salma Hayek's digital double based on a legal agreement, it's not hard to imagine the whole thing truly going off the rails.