Goto

Collaborating Authors

 Media


Neural models for Factual Inconsistency Classification with Explanations

arXiv.org Artificial Intelligence

Factual consistency is one of the most important requirements when editing high quality documents. It is extremely important for automatic text generation systems like summarization, question answering, dialog modeling, and language modeling. Still, automated factual inconsistency detection is rather under-studied. Existing work has focused on (a) finding fake news keeping a knowledge base in context, or (b) detecting broad contradiction (as part of natural language inference literature). However, there has been no work on detecting and explaining types of factual inconsistencies in text, without any knowledge base in context. In this paper, we leverage existing work in linguistics to formally define five types of factual inconsistencies. Based on this categorization, we contribute a novel dataset, FICLE (Factual Inconsistency CLassification with Explanation), with ~8K samples where each sample consists of two sentences (claim and context) annotated with type and span of inconsistency. When the inconsistency relates to an entity type, it is labeled as well at two levels (coarse and fine-grained). Further, we leverage this dataset to train a pipeline of four neural models to predict inconsistency type with explanations, given a (claim, context) sentence pair. Explanations include inconsistent claim fact triple, inconsistent context span, inconsistent claim component, coarse and fine-grained inconsistent entity types. The proposed system first predicts inconsistent spans from claim and context; and then uses them to predict inconsistency types and inconsistent entity types (when inconsistency is due to entities). We experiment with multiple Transformer-based natural language classification as well as generative models, and find that DeBERTa performs the best. Our proposed methods provide a weighted F1 of ~87% for inconsistency type classification across the five classes.


WebIE: Faithful and Robust Information Extraction on the Web

arXiv.org Artificial Intelligence

Extracting structured and grounded fact triples from raw text is a fundamental task in Information Extraction (IE). Existing IE datasets are typically collected from Wikipedia articles, using hyperlinks to link entities to the Wikidata knowledge base. However, models trained only on Wikipedia have limitations when applied to web domains, which often contain noisy text or text that does not have any factual information. We present WebIE, the first large-scale, entity-linked closed IE dataset consisting of 1.6M sentences automatically collected from the English Common Crawl corpus. WebIE also includes negative examples, i.e. sentences without fact triples, to better reflect the data on the web. We annotate ~21K triples from WebIE through crowdsourcing and introduce mWebIE, a translation of the annotated set in four other languages: French, Spanish, Portuguese, and Hindi. We evaluate the in-domain, out-of-domain, and zero-shot cross-lingual performance of generative IE models and find models trained on WebIE show better generalisability. We also propose three training strategies that use entity linking as an auxiliary task. Our experiments show that adding Entity-Linking objectives improves the faithfulness of our generative IE models.


Tool Learning with Foundation Models

arXiv.org Artificial Intelligence

Humans possess an extraordinary ability to create and utilize tools, allowing them to overcome physical limitations and explore new frontiers. With the advent of foundation models, AI systems have the potential to be equally adept in tool use as humans. This paradigm, i.e., tool learning with foundation models, combines the strengths of specialized tools and foundation models to achieve enhanced accuracy, efficiency, and automation in problem-solving. Despite its immense potential, there is still a lack of a comprehensive understanding of key challenges, opportunities, and future endeavors in this field. To this end, we present a systematic investigation of tool learning in this paper. We first introduce the background of tool learning, including its cognitive origins, the paradigm shift of foundation models, and the complementary roles of tools and models. Then we recapitulate existing tool learning research into tool-augmented and tool-oriented learning. We formulate a general tool learning framework: starting from understanding the user instruction, models should learn to decompose a complex task into several subtasks, dynamically adjust their plan through reasoning, and effectively conquer each sub-task by selecting appropriate tools. We also discuss how to train models for improved tool-use capabilities and facilitate the generalization in tool learning. Considering the lack of a systematic tool learning evaluation in prior works, we experiment with 18 representative tools and show the potential of current foundation models in skillfully utilizing tools. Finally, we discuss several open problems that require further investigation for tool learning. Overall, we hope this paper could inspire future research in integrating tools with foundation models.


Harvard professor believes aliens will make first contact with artificial intelligence - not humans

Daily Mail - Science & tech

A Harvard professor believes aliens will not make first contact with humans but instead will communicate with artificial intelligence. Avi Loeb shared the theory in a new documentary, God Versus Aliens, slated for July, in which he suggests extraterrestrials will send AI drones to Earth rather than'crewed' vehicles. Directed by British musician and TV director Mark Christopher Lee described Loeb as a'very active mind' but explained Loeb's suggestion is based on the vast distance aliens could have to travel to reach us. 'Loeb proposes that it's likely to be some form of AI because why would you send flesh and blood creatures?' Lee said. 'That means there's a possibility that their AI could just connect with AI and bypass humans, which is a bit scary to think about.


Reddit moderators vow to continue blackout in API access fees row

The Guardian

Reddit's battle with its own users over new access fees will continue beyond the planned two-day protest, as hundreds of volunteer moderators declared their intention to maintain a blackout indefinitely. The social network, which intends to begin levying swingeing data charges against developers of third-party tools used to browse the site, says it has no intention of backing down from its plans in the wake of the campaign. The new fees, payable by any service that uses the site's tools, or API, to access information, are in part intended to allow the company to monetise its popularity among artificial intelligence researchers, who use the database to train in tools such as GPT-4. "We're not planning any changes to the API updates we've previously announced," a Reddit spokesperson told the Guardian. "We're in contact with a number of communities to clarify any confusion around our data API terms, platform-wide policies, community support resources, and timing for new moderator tools."


The Morning After: OpenAI and Microsoft aren't happy

Engadget

Microsoft may own almost half of OpenAI, but a recent expose hints the pair aren't the happiest of bedfellows. The Wall Street Journal claims the AI company warned Microsoft not to incorporate GPT-4 into Bing search without further training, but it did so anyway. It resulted in several high-profile examples of odd behavior, including bots arguing with users, and at least one instance of a user being urged to dissolve their marriage and elope with Bing instead. There's resentment, too, on Microsoft's side, finding its own internal AI projects overlooked in favor of OpenAI. Which, despite the close financial ties, is very much free to work with Microsoft's rivals in plenty of fields.


AI helped create new Beatles song using Lennon's voice: McCartney

Al Jazeera

Artificial intelligence has helped create a final Beatles song set to be released this year, its member Paul McCartney has said. In an interview released on Tuesday by the BBC, McCartney said the technology was used to "extricate" John Lennon's voice from an old demo which was used to complete the song. "We just finished it up, and it'll be released this year," he said. McCartney, 84, said the song was made with the help of film director Peter Jackson, using the same AI technology employed for the Beatles documentary Get Back. During the making of that film, Jackson and his team were able to separate the voices from the instruments.


A step toward safe and reliable autopilots for flying

Robohub

MIT researchers developed a machine-learning technique that can autonomously drive a car or fly a plane through a very difficult "stabilize-avoid" scenario, in which the vehicle must stabilize its trajectory to arrive at and stay within some goal region, while avoiding obstacles. In the film "Top Gun: Maverick," Maverick, played by Tom Cruise, is charged with training young pilots to complete a seemingly impossible mission -- to fly their jets deep into a rocky canyon, staying so low to the ground they cannot be detected by radar, then rapidly climb out of the canyon at an extreme angle, avoiding the rock walls. Spoiler alert: With Maverick's help, these human pilots accomplish their mission. A machine, on the other hand, would struggle to complete the same pulse-pounding task. To an autonomous aircraft, for instance, the most straightforward path toward the target is in conflict with what the machine needs to do to avoid colliding with the canyon walls or staying undetected.


The Beatles Are Releasing One More Song--With Help From AI

TIME - Tech

Artificial intelligence has been used to extract John Lennon's voice from an old demo to create "the last Beatles record," decades after the band broke up, Paul McCartney said Tuesday. McCartney, 80, told the BBC that the technology was used to separate the Beatles' voices from background sounds during the making of director Peter Jackson's 2021 documentary series, "The Beatles: Get Back." The "new" song is set to be released later this year, he said. Jackson was "able to extricate John's voice from a ropey little bit of cassette and a piano," McCartney told BBC radio. "He could separate them with AI, he'd tell the machine'That's a voice, this is a guitar, lose the guitar'."


John Rich doesn't think AI could be any worse than the state of country music today

FOX News

John Rich shared his unfiltered opinions on the state of country music today and if AI would make it better or worse. John Rich is less concerned with the advancements in artificial intelligence than he is with the expansion of country music, questioning if the technology could produce better quality music than what country artists are releasing now. "Could AI do any worse than some of the country singers that are out there right now?" Rich wondered during an interview with Fox News Digital, without naming names. One thing's for sure, according to Rich: AI has nothing on the legends. "Listen, you can't replicate the great songwriters. I mean, you're talking about Albert Einstein honky-tonk songwriters."