Media
Artificial intelligence has begun to exceed expectations
In 2020 The Guardian published an article that had been written by AI. It was about the increasing use of AI in journalism, and how it is changing the landscape of the industry. It discussed how AI is being used to generate news stories, and how it is being used to help reporters with their work. It was so natural that it was hard to believe that it was written by a software called GPT-3 developed by OpenAI, a research company. The Guardian isn't the only news organization using algorithms to write articles.
Digital Media Firm Stability AI Raises Funds at $1 Billion Value
The parent company of Stable Diffusion, an artificial intelligence tool that makes digital art, has reached unicorn status after raising funds from some top names in venture capital. Stability AI Ltd. raised $101 million in a seed round led by Coatue Management and Lightspeed Venture Partners, according to a statement reviewed by Bloomberg News.
NT CONNECT Launches News Aggregator App
NT CONNECT, an international technology developer, announces the launch of its news aggregator application, NOOZ.AI, which features an AI-powered language analysis engine whose purpose is to bring transparency to the polarizing bias found in today's news media. This service aims to keep readers mindful of the news media's influence before reading an article. Authors or journalists tend to lean toward a particular bias--often without the reader's knowledge. By knowing the author and news source's historical bias, consumers can examine the article with a more objective mind and resist being manipulated to think in a certain way about any particular subject. NOOZ.AI is comprised of four key pillars.
Getting Started with Applied AI and NLP
Originally published on Towards AI the World's Leading AI and Technology News and Media Company. If you are building an AI-related product or service, we invite you to consider becoming an AI sponsor. At Towards AI, we help scale AI and technology startups. Let us help you unleash your technology to the masses. No need for complex self-made Machine Learning Projects -- it's time to use APIs instead.
Happy
Clap along, if you feel like that's what you wanna do Back in 2014, hip-hop producer Pharrell Williams wrote Happy for his friend Cee Lo Green, and had him record the song to include on Pharrell's upcoming album. Unfortunately, Cee Lo Green's record label executives vetoed the song's release, believing it would subtract attention from Green's own upcoming album. Upset but unfazed by this idiotic slight toward him, Pharrell recorded a new version of the song himself, and simply released that instead. "Happy" went on to become one of the singular most popular recorded songs in history, breaking every record one single song can break along the way and sending Pharrell's career to new and rarified heights. We are pleased to announce the Women Leaders of Conversational AI, Class of 2023: approximately 200 women who themselves have shown perseverance in their own careers as they've worked to impact the conversational AI / voice technology continuum.
Auditing YouTube's Recommendation Algorithm for Misinformation Filter Bubbles
Srba, Ivan, Moro, Robert, Tomlein, Matus, Pecher, Branislav, Simko, Jakub, Stefancova, Elena, Kompan, Michal, Hrckova, Andrea, Podrouzek, Juraj, Gavornik, Adrian, Bielikova, Maria
In this paper, we present results of an auditing study performed over YouTube aimed at investigating how fast a user can get into a misinformation filter bubble, but also what it takes to "burst the bubble", i.e., revert the bubble enclosure. We employ a sock puppet audit methodology, in which pre-programmed agents (acting as YouTube users) delve into misinformation filter bubbles by watching misinformation promoting content. Then they try to burst the bubbles and reach more balanced recommendations by watching misinformation debunking content. We record search results, home page results, and recommendations for the watched videos. Overall, we recorded 17,405 unique videos, out of which we manually annotated 2,914 for the presence of misinformation. The labeled data was used to train a machine learning model classifying videos into three classes (promoting, debunking, neutral) with the accuracy of 0.82. We use the trained model to classify the remaining videos that would not be feasible to annotate manually. Using both the manually and automatically annotated data, we observe the misinformation bubble dynamics for a range of audited topics. Our key finding is that even though filter bubbles do not appear in some situations, when they do, it is possible to burst them by watching misinformation debunking content (albeit it manifests differently from topic to topic). We also observe a sudden decrease of misinformation filter bubble effect when misinformation debunking videos are watched after misinformation promoting videos, suggesting a strong contextuality of recommendations. Finally, when comparing our results with a previous similar study, we do not observe significant improvements in the overall quantity of recommended misinformation content.
Topic Taxonomy Expansion via Hierarchy-Aware Topic Phrase Generation
Lee, Dongha, Shen, Jiaming, Lee, Seonghyeon, Yoon, Susik, Yu, Hwanjo, Han, Jiawei
Topic taxonomies display hierarchical topic structures of a text corpus and provide topical knowledge to enhance various NLP applications. To dynamically incorporate new topic information, several recent studies have tried to expand (or complete) a topic taxonomy by inserting emerging topics identified in a set of new documents. However, existing methods focus only on frequent terms in documents and the local topic-subtopic relations in a taxonomy, which leads to limited topic term coverage and fails to model the global topic hierarchy. In this work, we propose a novel framework for topic taxonomy expansion, named TopicExpan, which directly generates topic-related terms belonging to new topics. Specifically, TopicExpan leverages the hierarchical relation structure surrounding a new topic and the textual content of an input document for topic term generation. This approach encourages newly-inserted topics to further cover important but less frequent terms as well as to keep their relation consistency within the taxonomy. Experimental results on two real-world text corpora show that TopicExpan significantly outperforms other baseline methods in terms of the quality of output taxonomies.
A Second Wave of UD Hebrew Treebanking and Cross-Domain Parsing
Zeldes, Amir, Howell, Nick, Ordan, Noam, Moshe, Yifat Ben
Foundational Hebrew NLP tasks such as segmentation, tagging and parsing, have relied to date on various versions of the Hebrew Treebank (HTB, Sima'an et al. 2001). However, the data in HTB, a single-source newswire corpus, is now over 30 years old, and does not cover many aspects of contemporary Hebrew on the web. This paper presents a new, freely available UD treebank of Hebrew stratified from a range of topics selected from Hebrew Wikipedia. In addition to introducing the corpus and evaluating the quality of its annotations, we deploy automatic validation tools based on grew (Guillaume, 2021), and conduct the first cross domain parsing experiments in Hebrew. We obtain new state-of-the-art (SOTA) results on UD NLP tasks, using a combination of the latest language modelling and some incremental improvements to existing transformer based approaches. We also release a new version of the UD HTB matching annotation scheme updates from our new corpus.
UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models
Xie, Tianbao, Wu, Chen Henry, Shi, Peng, Zhong, Ruiqi, Scholak, Torsten, Yasunaga, Michihiro, Wu, Chien-Sheng, Zhong, Ming, Yin, Pengcheng, Wang, Sida I., Zhong, Victor, Wang, Bailin, Li, Chengzu, Boyle, Connor, Ni, Ansong, Yao, Ziyu, Radev, Dragomir, Xiong, Caiming, Kong, Lingpeng, Zhang, Rui, Smith, Noah A., Zettlemoyer, Luke, Yu, Tao
Structured knowledge grounding (SKG) leverages structured knowledge to complete user requests, such as semantic parsing over databases and question answering over knowledge bases. Since the inputs and outputs of SKG tasks are heterogeneous, they have been studied separately by different communities, which limits systematic and compatible research on SKG. In this paper, we overcome this limitation by proposing the UnifiedSKG framework, which unifies 21 SKG tasks into a text-to-text format, aiming to promote systematic SKG research, instead of being exclusive to a single task, domain, or dataset. We use UnifiedSKG to benchmark T5 with different sizes and show that T5, with simple modifications when necessary, achieves state-of-the-art performance on almost all of the 21 tasks. We further demonstrate that multi-task prefix-tuning improves the performance on most tasks, largely improving the overall performance. UnifiedSKG also facilitates the investigation of zero-shot and few-shot learning, and we show that T0, GPT-3, and Codex struggle in zero-shot and few-shot learning for SKG. We also use UnifiedSKG to conduct a series of controlled experiments on structured knowledge encoding variants across SKG tasks. UnifiedSKG is easily extensible to more tasks, and it is open-sourced at https://github.com/hkunlp/unifiedskg.