Media
Analysis: AI vigilantes fuel censorship fears in Russian cyberspace
Government enlists AI tools to flag'destructive' content Increasing number of Russians face court over social media posts Activists fear AI will be used to stifle dissent The company and law firm names shown above are generated automatically based on the text of the article. We are improving this feature as we continue to test and develop in beta. We welcome feedback, which you can provide using the feedback tab on the right of the page. TBILISI, Nov 30 (Thomson Reuters Foundation) - A woman posing in a thong outside a church; a single mother who berated Russian lawmakers and President Vladimir Putin; a saxophonist who criticised World War Two commemorations. They are among thousands of Russians who have faced court over their social media posts in the past year - a number digital rights groups say could soon turn into a deluge as authorities use artificial intelligence (AI) to police the web.
Researchers built an AI that automatically generates movie trailers
Tristan covers human-centric artificial intelligence advances, quantum computing, STEM, Spiderman, physics, and space stuff. Pronouns: He/hi (show all) Tristan covers human-centric artificial intelligence advances, quantum computing, STEM, Spiderman, physics, and space stuff. You can have Citizen Kane andThe Godfather. Keep your 3D films and IMAX experiences. The only real film-based art form is movie trailers.
MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions
Soldan, Mattia, Pardo, Alejandro, Alcázar, Juan León, Heilbron, Fabian Caba, Zhao, Chen, Giancola, Silvio, Ghanem, Bernard
The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at assessing the fitness of these datasets for the video-language grounding task. Recent works have begun to discover significant limitations in these datasets, suggesting that state-of-the-art techniques commonly overfit to hidden dataset biases. In this work, we present MAD (Movie Audio Descriptions), a novel benchmark that departs from the paradigm of augmenting existing video datasets with text annotations and focuses on crawling and aligning available audio descriptions of mainstream movies. MAD contains over 384,000 natural language sentences grounded in over 1,200 hours of video and exhibits a significant reduction in the currently diagnosed biases for video-language grounding datasets. MAD's collection strategy enables a novel and more challenging version of video-language grounding, where short temporal moments (typically seconds long) must be accurately grounded in diverse long-form videos that can last up to three hours.
NER-BERT: A Pre-trained Model for Low-Resource Entity Tagging
Liu, Zihan, Jiang, Feijun, Hu, Yuxiang, Shi, Chen, Fung, Pascale
Named entity recognition (NER) models generally perform poorly when large training datasets are unavailable for low-resource domains. Recently, pre-training a large-scale language model has become a promising direction for coping with the data scarcity issue. However, the underlying discrepancies between the language modeling and NER task could limit the models' performance, and pre-training for the NER task has rarely been studied since the collected NER datasets are generally small or large but with low quality. In this paper, we construct a massive NER corpus with a relatively high quality, and we pre-train a NER-BERT model based on the created dataset. Experimental results show that our pre-trained model can significantly outperform BERT (Devlin et al., 2019) as well as other strong baselines in low-resource scenarios across nine diverse domains. Moreover, a visualization of entity representations further indicates the effectiveness of NER-BERT for categorizing a variety of entities.
The Effect of Iterativity on Adversarial Opinion Forming
Panagiotou, Konstantinos, Reisser, Simon
Understanding how opinions are formed is as important as ever, as the spread of misinformation becomes more prevalent every day. Assume there is some new innovation being either good or bad that is introduced to a group of people who want to form their (binary) opinion about it. Following a key insight by Rogers [22], the opining forming process can be modelled as follows. At first, a small set of so-called early adopters, or experts, forms their opinion about the newly introduced innovation. Afterwards, they disseminate their opinion to all other non-experts in the network. When looking at that network from the outside an observer wants to infer the quality of the new innovation by observing the opinion of all individuals, but without taking the actual structure of the network into consideration (maybe by doing a poll). One popular method to achieve this is using the wisdom of the crowd. In this case that corresponds to a simple majority rule, that is, the observer takes the majority of opinions as an estimate. Wisdom of the crowd has been shown to have a plethora of useful applications in decision making, see e.g.
Hallucinated Neural Radiance Fields in the Wild
Chen, Xingyu, Zhang, Qi, Li, Xiaoyu, Chen, Yue, Feng, Ying, Wang, Xuan, Wang, Jue
Neural Radiance Fields (NeRF) has recently gained popularity for its impressive novel view synthesis ability. This paper studies the problem of hallucinated NeRF: i.e. recovering a realistic NeRF at a different time of day from a group of tourism images. Existing solutions adopt NeRF with a controllable appearance embedding to render novel views under various conditions, but cannot render view-consistent images with an unseen appearance. To solve this problem, we present an end-to-end framework for constructing a hallucinated NeRF, dubbed as Ha-NeRF. Specifically, we propose an appearance hallucination module to handle time-varying appearances and transfer them to novel views. Considering the complex occlusions of tourism images, an anti-occlusion module is introduced to decompose the static subjects for visibility accurately. Experimental results on synthetic data and real tourism photo collections demonstrate that our method can not only hallucinate the desired appearances, but also render occlusion-free images from different views. The project and supplementary materials are available at https://rover-xingyu.github.io/Ha-NeRF/.
Training computers to tease out subtext behind text
WEST LAFAYETTE, Ind. – It is hard enough for humans to interpret the deeper meaning and context of social media and news articles. Asking computers to do it is a nearly impossible task. Even C-3PO, fluent in over 6 million forms of communication, misses the subtext much of the time. Natural language processing, the subfield of artificial intelligence connecting computers with human languages, uses statistical methods to analyze language, often without incorporating the real-world context needed for understanding the shifts and currents of human society. To do that, you have to translate online communication, and the context from which it emerges, into something the computers can parse and reason over.