Media
What Should Happen To Our Data When We Die?
The new Anthony Bourdain documentary, "Roadrunner," is one of many projects dedicated to the larger-than-life chef, writer and television personality. But the film has drawn outsize attention, in part because of its subtle reliance on artificial intelligence technology. Using several hours of Bourdain's voice recordings, a software company created 45 seconds of new audio for the documentary. The AI voice sounds just like Bourdain speaking from the great beyond; at one point in the movie, it reads an email he sent before his death by suicide in 2018. "If you watch the film, other than that line you mentioned, you probably don't know what the other lines are that were spoken by the AI, and you're not going to know," Morgan Neville, the director, said in an interview with The New Yorker.
Is Artificial Intelligence the Biggest Threat to Humanity? -- Sentient Machines
The creators of these films imagine a world where humans have lost control of the technology they developed and must fight for survival of the human species. It could also be suggested the writers and directors of these films are predicting a world where these things happen. After all, many notable figures in the world of tech such as Bill Gates, Elon Musk, and even Stephen Hawking have all made warnings against the potential consequences of AI. But how accurate is the silver screen's depiction and are the fears of my friend based off of these films warranted? Is AI going rogue, building a robot army and attempting to eradicate humans as likely as finding a dead body on top of a lift or being attacked by a giant shark off the Isle of Wight?
An exploration into Natural Language Processing with the Universal Studios dataset
I have not posted a lot about Natural Language Processing, NLP, because there are not a lot of data science competitions concerning this genre of machine learning. I have, however, discovered that Kaggle, the premier data science website, does have text based datasets in their dataset section of their website.
Towards Propagation Uncertainty: Edge-enhanced Bayesian Graph Convolutional Networks for Rumor Detection
Wei, Lingwei, Hu, Dou, Zhou, Wei, Yue, Zhaojuan, Hu, Songlin
Detecting rumors on social media is a very critical task with significant implications to the economy, public health, etc. Previous works generally capture effective features from texts and the propagation structure. However, the uncertainty caused by unreliable relations in the propagation structure is common and inevitable due to wily rumor producers and the limited collection of spread data. Most approaches neglect it and may seriously limit the learning of features. Towards this issue, this paper makes the first attempt to explore propagation uncertainty for rumor detection. Specifically, we propose a novel Edge-enhanced Bayesian Graph Convolutional Network (EBGCN) to capture robust structural features. The model adaptively rethinks the reliability of latent relations by adopting a Bayesian approach. Besides, we design a new edge-wise consistency training framework to optimize the model by enforcing consistency on relations. Experiments on three public benchmark datasets demonstrate that the proposed model achieves better performance than baseline methods on both rumor detection and early rumor detection tasks.
A Review of Bangla Natural Language Processing Tasks and the Utility of Transformer Models
Alam, Firoj, Hasan, Arid, Alam, Tanvirul, Khan, Akib, Tajrin, Janntatul, Khan, Naira, Chowdhury, Shammur Absar
Bangla -- ranked as the 6th most widely spoken language across the world (https://www.ethnologue.com/guides/ethnologue200), with 230 million native speakers -- is still considered as a low-resource language in the natural language processing (NLP) community. With three decades of research, Bangla NLP (BNLP) is still lagging behind mainly due to the scarcity of resources and the challenges that come with it. There is sparse work in different areas of BNLP; however, a thorough survey reporting previous work and recent advances is yet to be done. In this study, we first provide a review of Bangla NLP tasks, resources, and tools available to the research community; we benchmark datasets collected from various platforms for nine NLP tasks using current state-of-the-art algorithms (i.e., transformer-based models). We provide comparative results for the studied NLP tasks by comparing monolingual vs. multilingual models of varying sizes. We report our results using both individual and consolidated datasets and provide data splits for future research. We reviewed a total of 108 papers and conducted 175 sets of experiments. Our results show promising performance using transformer-based models while highlighting the trade-off with computational costs. We hope that such a comprehensive survey will motivate the community to build on and further advance the research on Bangla NLP.