Media
Whispers of A.I.'s Modular Future
One day in late December, I downloaded a program called Whisper.cpp onto my laptop, hoping to use it to transcribe an interview I'd done. I fed it an audio file and, every few seconds, it produced one or two lines of eerily accurate transcript, writing down exactly what had been said with a precision I'd never seen before. As the lines piled up, I could feel my computer getting hotter. This was one of the few times in recent memory that my laptop had actually computed something complicated--mostly I just use it to browse the Web, watch TV, and write. Now it was running cutting-edge A.I. Despite being one of the more sophisticated programs ever to run on my laptop, Whisper.cpp is also one of the simplest.
ValueBase, backed by Sam Altman's Hydrazine, raises $1.6 million seed round • TechCrunch
OpenAI CEO Sam Altman believes AI can help usher in "unbelievable abundance," but he says he wants to ensure that such abundance is shared. Toward that end, Altman has embraced a theory of 19th century political economist Henry George, who in his own lifetime worried about wealth amassing in the hands of the few following the Industrial Revolution. George posited that greater equality could be enjoyed if the economic value of land belonged equally to all members of society. Altman similarly believes that in a world where jobs may create less economic value, a land tax could make up for income tax and guarantee that all individuals' assets rise as land -- a fixed asset -- grows in value. He's putting his money where his mouth is, too, leading a seed round in a six-month-old startup that represents a step in that same direction.
How the Supreme Court ruling on Section 230 could end Reddit as we know it
But another big issue is at stake that has received much less attention: depending on the outcome of the case, individual users of sites may suddenly be liable for run-of-the-mill content moderation. Many sites rely on users for community moderation to edit, shape, remove, and promote other users' content online--think Reddit's upvote, or changes to a Wikipedia page. What might happen if those users were forced to take on legal risk every time they made a content decision? In short, the court could change Section 230 in ways that won't just impact big platforms; smaller sites like Reddit and Wikipedia that rely on community moderation will be hit too, warns Emma Llansó, director of the Center for Democracy and Technology's Free Expression Project. "It would be an enormous loss to online speech communities if suddenly it got really risky for mods themselves to do their work," she says.
OpenAI launches AI classifier tool to detect AI generated text
OpenAI, the company behind ChatGPT has launched a new AI classifier tool to determine if a text has been written by a person or by Artificial Intelligence. The tool comes with the caveat that it is not 100 percent reliable and can label human-written text as AI written, you can see more information below. We've trained a classifier to distinguish between text written by a human and text written by AIs from a variety of providers. While it is impossible to reliably detect all AI-written text, we believe good classifiers can inform mitigations for false claims that AI-generated text was written by a human: for example, running automated misinformation campaigns, using AI tools for academic dishonesty, and positioning an AI chatbot as a human. Our classifier is not fully reliable.
New AI classifier for indicating AI-written text
We're launching a classifier trained to distinguish between AI-written and human-written text. We've trained a classifier to distinguish between text written by a human and text written by AIs from a variety of providers. While it is impossible to reliably detect all AI-written text, we believe good classifiers can inform mitigations for false claims that AI-generated text was written by a human: for example, running automated misinformation campaigns, using AI tools for academic dishonesty, and positioning an AI chatbot as a human. Our classifier is not fully reliable. In our evaluations on a "challenge set" of English texts, our classifier correctly identifies 26% of AI-written text (true positives) as "likely AI-written," while incorrectly labeling human-written text as AI-written 9% of the time (false positives).
'Forrest Gump' stars Tom Hanks, Robin Wright to be 'de-aged' in new movie
Tom Hanks was seen speaking at the Australian premiere of "Elvis" earlier this month. Tom Hanks and Robin Wright will be reuniting and going back in time in an upcoming film. "Forrest Gump" director Robert Zemeckis' "Here" will star Wright and Hanks, digitally "de-aged," thanks to Metaphysic, an AI company that will bring the experience to life. The film adaption of Richard McGuire's novel will also star Paul Bettany and Kelly Reilly. "Here" is set to be released 30 years later in 2024.
Multimodality Representation Learning: A Survey on Evolution, Pretraining and Its Applications
Manzoor, Muhammad Arslan, Albarri, Sarah, Xian, Ziting, Meng, Zaiqiao, Nakov, Preslav, Liang, Shangsong
Multimodality Representation Learning, as a technique of learning to embed information from different modalities and their correlations, has achieved remarkable success on a variety of applications, such as Visual Question Answering (VQA), Natural Language for Visual Reasoning (NLVR), and Vision Language Retrieval (VLR). Among these applications, cross-modal interaction and complementary information from different modalities are crucial for advanced models to perform any multimodal task, e.g., understand, recognize, retrieve, or generate optimally. Researchers have proposed diverse methods to address these tasks. The different variants of transformer-based architectures performed extraordinarily on multiple modalities. This survey presents the comprehensive literature on the evolution and enhancement of deep learning multimodal architectures to deal with textual, visual and audio features for diverse cross-modal and modern multimodal tasks. This study summarizes the (i) recent task-specific deep learning methodologies, (ii) the pretraining types and multimodal pretraining objectives, (iii) from state-of-the-art pretrained multimodal approaches to unifying architectures, and (iv) multimodal task categories and possible future improvements that can be devised for better multimodal learning. Moreover, we prepare a dataset section for new researchers that covers most of the benchmarks for pretraining and finetuning. Finally, major challenges, gaps, and potential research topics are explored. A constantly-updated paperlist related to our survey is maintained at https://github.com/marslanm/multimodality-representation-learning.
Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization
Haas, Lukas, Alberti, Silas, Skreta, Michal
By understanding the hidden locational clues in images, entirely new approaches of analyzing the natural and built environment are being opened up with profound implications for a number of fields, ranging from the recognition of weather, season, and climate patterns to rural and urban scene understanding, and improvements in navigation and self-driving car technology. Since the beginning of 2022, image geolocalization has additionally garnered extensive media coverage for becoming an immediate priority of investigative journalists and open source intelligence (OSINT) researchers in their attempt to verify information and to document war atrocities in Ukraine, extracting geolocational information from social media content. Despite high academic and public interest, image geolocalization remains an extremely challenging problem. This is because training datasets are geographically sparse, often limited to specific countries, and biased towards urban or rural scenes. The task is further complicated by the fact that geolocalization requires reasoning on multiple levels of geographic granularity (e.g.
DANES: Deep Neural Network Ensemble Architecture for Social and Textual Context-aware Fake News Detection
Truică, Ciprian-Octavian, Apostol, Elena-Simona, Karras, Panagiotis
The growing popularity of social media platforms has simplified the creation and distribution of news articles but also creates a conduit for spreading fake news. In consequence, the need arises for effective context-aware fake news detection mechanisms, where the contextual information can be built either from the textual content of posts or from available social data (e.g., information about the users, reactions to posts, or the social network). In this paper, we propose DANES, a Deep Neural Network Ensemble Architecture for Social and Textual Context-aware Fake News Detection. DANES comprises a Text Branch for a textual content-based context and a Social Branch for the social context. These two branches are used to create a novel Network Embedding. Preliminary ablation results on 3 real-world datasets, i.e., BuzzFace, Twitter15, and Twitter16, are promising, with an accuracy that outperforms state-of-the-art solutions when employing both social and textual content features.