Goto

Collaborating Authors

 Media


Search engines that don't pay up can't index Reddit content

Engadget

When Reddit said last month that it would block unauthorized data scraping from its site, everyone's (rightful) first reaction was "AI, AI, AI." However, now that the change has taken effect, chatbot makers aren't the only ones being locked out. The widely used forum also appears to be blocking all search engines other than Google, which reportedly inked a deal earlier this year with Reddit worth 60 million annually. The publication reported that DuckDuckGo produced seven links without any descriptions, only providing the note, "We would like to show you a description here but the site won't allow us." The engine now appears to have removed even those, as our test only produced an empty page, reading, "no results found."


Why Colin Kaepernick Is Starting an AI Company

TIME - Tech

When NFL quarterback Colin Kaepernick began kneeling during the national anthem to protest police brutality and racial injustice in 2016, he soon found himself out of a job, eventually moving onto other ventures in media and entertainment. Today, he's entering the AI industry by launching a project he says he hopes will allow others to bypass "gatekeeping:" an artificial intelligence platform called Lumi. The new subscription-based platform aims to provide tools for storytellers to create, illustrate, publish and monetize their ideas. The company has raised 4 million in funding led by Alexis Ohanian's Seven Seven Six, and its product went live today, July 24. In an interview with TIME, Kaepernick says this project can be viewed as an extension of his activism.


Fox News AI Newsletter: Waymo's robotaxi launches citywide in San Francisco

FOX News

UPenn Wharton School Associate Professor Ethan Mollick weighs in on the Biden White House's new guidelines for artificial intelligence in the workplace on'Fox News Live.' DRIVERLESS TAXIS ARRIVE: The future of urban transportation is here, and it's taking the form of sleek, autonomous vehicles traveling through city streets. Across the United States, self-driving car companies are racing to revolutionize how we move, promising safer roads, reduced traffic and a new era of mobility. But it's in San Francisco that this future is suddenly now a reality for thousands. 'SHADOWY ECOSYSTEM': The Federal Trade Commission on Tuesday announced that it launched a probe of eight companies that offer "surveillance pricing" tools that use artificial intelligence and other technology to analyze consumer data to help set price targets for products and services. AI IN THE SKY: The U.S. Air Force has just unveiled a new aircraft that's turning heads and raising eyebrows across the globe.


The Morning After: Netflix's new gaming boss is a former Epic Games exec

Engadget

Netflix has hired Alain Tascan as its new president of games. Before joining Netflix, Tascan was executive vice president for Epic Games and oversaw first-party development for some of the company's (and gaming's) most successful titles, like Fortnite, Rocket League and Fall Guys. Since launching its games project in 2021, Netflix has acquired notable indie studios Night School, Boss Fight, Next Games and Spry Fox and has brought many great indie games to mobile -- seriously, search the app store, if only for Into The Breach. Netflix recently said it has 80-plus games currently in development. A multiplayer Squid Game project will be part of that, coinciding with the hit show's next season, later this year.


California's news industry is shrinking while misinformation spreads. Here's what the numbers tell us

Los Angeles Times

As the world turned digital, people were quick to drop their Sunday papers and pick up their smartphones for news. Advertisers followed suit as digital platforms became more valuable real estate than print newspapers, leaving California news outlets desperate to find ways to stay profitable and relevant. News outlets must spend at least 70% of the received funds on their staff. A second bill being considered by California lawmakers, Senate Bill 1327, would charge Amazon, Meta and Google a "data extraction mitigation fee" for data they collect from users. The funds would go toward supporting local newsrooms.


Steven Pinker: Young people sick and tired of being told, 'you can't say that, you can't think that' on campus

FOX News

Dr. Steven Pinker, a Harvard psychologist and prolific author, has often been described as a cheerleader for science, reason, and humanism. He is often maligned by his critics as a defender of the status quo. Much of his research focuses on slow and steady incremental improvements that have defined rapid human development, both in the United States and globally, over the past century. His 2018 book, "Enlightenment Now" was famously cited by Bill Gates as "his new favorite book," and became a focal point for global policymakers. He is a fierce defender of liberalism, democracy, and market economies, and believes a variety of forces are conspiring against them: populism of both the right and left, religious fundamentalism, and political correctness, among others. He also has emerged as a champion of reasoned, civil debate on college campuses, pushing back against cancel culture, and what he views as a'political monoculture' in academia.


Multimodal Detection of Bots on X (Twitter) using Transformers

arXiv.org Artificial Intelligence

Although not all bots are malicious, the vast majority of them are responsible for spreading misinformation and manipulating the public opinion about several issues, i.e., elections and many more. Therefore, the early detection of bots is crucial. Although there have been proposed methods for detecting bots in social media, there are still substantial limitations. For instance, existing research initiatives still extract a large number of features and train traditional machine learning algorithms or use GloVe embeddings and train LSTMs. However, feature extraction is a tedious procedure demanding domain expertise. Also, language models based on transformers have been proved to be better than LSTMs. Other approaches create large graphs and train graph neural networks requiring in this way many hours for training and access to computational resources. To tackle these limitations, this is the first study employing only the user description field and images of three channels denoting the type and content of tweets posted by the users. Firstly, we create digital DNA sequences, transform them to 3d images, and apply pretrained models of the vision domain, including EfficientNet, AlexNet, VGG16, etc. Next, we propose a multimodal approach, where we use TwHIN-BERT for getting the textual representation of the user description field and employ VGG16 for acquiring the visual representation for the image modality. We propose three different fusion methods, namely concatenation, gated multimodal unit, and crossmodal attention, for fusing the different modalities and compare their performances. Finally, we present a qualitative analysis of the behavior of our best performing model. Extensive experiments conducted on the Cresci'17 and TwiBot-20 datasets demonstrate valuable advantages of our introduced approaches over state-of-the-art ones.


Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model

arXiv.org Artificial Intelligence

In this paper, we experimented with the SpeechT5 model pre-trained on large-scale datasets. We pre-trained the foundation model from scratch and fine-tuned it on a large-scale robust multi-speaker text-to-speech (TTS) task. We tested the model capabilities in a zero- and few-shot scenario. Based on two listening tests, we evaluated the synthetic audio quality and the similarity of how synthetic voices resemble real voices. Our results showed that the SpeechT5 model can generate a synthetic voice for any speaker using only one minute of the target speaker's data. We successfully demonstrated the high quality and similarity of our synthetic voices on publicly known Czech politicians and celebrities.


Speech Editing -- a Summary

arXiv.org Artificial Intelligence

With the rise of video production and social media, speech editing has become crucial for creators to address issues like mispronunciations, missing words, or stuttering in audio recordings. This paper explores text-based speech editing methods that modify audio via text transcripts without manual waveform editing. These approaches ensure edited audio is indistinguishable from the original by altering the mel-spectrogram. Recent advancements, such as context-aware prosody correction and advanced attention mechanisms, have improved speech editing quality. This paper reviews state-of-the-art methods, compares key metrics, and examines widely used datasets. The aim is to highlight ongoing issues and inspire further research and innovation in speech editing.


WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries

arXiv.org Artificial Intelligence

While hallucinations of large language models (LLMs) prevail as a major challenge, existing evaluation benchmarks on factuality do not cover the diverse domains of knowledge that the real-world users of LLMs seek information about. To bridge this gap, we introduce WildHallucinations, a benchmark that evaluates factuality. It does so by prompting LLMs to generate information about entities mined from user-chatbot conversations in the wild. These generations are then automatically fact-checked against a systematically curated knowledge source collected from web search. Notably, half of these real-world entities do not have associated Wikipedia pages. We evaluate 118,785 generations from 15 LLMs on 7,919 entities. We find that LLMs consistently hallucinate more on entities without Wikipedia pages and exhibit varying hallucination rates across different domains. Finally, given the same base models, adding a retrieval component only slightly reduces hallucinations but does not eliminate hallucinations.