Goto

Collaborating Authors

 Large Language Model


British AI startup beats humans in international forecasting competition

The Guardian

The Metaculus Cup required entrants to forecast the likelihood of 60 events over the summer. The Metaculus Cup required entrants to forecast the likelihood of 60 events over the summer. ManticAI ranked eighth in the Metaculus Cup, leaving some believing bots' prediction skills could soon overtake experts An artificial intelligence system has beaten scores of forecasting enthusiasts, including several professionals, in a contest to predict events ranging from bust-ups between Donald Trump and Elon Musk to Kemi Badenoch being removed from the Conservative party leadership. A British AI startup, co-founded by a former Google DeepMind researcher, has ranked in the top 10 of an international forecasting competition, which requires entrants to forecast the likelihood of 60 events over the summer. ManticAI came eighth in the Metaculus Cup, run by a San Francisco-based forecasting company that tries to predict the future for investment funds and corporations.


YouTube Thinks AI Is Its Next Big Bang

WIRED

On its 20th anniversary, YouTube is venturing into an era of AI-generated video, and may never be the same. Google figured out early on that video would be a great addition to its search business, so in 2005 it launched Google Video. Focused on making deals with the entertainment industry for second-rate content, and overly cautious on what users could upload, it flopped . In 2006, Google snapped up that year-old company, figuring it would sort out the IP stuff later. Though the $1.65 billion purchase price for YouTube was about a billion dollars more than its valuation, it was one of the greatest bargains ever.


Unlocking the Potential of Arabic Voice-Generation Technologies

Communications of the ACM

Membership in ACM includes a subscription to Communications of the ACM (CACM), the computing industry's most trusted source for staying connected to the world of advanced computing. Addressing linguistic complexities, the scarcity of high-quality datasets, and other challenges is crucial for advancing Arabic text-to-speech technology. Voice-generation technology enables machines to synthesize human-like speech--text-to-speech (TTS)--revolutionizing digital communication by fostering more inclusive and accessible experiences. What began as simple robotic speech synthesis has evolved into highly sophisticated voice-cloning systems that can produce natural, coherent, expressive, and personalized voices using minimal data. These technologies empower individuals with cross-lingual communication through virtual agents, assist in overcoming visual or speech impairments or literacy challenges via assistive tools, and support educators and industries such as entertainment with creative content generation.


The Landscape of Arabic Large Language Models

Communications of the ACM

Membership in ACM includes a subscription to Communications of the ACM (CACM), the computing industry's most trusted source for staying connected to the world of advanced computing. The emergence of ChatGPT marked a transformative milestone for artificial intelligence (AI), showcasing the remarkable potential of large language models (LLMs) to generate human-like text. This wave of innovation has revolutionized how we interact with technology, seamlessly integrating LLMs into everyday tasks such as vacation planning, email drafting, and content creation. While English-speaking users have significantly benefited from these advancements, the Arabic world faces distinct challenges in developing Arabic-specific LLMs. Arabic, one of the languages spoken most widely around the world, serves more than 422 million native speakers in 27 countries and is deeply rooted in a rich linguistic and cultural heritage. Developing Arabic LLMs (ALLMs) presents an unparalleled opportunity to bridge technological gaps and empower communities. The journey of ALLMs has been both fascinating and complex, evolving from rudimentary text-processing systems to sophisticated AI-driven models. This article explores the trajectory of ALLMs, from their inception to the present day, highlighting the efforts to evaluate these models through benchmarks and public leaderboards.


AI-Driven Disaster Response and Displacement Monitoring

Communications of the ACM

The 2023 Türkiye-Syria earthquakes, also known as the 2023 Kahramanmaraş earthquakes, were two catastrophic events that struck nine hours apart on February 6, 2023, with epicenters in the Pazarcık and Elbistan districts of Kahramanmaraş, and magnitudes of 7.8 Mw and 7.5 Mw, respectively (see Figure 1).


The Download: the CDC's vaccine chaos

MIT Technology Review

This week has been an eventful one for America's public health agency. Two former leaders of the US Centers for Disease Control and Prevention explained why they suddenly departed in a Senate hearing. They also described how CDC employees are being instructed to turn their backs on scientific evidence. They painted a picture of a health agency in turmoil--and at risk of harming the people it is meant to serve. And, just hours afterwards, a panel of CDC advisers voted to stop recommending the MMRV vaccine for children under four. This article first appeared in The Checkup, MIT Technology Review's weekly biotech newsletter.


Meta Accused of Torrenting Porn to Advance Its Goal of AI 'Superintelligence'

WIRED

The complaint, filed in July, alleges Meta has been torrenting and seeding Strike 3's videos since 2018. Associated exhibits and details of the complaint were unsealed last week. Strike 3 alleges Meta's motive was partly to obtain otherwise difficult to scrape visual angles, parts of the human body, and extended, uninterrupted scenes--rare in mainstream movies and TV--to help it create what Mark Zuckerberg calls AI "superintelligence." "They have an interest in getting our content because it can give them a competitive advantage for the quality, fluidity, and humanity of the AI," alleges Christian Waugh, an attorney for Strike 3. This process made Strike 3's porn videos accessible to minors, the complaint alleges, since BitTorrent does not have age verification.


ChatGPT was used 'to help scammers do their thing' at Asia fraud compound

The Japan Times

ChatGPT was used'to help scammers do their thing' at Asia fraud compound ChatGPT owner OpenAI says it actively works to identify and disrupt scam-related misuse of ChatGPT." Duncan Okindo says he was lured to Southeast Asia last year by the promise of a customer service job in Thailand. Instead, he ended up spending four months in a scam compound on the lawless Myanmar-Thai border, where he saw first-hand how criminal groups are at scale. Okindo, 26, says he was struggling to find a job as the breadwinner for his family in his native Kenya when a local recruitment agency promised him work in Bangkok. The flight was his first trip overseas. On landing, he says, he was abducted at the airport and spirited across the border, into the notorious KK Park complex, guarded by heavily armed men and fortified like it was meant for war."


Simulating a Bias Mitigation Scenario in Large Language Models

arXiv.org Artificial Intelligence

Large Language Models (LLMs) have fundamentally transformed the field of natural language processing; however, their vulnerability to biases presents a notable obstacle that threatens both fairness and trust. This review offers an extensive analysis of the bias landscape in LLMs, tracing its roots and expressions across various NLP tasks. Biases are classified into implicit and explicit types, with particular attention given to their emergence from data sources, architectural designs, and contextual deployments. This study advances beyond theoretical analysis by implementing a simulation framework designed to evaluate bias mitigation strategies in practice. The framework integrates multiple approaches including data curation, debiasing during model training, and post-hoc output calibration and assesses their impact in controlled experimental settings. In summary, this work not only synthesizes existing knowledge on bias in LLMs but also contributes original empirical validation through simulation of mitigation strategies.


When Content is Goliath and Algorithm is David: The Style and Semantic Effects of Generative Search Engine

arXiv.org Artificial Intelligence

Generative search engines (GEs) leverage large language models (LLMs) to deliver AI-generated summaries with website citations, establishing novel traffic acquisition channels while fundamentally altering the search engine optimization landscape. To investigate the distinctive characteristics of GEs, we collect data through interactions with Google's generative and conventional search platforms, compiling a dataset of approximately ten thousand websites across both channels. Our empirical analysis reveals that GEs exhibit preferences for citing content characterized by significantly higher predictability for underlying LLMs and greater semantic similarity among selected sources. Through controlled experiments utilizing retrieval augmented generation (RAG) APIs, we demonstrate that these citation preferences emerge from intrinsic LLM tendencies to favor content aligned with their generative expression patterns. Motivated by applications of LLMs to optimize website content, we conduct additional experimentation to explore how LLM-based content polishing by website proprietors alters AI summaries, finding that such polishing paradoxically enhances information diversity within AI summaries. Finally, to assess the user-end impact of LLM-induced information increases, we design a generative search engine and recruit Prolific participants to conduct a randomized controlled experiment involving an information-seeking and writing task. We find that higher-educated users exhibit minimal changes in their final outputs' information diversity but demonstrate significantly reduced task completion time when original sites undergo polishing. Conversely, lower-educated users primarily benefit through enhanced information density in their task outputs while maintaining similar completion times across experimental groups.