Media
BEADs: Bias Evaluation Across Domains
Raza, Shaina, Rahman, Mizanur, Zhang, Michael R.
Recent improvements in large language models (LLMs) have significantly enhanced natural language processing (NLP) applications. However, these models can also inherit and perpetuate biases from their training data. Addressing this issue is crucial, yet many existing datasets do not offer evaluation across diverse NLP tasks. To tackle this, we introduce the Bias Evaluations Across Domains (BEADs) dataset, designed to support a wide range of NLP tasks, including text classification, bias entity recognition, bias quantification, and benign language generation. BEADs uses AI-driven annotation combined with experts' verification to provide reliable labels. This method overcomes the limitations of existing datasets that typically depend on crowd-sourcing, expert-only annotations with limited bias evaluations, or unverified AI labeling. Our empirical analysis shows that BEADs is effective in detecting and reducing biases across different language models, with smaller models fine-tuned on BEADs often outperforming LLMs in bias classification tasks. However, these models may still exhibit biases towards certain demographics. Fine-tuning LLMs with our benign language data also reduces biases while preserving the models' knowledge. Our findings highlight the importance of comprehensive bias evaluation and the potential of targeted fine-tuning for reducing the bias of LLMs. We are making BEADs publicly available at https://huggingface.co/datasets/shainar/BEAD Warning: This paper contains examples that may be considered offensive.
Sora as an AGI World Model? A Complete Survey on Text-to-Video Generation
Cho, Joseph, Puspitasari, Fachrina Dewi, Zheng, Sheng, Zheng, Jingyao, Lee, Lik-Hang, Kim, Tae-Ho, Hong, Choong Seon, Zhang, Chaoning
The evolution of video generation from text, starting with animating MNIST numbers to simulating the physical world with Sora, has progressed at a breakneck speed over the past seven years. While often seen as a superficial expansion of the predecessor text-to-image generation model, text-to-video generation models are developed upon carefully engineered constituents. Here, we systematically discuss these elements consisting of but not limited to core building blocks (vision, language, and temporal) and supporting features from the perspective of their contributions to achieving a world model. We employ the PRISMA framework to curate 97 impactful research articles from renowned scientific databases primarily studying video synthesis using text conditions. Upon minute exploration of these manuscripts, we observe that text-to-video generation involves more intricate technologies beyond the plain extension of text-to-image generation. Our additional review into the shortcomings of Sora-generated videos pinpoints the call for more in-depth studies in various enabling aspects of video generation such as dataset, evaluation metric, efficient architecture, and human-controlled generation. Finally, we conclude that the study of the text-to-video generation may still be in its infancy, requiring contribution from the cross-discipline research community towards its advancement as the first step to realize artificial general intelligence (AGI).
AI language models are running out of human-written text to learn from
UPenn Wharton School Associate Professor Ethan Mollick weighs in on the Biden White House's new guidelines for artificial intelligence in the workplace on'Fox News Live.' Artificial intelligence systems like ChatGPT could soon run out of what keeps making them smarter -- the tens of trillions of words people have written and shared online. A new study released Thursday by research group Epoch AI projects that tech companies will exhaust the supply of publicly available training data for AI language models by roughly the turn of the decade -- sometime between 2026 and 2032. Comparing it to a "literal gold rush" that depletes finite natural resources, Tamay Besiroglu, an author of the study, said the AI field might face challenges in maintaining its current pace of progress once it drains the reserves of human-generated writing. In the short term, tech companies like ChatGPT-maker OpenAI and Google are racing to secure and sometimes pay for high-quality data sources to train their AI large language models โ for instance, by signing deals to tap into the steady flow of sentences coming out of Reddit forums and news media outlets. In the longer term, there won't be enough new blogs, news articles and social media commentary to sustain the current trajectory of AI development, putting pressure on companies to tap into sensitive data now considered private -- such as emails or text messages -- or relying on less-reliable "synthetic data" spit out by the chatbots themselves.
Tokyo City Hall is creating a dating app to encourage marriage amid Japan's historically low birth rate
Called "Tokyo Futari Story," the city hall's new initiative is just that: An effort to create couples, "futari," in a country where it is increasingly common to be "hitori," or alone. While a site offering counsel and general information for potential lovebirds is online, a dating app is also in development. City hall hopes to offer it later this year, accessible through phone or web, a city official said Thursday. City Hall declined to comment on Japanese media reports that said the app will require a confirmation of identity, such as a driver's license, your tax records to prove income and a signed form that says you are ready to get married. 'MEET HOT, SINGLE FIREMEN, SCORE A PRIZE': NEWEST WAY WOMEN ARE FINDING THEIR LOVE MATCHES Marriage is on the decline in Japan as the country's birth rate fell to an all-time low, according to health ministry data on Wednesday.
Popular US news app accused of using AI to make up fake stories
NewsBreak, a popular free news app in the US, has been publishing fictitious stories written by AI since 2021, according to Reuters. The app publishes licensed content from legitimate news sources, such as CNN, AP and Reuters itself, but it also uses artificial intelligence tools to rewrite press releases and local news. One of the most egregious examples of a false news story by NewsBreak was published on Christmas Eve last year. The app's writeup claimed that there was a shooting in Bridgeton, New Jersey when no such incident took place. New Jersey's police department dismissed the claims made in the article before the app, which said it got the information from another website, took it down four days later.
An AI Cartoon May Interview You For Your Next Job
The cartoon interviewer greets you on screen. He looks a little young to be asking questions about a job--sort of a cartoon version of Harry Potter, with dark hair and glasses. You can choose other interviewers to speak with instead, representing various genders and races with names like Benjamin, Leslie, and Kristin. Alex, the name given to this AI interviewer, asks about your professional experience, theoretical questions about programming, and then gives out a coding exercise. Alex is an AI interviewer developed by micro1, a US company that describes itself as an AI recruitment engine for engineers.
How to Lead an Army of Digital Sleuths in the Age of AI
Ten years ago, Eliot Higgins could eat room service meals at a hotel without fear of being poisoned. He hadn't yet been declared a foreign agent by Russia; in fact, he wasn't even a blip on the radar of security agencies in that country or anywhere else. He was just a British guy with an unfulfilling admin job who'd been blogging under the pen name Brown Moses--after a Frank Zappa song--and was in the process of turning his blog into a full-fledged website. He was an open source intelligence analyst avant la lettre, poring over social media photos and videos and other online jetsam to investigate wartime atrocities in Libya and Syria. In its disorganized way, the internet supplied him with so much evidence that he was beating UN investigators to their conclusions.
Unintended Impacts of LLM Alignment on Global Representation
Ryan, Michael J., Held, William, Yang, Diyi
Before being deployed for user-facing applications, developers align Large Language Models (LLMs) to user preferences through a variety of procedures, such as Reinforcement Learning From Human Feedback (RLHF) and Direct Preference Optimization (DPO). Current evaluations of these procedures focus on benchmarks of instruction following, reasoning, and truthfulness. However, human preferences are not universal, and aligning to specific preference sets may have unintended effects. We explore how alignment impacts performance along three axes of global representation: English dialects, multilingualism, and opinions from and about countries worldwide. Our results show that current alignment procedures create disparities between English dialects and global opinions. We find alignment improves capabilities in several languages. We conclude by discussing design decisions that led to these unintended impacts and recommendations for more equitable preference tuning. We make our code and data publicly available on Github.
A + B: A General Generator-Reader Framework for Optimizing LLMs to Unleash Synergy Potential
Tang, Wei, Cao, Yixin, Ying, Jiahao, Wang, Bo, Zhao, Yuyue, Liao, Yong, Zhou, Pengyuan
Retrieval-Augmented Generation (RAG) is an effective solution to supplement necessary knowledge to large language models (LLMs). Targeting its bottleneck of retriever performance, "generate-then-read" pipeline is proposed to replace the retrieval stage with generation from the LLM itself. Although promising, this research direction is underexplored and still cannot work in the scenario when source knowledge is given. In this paper, we formalize a general "A + B" framework with varying combinations of foundation models and types for systematic investigation. We explore the efficacy of the base and chat versions of LLMs and found their different functionalities suitable for generator A and reader B, respectively. Their combinations consistently outperform single models, especially in complex scenarios. Furthermore, we extend the application of the "A + B" framework to scenarios involving source documents through continuous learning, enabling the direct integration of external knowledge into LLMs. This approach not only facilitates effective acquisition of new knowledge but also addresses the challenges of safety and helpfulness post-adaptation. The paper underscores the versatility of the "A + B" framework, demonstrating its potential to enhance the practical application of LLMs across various domains.
Assessing LLMs for Zero-shot Abstractive Summarization Through the Lens of Relevance Paraphrasing
Askari, Hadi, Chhabra, Anshuman, Chen, Muhao, Mohapatra, Prasant
Large Language Models (LLMs) have achieved state-of-the-art performance at zero-shot generation of abstractive summaries for given articles. However, little is known about the robustness of such a process of zero-shot summarization. To bridge this gap, we propose relevance paraphrasing, a simple strategy that can be used to measure the robustness of LLMs as summarizers. The relevance paraphrasing approach identifies the most relevant sentences that contribute to generating an ideal summary, and then paraphrases these inputs to obtain a minimally perturbed dataset. Then, by evaluating model performance for summarization on both the original and perturbed datasets, we can assess the LLM's one aspect of robustness. We conduct extensive experiments with relevance paraphrasing on 4 diverse datasets, as well as 4 LLMs of different sizes (GPT-3.5-Turbo, Llama-2-13B, Mistral-7B, and Dolly-v2-7B). Our results indicate that LLMs are not consistent summarizers for the minimally perturbed articles, necessitating further improvements.