Goto

Collaborating Authors

 Government


Russian bots use fake Tom Cruise for Olympic disinformation

The Japan Times

A pro-Russian propaganda effort is using artificial intelligence as part of a vast operation to suggest that violence is likely to occur at the upcoming Olympic Games in Paris, according to Microsoft findings released Sunday. Researchers found that one disinformation group used AI-generated audio to make it appear as if actor Tom Cruise had narrated a video titled Olympics Has Fallen, modeled after the 2013 action movie Olympus Has Fallen. The video, which spread in the fall of 2023, presented itself as a Netflix documentary, including the use of Netflix's signature introduction that the company uses on all of its streaming videos. The video also included falsified endorsements from well-known media outlets including The New York Times and the BBC.


The Inconvenient Truth About Elon Musk's New Love Affair With Trump

Slate

On Wednesday, the Wall Street Journal reported on Elon Musk's increasingly close relationship with Republican presidential candidate Donald Trump, which has flourished to the point that the two "talk on the phone several times a month." The conversation subjects tend to cover Trump's attempt to regain White House control and the potential opportunities for Musk and his companies, like Tesla and SpaceX, under another potential Trump administration. Musk rejected the report's central conceit--that Trump had discussed an advisory role for him should the former president be reelected--but he certainly keeps behaving like a typical Trump supplicant. Just look at his X posts following Trump's 34-count conviction in the New York hush money trial, in which he refers to the process as "troubling," endorses a Sequoia Capital partner's 300,000 donation to Trump's campaign, proclaims that "great damage was done today to the public's faith in the American legal system," and reply-guys a couple of characteristically lame Babylon Bee headlines. President Joe Biden's reelection campaign has responded with a scoff, declaring that, "Despite what Donald Trump thinks, America is not for sale to billionaires, oil and gas executives, or even Elon Musk."


The Morning After: Starliner's crewed flight gets scrubbed

Engadget

The first crewed launch of Boeing's Starliner was scrubbed less than four minutes before liftoff after a computer failed to launch the correct countdown. It's the squillionth setback for the craft, (our math may be out a little) which should support the next generation of spaceflight. NASA says it'll target June 5 for its next launch attempt. At this point, we'll believe it when we see it. This tool unlocks Windows' AI-powered Recall feature for unsupported PCs Marvel's "What If...?" for Apple Vision Pro looks incredible, but plays terribly You can get these reports delivered daily direct to your inbox.


Are We Doomed? Here's How to Think About It

The New Yorker

A course at the University of Chicago taught by Daniel Holz and James Evans considers threats posed by climate change, artificial intelligence, nuclear annihilation, and biological warfare, Rivka Galchen writes.


On this day in history, June 3, 1965, Ed White becomes first American to walk in space: 'Just tremendous'

FOX News

The meeting is expected to help the agency's independent study team determine how to evaluate these mysterious sightings going forward. Astronaut Ed White became the first American to walk in space on this day in history, June 3, 1965. White, an engineer, a Lieutenant Colonel in the U.S. Air Force, a test pilot and NASA astronaut, made the spacewalk -- technically known as "Extravehicular Activity" or "EVA" -- while serving as the pilot on the Gemini 4 mission. Command pilot James McDivitt was the other member of the crew, and took pictures of White outside the vehicle. ON THIS DAY IN HISTORY, JUNE 2, 1953, QUEEN ELIZABETH II IS CROWNED IN LONDON'S WESTMINSTER ABBEY White spent about 20 minutes floating outside the Gemini 4 capsule, nearly double the time initially allowed by NASA for the spacewalk.


Using Artificial Intelligence to Accelerate Collective Intelligence: Policy Synth and Smarter Crowdsourcing

arXiv.org Artificial Intelligence

In an era characterized by rapid societal changes and complex challenges, institutions' traditional methods of problem-solving in the public sector are increasingly proving inadequate. In this study, we present an innovative and effective model for how institutions can use artificial intelligence to enable groups of people to generate effective solutions to urgent problems more efficiently. We describe a proven collective intelligence method, called Smarter Crowdsourcing, which is designed to channel the collective intelligence of those with expertise about a problem into actionable solutions through crowdsourcing. Then we introduce Policy Synth, an innovative toolkit which leverages AI to make the Smarter Crowdsourcing problem-solving approach both more scalable, more effective and more efficient. Policy Synth is crafted using a human-centric approach, recognizing that AI is a tool to enhance human intelligence and creativity, not replace it. Based on a real-world case study comparing the results of expert crowdsourcing alone with expert sourcing supported by Policy Synth AI agents, we conclude that Smarter Crowdsourcing with Policy Synth presents an effective model for integrating the collective wisdom of human experts and the computational power of AI to enhance and scale up public problem-solving processes. While many existing approaches view AI as a tool to make crowdsourcing and deliberative processes better and more efficient, Policy Synth goes a step further, recognizing that AI can also be used to synthesize the findings from engagements together with research to develop evidence-based solutions and policies. The study offers practical tools and insights for institutions looking to engage communities effectively in addressing urgent societal challenges.


Evolutionary Computation for the Design and Enrichment of General-Purpose Artificial Intelligence Systems: Survey and Prospects

arXiv.org Artificial Intelligence

In Artificial Intelligence, there is an increasing demand for adaptive models capable of dealing with a diverse spectrum of learning tasks, surpassing the limitations of systems devised to cope with a single task. The recent emergence of General-Purpose Artificial Intelligence Systems (GPAIS) poses model configuration and adaptability challenges at far greater complexity scales than the optimal design of traditional Machine Learning models. Evolutionary Computation (EC) has been a useful tool for both the design and optimization of Machine Learning models, endowing them with the capability to configure and/or adapt themselves to the task under consideration. Therefore, their application to GPAIS is a natural choice. This paper aims to analyze the role of EC in the field of GPAIS, exploring the use of EC for their design or enrichment. We also match GPAIS properties to Machine Learning areas in which EC has had a notable contribution, highlighting recent milestones of EC for GPAIS. Furthermore, we discuss the challenges of harnessing the benefits of EC for GPAIS, presenting different strategies to both design and improve GPAIS with EC, covering tangential areas, identifying research niches, and outlining potential research directions for EC and GPAIS.


TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability

arXiv.org Artificial Intelligence

The typical benchmark evaluations have begun to However, it remains unclear if the model's fall short and do not cover the nuances of LLMs' responses bear useful meaning - whether the model abilities (Zoph et al., 2022). Did the model provide understands the topic or is responding probabilistically a certain answer simply because of the huge purely based on training data. TruthfulQA amount of similar text it saw during training? Or (Lin et al., 2021) comes close to assessing a did the model register a piece of knowledge and use model's understanding of the world but it is designed that to answer the question? It is impossible to tell to exploit the imitative weaknesses of models them apart without analyzing the training dataset, and relies on a model's elaborate response and which, given the current trend, is not available for text-matching metrics. In contrast, our work intends most models. Current RAG (Retrieval Augmented to extract knowledge and understanding from Generation) systems rely on LLM's prompt memory LLMs without intentionally tricking or confusing to register some facts and expect the model the model.


Eliciting the Priors of Large Language Models using Iterated In-Context Learning

arXiv.org Artificial Intelligence

As Large Language Models (LLMs) are increasingly deployed in real-world settings, understanding the knowledge they implicitly use when making decisions is critical. One way to capture this knowledge is in the form of Bayesian prior distributions. We develop a prompt-based workflow for eliciting prior distributions from LLMs. Our approach is based on iterated learning, a Markov chain Monte Carlo method in which successive inferences are chained in a way that supports sampling from the prior distribution. We validated our method in settings where iterated learning has previously been used to estimate the priors of human participants -- causal learning, proportion estimation, and predicting everyday quantities. We found that priors elicited from GPT-4 qualitatively align with human priors in these settings. We then used the same method to elicit priors from GPT-4 for a variety of speculative events, such as the timing of the development of superhuman AI.


MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures

arXiv.org Artificial Intelligence

Evaluating large language models (LLMs) is challenging. Traditional ground-truth-based benchmarks fail to capture the comprehensiveness and nuance of real-world queries, while LLM-as-judge benchmarks suffer from grading biases and limited query quantity. Both of them may also become contaminated over time. User-facing evaluation, such as Chatbot Arena, provides reliable signals but is costly and slow. In this work, we propose MixEval, a new paradigm for establishing efficient, gold-standard LLM evaluation by strategically mixing off-the-shelf benchmarks. It bridges (1) comprehensive and well-distributed real-world user queries and (2) efficient and fairly-graded ground-truth-based benchmarks, by matching queries mined from the web with similar queries from existing benchmarks. Based on MixEval, we further build MixEval-Hard, which offers more room for model improvement. Our benchmarks' advantages lie in (1) a 0.96 model ranking correlation with Chatbot Arena arising from the highly impartial query distribution and grading mechanism, (2) fast, cheap, and reproducible execution (6% of the time and cost of MMLU), and (3) dynamic evaluation enabled by the rapid and stable data update pipeline. We provide extensive meta-evaluation and analysis for our and existing LLM benchmarks to deepen the community's understanding of LLM evaluation and guide future research directions.