Goto

Collaborating Authors

 archer



AI models misrepresent news events nearly half the time, study says

Al Jazeera

AI models such as ChatGPT routinely misrepresent news events, providing faulty responses to questions almost half the time, a study has found. The study published on Wednesday by the European Broadcasting Union (EBU) and the BBC assessed the accuracy of more than 2,700 responses given by OpenAI's ChatGPT, Google's Gemini, Microsoft's Copilot, and Perplexity. Overall, 45 percent of responses had at least one "significant" issue, according to the research. Sourcing was the most common problem, with 31 percent of responses including information not supported by the cited source, or incorrect or unverifiable attribution, among other issues. A lack of accuracy was the next biggest contributor to faulty answers, affecting 20 percent of responses, followed by the absence of appropriate context, with 14 percent.


Harnessing Language for Coordination: A Framework and Benchmark for LLM-Driven Multi-Agent Control

arXiv.org Artificial Intelligence

Large Language Models (LLMs) have demonstrated remarkable performance across various tasks. A promising but largely under-explored area is their potential to facilitate human coordination with many agents. Such capabilities would be useful in domains including disaster response, urban planning, and real-time strategy scenarios. In this work, we introduce (1) a real-time strategy game benchmark designed to evaluate these abilities and (2) a novel framework we term HIVE. HIVE empowers a single human to coordinate swarms of up to 2,000 agents using natural language dialog with an LLM. We present promising results on this multi-agent benchmark, with our hybrid approach solving tasks such as coordinating agent movements, exploiting unit weaknesses, leveraging human annotations, and understanding terrain and strategic points. However, our findings also highlight critical limitations of current models, including difficulties in processing spatial visual information and challenges in formulating long-term strategic plans. This work sheds light on the potential and limitations of LLMs in human-swarm coordination, paving the way for future research in this area. The HIVE project page, which includes videos of the system in action, can be found here: hive.syrkis.com.


Safety-Critical Controller Synthesis with Reduced-Order Models

arXiv.org Artificial Intelligence

Reduced-order models (ROMs) provide lower dimensional representations of complex systems, capturing their salient features while simplifying control design. Building on previous work, this paper presents an overarching framework for the integration of ROMs and control barrier functions, enabling the use of simplified models to construct safety-critical controllers while providing safety guarantees for complex full-order models. To achieve this, we formalize the connection between full and ROMs by defining projection mappings that relate the states and inputs of these models and leverage simulation functions to establish conditions under which safety guarantees may be transferred from a ROM to its corresponding full-order model. The efficacy of our framework is illustrated through simulation results on a drone and hardware demonstrations on ARCHER, a 3D hopping robot.


ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

arXiv.org Artificial Intelligence

A broad use case of large language models (LLMs) is in goal-directed decision-making tasks (or "agent" tasks), where an LLM needs to not just generate completions for a given prompt, but rather make intelligent decisions over a multi-turn interaction to accomplish a task (e.g., when interacting with the web, using tools, or providing customer support). Reinforcement learning (RL) provides a general paradigm to address such agent tasks, but current RL methods for LLMs largely focus on optimizing single-turn rewards. By construction, most single-turn RL methods cannot endow LLMs with the ability to intelligently seek information over multiple turns, perform credit assignment, or reason about their past actions -- all of which are critical in agent tasks. This raises the question: how can we design effective and efficient multi-turn RL algorithms for LLMs? In this paper, we develop a framework for building multi-turn RL algorithms for fine-tuning LLMs, that preserves the flexibility of existing single-turn RL methods for LLMs (e.g., proximal policy optimization), while accommodating multiple turns, long horizons, and delayed rewards effectively. To do this, our framework adopts a hierarchical RL approach and runs two RL algorithms in parallel: a high-level off-policy value-based RL algorithm to aggregate reward over utterances, and a low-level RL algorithm that utilizes this high-level value function to train a token policy within each utterance or turn. Our hierarchical framework, Actor-Critic Framework with a Hierarchical Structure (ArCHer), can also give rise to other RL methods. Empirically, we find that ArCHer significantly improves efficiency and performance on agent tasks, attaining a sample efficiency of about 100x over existing methods, while also improving with larger model capacity (upto the 7 billion scale that we tested on).


Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning

arXiv.org Artificial Intelligence

We present Archer, a challenging bilingual text-to-SQL dataset specific to complex reasoning, including arithmetic, commonsense and hypothetical reasoning. It contains 1,042 English questions and 1,042 Chinese questions, along with 521 unique SQL queries, covering 20 English databases across 20 domains. Notably, this dataset demonstrates a significantly higher level of complexity compared to existing publicly available datasets. Our evaluation shows that Archer challenges the capabilities of current state-of-the-art models, with a high-ranked model on the Spider leaderboard achieving only 6.73% execution accuracy on Archer test set. Thus, Archer presents a significant challenge for future research in this field.


Assassin's Creed Mirage preview: Finally, a return to stealth roots

PCWorld

Assassin's Creed Mirage is a dream for stealth kings. People who loved Sam Fisher in Splinter Cell or simply the old Assassin's Creeds will have a tremendous fun in beautiful 9th century Baghdad, our recent hands-on with the game revealed. We throw coins, briefly distract a guard, dart around corners. In that game, we are a bear of a man, with arms like tree trunks as we swing the axe and make the English army tremble. Valhalla also had its moments, but in Mirage there is much more of a hand-built feel.


MHfit: Mobile Health Data for Predicting Athletics Fitness Using Machine Learning

arXiv.org Artificial Intelligence

Mobile phones and other electronic gadgets or devices have aided in collecting data without the need for data entry. This paper will specifically focus on Mobile health data. Mobile health data use mobile devices to gather clinical health data and track patient vitals in real-time. Our study is aimed to give decisions for small or big sports teams on whether one athlete good fit or not for a particular game with the compare several machine learning algorithms to predict human behavior and health using the data collected from mobile devices and sensors placed on patients. In this study, we have obtained the dataset from a similar study done on mhealth. The dataset contains vital signs recordings of ten volunteers from different backgrounds. They had to perform several physical activities with a sensor placed on their bodies. Our study used 5 machine learning algorithms (XGBoost, Naive Bayes, Decision Tree, Random Forest, and Logistic Regression) to analyze and predict human health behavior. XGBoost performed better compared to the other machine learning algorithms and achieved 95.2% accuracy, 99.5% in sensitivity, 99.5% in specificity, and 99.66% in F1 score. Our research indicated a promising future in mhealth being used to predict human behavior and further research and exploration need to be done for it to be available for commercial use specifically in the sports industry.


OpenAI and Figure join the race to humanoid robot workers

#artificialintelligence

The jarring emergence of ChatGPT has made it clear: AIs are advancing at a wild and accelerating pace, and they're beginning to transform industries based around desk jobs that typically marshall human intelligence. They'll begin taking over portions of many white-collar jobs in the coming years, leading initially to huge increases in productivity, and eventually, many believe, to huge increases in unemployment. If you're coming out of school right now and looking to be useful, blue collar work involving actual physical labor might be a better bet than anything that'd put you behind a desk. But on the other hand, it's starting to look like a general-purpose humanoid robot worker might be closer than anyone thinks, imbued with light-speed, swarm-based learning capabilities to go along with GPT-version-X communication abilities, a whole internet's worth of knowledge, and whatever physical attributes you need for a given job. Such humanoids will begin as dumbass job-site apprentices with zero common sense, but they'll learn – at a frightening pace, if the last few months in AI has been any kind of indication.


Anyone for AIPA? Detroit brewery lets ChatGPT create its latest BEER

Daily Mail - Science & tech

AI tool ChatGPT has already been used to write essays, prescribe antibiotics and even fool job recruiters. Now, Detroit-based Atwater Brewery has created the first ever beer using a recipe fully generated by the chatbot sensation. The new brew, called Artificial Intelligence IPA, or AI IPA for short, contains three types of malt and a whopping eight varieties of hops. Although human brewers had to make the beer themselves, the entire process was based on ChatGPT's detailed recipe and instructions. ChatGPT, created by San Francisco-based company OpenAI, has been trained on a massive amount of text so it can generate human-like answers to questions.