Africa
AUTOACT: Automatic Agent Learning from Scratch via Self-Planning
Qiao, Shuofei, Zhang, Ningyu, Fang, Runnan, Luo, Yujie, Zhou, Wangchunshu, Jiang, Yuchen Eleanor, Lv, Chengfei, Chen, Huajun
Language agents have achieved considerable performance on various complex tasks. Despite the incessant exploration in this field, existing language agent systems still struggle with costly, non-reproducible data reliance and face the challenge of compelling a single model for multiple functions. To this end, we introduce AutoAct, an automatic agent learning framework that does not rely on large-scale annotated data and synthetic trajectories from closed-source models (e.g., GPT-4). Given limited data with a tool library, AutoAct first automatically synthesizes planning trajectories without any assistance from humans or strong closed-source models. Then, AutoAct leverages a division-of-labor strategy to automatically differentiate based on the target task information and synthesized trajectories, producing a sub-agent group to complete the task. We conduct comprehensive experiments with different LLMs, which demonstrates that AutoAct yields better or parallel performance compared to various strong baselines. We even notice that AutoAct, when using the Llama-2-13b model, can achieve performance comparable to that of the zero-shot GPT-3.5-Turbo agent. Code will be available at https://github.com/zjunlp/AutoAct.
MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
Pi, Renjie, Han, Tianyang, Xie, Yueqi, Pan, Rui, Lian, Qing, Dong, Hanze, Zhang, Jipeng, Zhang, Tong
The deployment of multimodal large language models (MLLMs) has brought forth a unique vulnerability: susceptibility to malicious attacks through visual inputs. We delve into the novel challenge of defending MLLMs against such attacks. We discovered that images act as a "foreign language" that is not considered during alignment, which can make MLLMs prone to producing harmful responses. Unfortunately, unlike the discrete tokens considered in text-based LLMs, the continuous nature of image signals presents significant alignment challenges, which poses difficulty to thoroughly cover the possible scenarios. This vulnerability is exacerbated by the fact that open-source MLLMs are predominantly fine-tuned on limited image-text pairs that is much less than the extensive text-based pretraining corpus, which makes the MLLMs more prone to catastrophic forgetting of their original abilities during explicit alignment tuning. To tackle these challenges, we introduce MLLM-Protector, a plug-and-play strategy combining a lightweight harm detector and a response detoxifier. The harm detector's role is to identify potentially harmful outputs from the MLLM, while the detoxifier corrects these outputs to ensure the response stipulates to the safety standards. This approach effectively mitigates the risks posed by malicious visual inputs without compromising the model's overall performance. Our results demonstrate that MLLM-Protector offers a robust solution to a previously unaddressed aspect of MLLM security.
CLadder: Assessing Causal Reasoning in Language Models
Jin, Zhijing, Chen, Yuen, Leeb, Felix, Gresele, Luigi, Kamal, Ojasv, Lyu, Zhiheng, Blin, Kevin, Adauto, Fernando Gonzalez, Kleiman-Weiner, Max, Sachan, Mrinmaya, Schรถlkopf, Bernhard
The ability to perform causal reasoning is widely considered a core feature of intelligence. In this work, we investigate whether large language models (LLMs) can coherently reason about causality. Much of the existing work in natural language processing (NLP) focuses on evaluating commonsense causal reasoning in LLMs, thus failing to assess whether a model can perform causal inference in accordance with a set of well-defined formal rules. To address this, we propose a new NLP task, causal inference in natural language, inspired by the "causal inference engine" postulated by Judea Pearl et al. We compose a large dataset, CLadder, with 10K samples: based on a collection of causal graphs and queries (associational, interventional, and counterfactual), we obtain symbolic questions and ground-truth answers, through an oracle causal inference engine. These are then translated into natural language. We evaluate multiple LLMs on our dataset, and we introduce and evaluate a bespoke chain-of-thought prompting strategy, CausalCoT. We show that our task is highly challenging for LLMs, and we conduct an in-depth analysis to gain deeper insights into the causal reasoning abilities of LLMs. Our data is open-sourced at https://huggingface.co/datasets/causalNLP/cladder, and our code can be found at https://github.com/causalNLP/cladder.
Genetic Algorithm enhanced by Deep Reinforcement Learning in parent selection mechanism and mutation : Minimizing makespan in permutation flow shop scheduling problems
Irmouli, Maissa, Benazzoug, Nourelhouda, Adimi, Alaa Dania, Rezkellah, Fatma Zohra, Hamzaoui, Imane, Hamitouche, Thanina, Bessedik, Malika, Tayeb, Fatima Si
This paper introduces a reinforcement learning (RL) approach to address the challenges associated with configuring and optimizing genetic algorithms (GAs) for solving difficult combinatorial or non-linear problems. The proposed RL+GA method was specifically tested on the flow shop scheduling problem (FSP). The hybrid algorithm incorporates neural networks (NN) and uses the off-policy method Q-learning or the on-policy method Sarsa(0) to control two key genetic algorithm (GA) operators: parent selection mechanism and mutation. At each generation, the RL agent's action is determining the selection method, the probability of the parent selection and the probability of the offspring mutation. This allows the RL agent to dynamically adjust the selection and mutation based on its learned policy. The results of the study highlight the effectiveness of the RL+GA approach in improving the performance of the primitive GA. They also demonstrate its ability to learn and adapt from population diversity and solution improvements over time. This adaptability leads to improved scheduling solutions compared to static parameter configurations while maintaining population diversity throughout the evolutionary process.
Watch Your Language: Investigating Content Moderation with Large Language Models
Kumar, Deepak, AbuHashem, Yousef, Durumeric, Zakir
Large language models (LLMs) have exploded in popularity due to their ability to perform a wide array of natural language tasks. Text-based content moderation is one LLM use case that has received recent enthusiasm, however, there is little research investigating how LLMs perform in content moderation settings. In this work, we evaluate a suite of commodity LLMs on two common content moderation tasks: rule-based community moderation and toxic content detection. For rule-based community moderation, we instantiate 95 subcommunity specific LLMs by prompting GPT-3.5 with rules from 95 Reddit subcommunities. We find that GPT-3.5 is effective at rule-based moderation for many communities, achieving a median accuracy of 64% and a median precision of 83%. For toxicity detection, we evaluate a suite of commodity LLMs (GPT-3, GPT-3.5, GPT-4, Gemini Pro, LLAMA 2) and show that LLMs significantly outperform currently widespread toxicity classifiers. However, recent increases in model size add only marginal benefit to toxicity detection, suggesting a potential performance plateau for LLMs on toxicity detection tasks. We conclude by outlining avenues for future work in studying LLMs and content moderation.
It is time to use Russia's frozen assets to help Ukraine
An estimated 350bn in Russian government assets have been frozen in Western accounts since Russian President Vladimir Putin ordered a full-scale invasion of Ukraine on February 24, 2022. These are not idle funds. In 2023, Belgium-based financial services company Euroclear, whose settling and clearance role mean that it holds 197 billion euros ( 214bn) in such assets, reported that they produced at least 3 billion euros ( 3.26bn) from interest. Given that the sanctions on the Kremlin remain firmly in place and Putin has shown no willingness to negotiate on his demand to annex one-quarter of Ukraine's territory or to cease his attacks, how these assets can be harnessed to push for an end to the war or help Ukraine resist has become a key question for Kyiv's Western allies. British Foreign Secretary David Cameron publicly opened the doors to the idea last December by stating: "Instead of just freezing that money, let's take that money, [and] spend it on rebuilding Ukraine."
Zelenskyy makes urgent call for support at World Economic Forum at Davos
Ukrainian President Volodomyr Zelenskyy gives his outlook on the conflict and offers an update on his country's counter-offensive on'Special Report.' Ukrainian President Volodymyr Zelenskyy huddled with corporate executives and world leaders in a frenzied first full day of the World Economic Forum's annual meeting in the Swiss ski resort of Davos, where top officials from the United States, European Union, China, the Middle East and beyond spoke Tuesday about tackling conflict and embracing technology like artificial intelligence. Zelenskyy is endeavoring to keep his country's long and largely stalemated defense against Russia on the minds of political leaders, just as Israel's war with Hamas, which passed the 100-day mark this week, has siphoned off much of the world's attention and sparked concerns about a wider conflict in the Middle East. "It is important that you stand with us, I thank you for your support. It is very important to be here, to boost investment in Ukraine and support our economy," Zelenskyy said at an invitation-only "CEOs for Ukraine" session, according to his office.
US Navy announces first seizure of Iranian weapons bound for Yemen as two SEALs remain lost from mission
The U.S. Navy on Tuesday announced what's considered the first seizure of Iranian weapons bound for Yemen since Houthi rebels began their campaign of attacks against international merchant shipping in the Red Sea two months ago โ yet the two Navy SEALs lost at sea during the mission carried out last week still remain missing amid search and rescue efforts. On Jan. 11, 2024, while conducting a flag verification, U.S. CENTCOM Navy forces "conducted a night-time seizure of a dhow conducting illegal transport of advanced lethal aid from Iran to resupply Houthi forces in Yemen as part of the Houthis' ongoing campaign of attacks against international merchant shipping," U.S. Central Command said in a statement Tuesday. "U.S. Navy SEALs operating from USS Lewis B Puller (ESB 3), supported by helicopters and unmanned aerial vehicles (UAVs), executed a complex boarding of the dhow near the coast of Somalia in international waters of the Arabian Sea, seizing Iranian-made ballistic missile and cruise missiles components," the statement said. "Seized items include propulsion, guidance, and warheads for Houthi medium range ballistic missiles (MRBMs) and anti-ship cruise missiles (ASCMs), as well as air defense associated components." On Jan. 10, 2024, a dhow was identified, and an assessment was made that the dhow was in the process of smuggling.
Three armed drones intercepted and shot down near US base in northern Iraq
Senior foreign affairs correspondent Greg Palkot provides details on the major strike on an Iraqi militia leader and the U.S.'s response to Houthi attacks in the Red Sea Three armed drones were shot down in Iraq on Tuesday, near where U.S. and other international forces are stationed, officials said. Iraqi Kurdistan's counter-terrorism service said its forces intercepted and shot down the drones over Erbil airport in northern Iraq at around 5:05 a.m. It did not say if there were any casualties or damage to infrastructure. There was no immediate claim of responsibility. Similar previous attacks have been claimed by a group called the Islamic Resistance in Iraq, an umbrella group of Iran-aligned Iraqi militias.
What doom loop? With AI, a 'spirit of optimism' returns to San Francisco start-ups
Far from the palm trees of Miami or Austin's taco trucks, Catalin Voss has headquartered his literacy start-up between a cannabis club and pawn shop in the heart of the Mission District. Voss rents a nondescript office building in one of San Francisco's most vibrant neighborhoods as a home base for Ello, a company he co-founded in 2020 that uses speech recognition technology, powered by artificial intelligence, to help struggling students develop their reading skills. The office is within walking distance of his Noe Valley apartment and only steps away from some of the city's best taquerias and cocktail bars. And those are just a few of the perks he recited in explaining why he is headquartered in San Francisco. Voss is part of a sizable cohort of San Francisco loyalists -- old and new -- who say they are flummoxed by the "all is lost" narrative propagated by conservative media hosts and more recently a vocal contingent of tech leaders that includes billionaire entrepreneur-turned-agitator Elon Musk.