Africa
Adaptation with Self-Evaluation to Improve Selective Prediction in LLMs
Chen, Jiefeng, Yoon, Jinsung, Ebrahimi, Sayna, Arik, Sercan O, Pfister, Tomas, Jha, Somesh
Large language models (LLMs) have recently shown great advances in a variety of tasks, including natural language understanding and generation. However, their use in high-stakes decision-making scenarios is still limited due to the potential for errors. Selective prediction is a technique that can be used to improve the reliability of the LLMs by allowing them to abstain from making predictions when they are unsure of the answer. In this work, we propose a novel framework for adaptation with self-evaluation to improve the selective prediction performance of LLMs. Our framework is based on the idea of using parameter-efficient tuning to adapt the LLM to the specific task at hand while improving its ability to perform self-evaluation. We evaluate our method on a variety of question-answering (QA) datasets and show that it outperforms state-of-the-art selective prediction methods. For example, on the CoQA benchmark, our method improves the AUACC from 91.23% to 92.63% and improves the AUROC from 74.61% to 80.25%.
MixPro: Simple yet Effective Data Augmentation for Prompt-based Learning
Li, Bohan, Dou, Longxu, Hou, Yutai, Feng, Yunlong, Mu, Honglin, Zhu, Qingfu, Sun, Qinghua, Che, Wanxiang
Prompt-based learning has shown considerable promise in reformulating various downstream tasks as cloze problems by combining original input with a predetermined template. This approach demonstrates its effectiveness, especially in few-shot learning scenarios, where the model is trained on a scarce amount of data. Despite its successes, the limited templates and text in few-shot prompt-based learning scenarios leave significant room for performance improvement. Moreover, existing methods sometimes resort to model ensembles, which, while effective, could potentially hamper model efficiency due to increased computational demands. To address these issues, we introduce MixPro, an augmentation method designed to augment both the vanilla input text and the templates. We implement this through the token-level, the sentence-level, and the template-level Mixup strategies. The experimental results on five few-shot datasets show that MixPro outperforms other augmentation baselines, improving model performance by an average of 5.08% compared to before augmentation.
Semantics-Empowered Communication: A Tutorial-cum-Survey
Lu, Zhilin, Li, Rongpeng, Lu, Kun, Chen, Xianfu, Hossain, Ekram, Zhao, Zhifeng, Zhang, Honggang
Along with the springing up of the semantics-empowered communication (SemCom) research, it is now witnessing an unprecedentedly growing interest towards a wide range of aspects (e.g., theories, applications, metrics and implementations) in both academia and industry. In this work, we primarily aim to provide a comprehensive survey on both the background and research taxonomy, as well as a detailed technical tutorial. Specifically, we start by reviewing the literature and answering the "what" and "why" questions in semantic transmissions. Afterwards, we present the ecosystems of SemCom, including history, theories, metrics, datasets and toolkits, on top of which the taxonomy for research directions is presented. Furthermore, we propose to categorize the critical enabling techniques by explicit and implicit reasoning-based methods, and elaborate on how they evolve and contribute to modern content & channel semantics-empowered communications. Besides reviewing and summarizing the latest efforts in SemCom, we discuss the relations with other communication levels (e.g., conventional communications) from a holistic and unified viewpoint. Subsequently, in order to facilitate future developments and industrial applications, we also highlight advanced practical techniques for boosting semantic accuracy, robustness, and large-scale scalability, just to mention a few. Finally, we discuss the technical challenges that shed light on future research opportunities.
Israel strikes Iran-backed terrorists in ongoing effort to stop new war front in West Bank
The IDF says it forces "destroyed an underground tunnel shaft containing ready-to-use explosive devices. It also said that "additional weapons were found, as well as ammunition and military equipment." JERUSALEM - Israel Defense Forces (IDF) on Wednesday launched a raid on the city of Jenin and its refugee camp - two strongholds of Palestinian terrorist activity - in the West Bank. The IDF operation in the West Bank, known by Israelis by its biblical name Judea and Samaria, raises questions about the opening of a third front in Israel's response to Hamas' multipronged attack against the Jewish state on Oct. 7, resulting in the massacre of 1,400 people in southern Israel. The IDF said in a statement that its counterterrorism forces "exchanged fire with armed terrorists, over ten terrorists were killed, and over 20 wanted suspects were apprehended, among them Nur and Minur Salma, Palestinian Islamic Jihad terrorists." The U.S. has designated the Iran-backed Palestinian Islamic Jihad a foreign terrorist organization. The fighting comes at a time when the Biden administration is cautioning Israeli actions in the West Bank, especially when it comes to violence from a small group of extremist settlers who have been involved in armed confrontations with Palestinian villagers in the area. NETANYAHU TELLS BRET BAIER CEASE-FIRE'MEANS SURRENDER,' INSISTS SQUAD MEMBER IS CALLING FOR'GENOCIDE' Palestinian terrorists take up position during a confrontation with the Israeli army in Jenin on July 3, 2023. The Israeli army said it had launched drone strikes in Jenin as part of an "extensive counterterrorism effort." U.S. Secretary of State Antony Blinken said on Monday in Tokyo at the G-7 meeting that "I briefed by (sic) colleagues about my conversations with Israeli leaders on pauses, and on concrete steps to minimize harm to Palestinian civilians in Gaza and to stop extremist violence in the West Bank." The Associated Press reported that President Biden said in late October the attacks by "extremist settlers" amounted to "pouring gasoline" on the already burning fires in the Middle East since the Hamas attack. The administration refers to Jewish residents who live in the disputed West Bank territory as settlers. Following Thursday's raid, the IDF added that "Two M-16 rifles, a'Carlo' gun, three handguns, ammunition, and military equipment were seized." The Palestinian-manufactured "Carlo" gun has its origins in the 2016 terrorism wave against Israelis. The weapon is a watered-down version of the Carl Gustav submachine gun - hence its name, the "Carlo" gun. "The initiative is always ours to prevent a third front.
US troops face further attacks in Iraq
United States troops in Iraq have been targeted in new attacks using drones and explosives, according to military and security sources. Three attacks took place on Thursday, the sources said, adding to the more than 40 assaults that US and allied troops based across the Middle East have come under since the Israel-Hamas war started on October 7. As well as two drone assaults at bases, a US-led coalition convoy was hit by an improvised explosive device (IED) blast in the vicinity of Mosul Dam. The security sources said the patrol was accompanied by Iraqi counterterrorism forces and that a vehicle in the patrol was damaged. Three US troops sustained minor injuries but had returned to duty, the official added.
Preference-conditioned Pixel-based AI Agent For Game Testing
Abdelfattah, Sherif, Brown, Adrian, Zhang, Pushi
The game industry is challenged to cope with increasing growth in demand and game complexity while maintaining acceptable quality standards for released games. Classic approaches solely depending on human efforts for quality assurance and game testing do not scale effectively in terms of time and cost. Game-testing AI agents that learn by interaction with the environment have the potential to mitigate these challenges with good scalability properties on time and costs. However, most recent work in this direction depends on game state information for the agent's state representation, which limits generalization across different game scenarios. Moreover, game test engineers usually prefer exploring a game in a specific style, such as exploring the golden path. However, current game testing AI agents do not provide an explicit way to satisfy such a preference. This paper addresses these limitations by proposing an agent design that mainly depends on pixel-based state observations while exploring the environment conditioned on a user's preference specified by demonstration trajectories. In addition, we propose an imitation learning method that couples self-supervised and supervised learning objectives to enhance the quality of imitation behaviors. Our agent significantly outperforms state-of-the-art pixel-based game testing agents over exploration coverage and test execution quality when evaluated on a complex open-world environment resembling many aspects of real AAA games.
Towards a Feminist Metaethics of AI
The proliferation of Artificial Intelligence (AI) has sparked an overwhelming number of AI ethics guidelines, boards and codes of conduct. These outputs primarily analyse competing theories, principles and values for AI development and deployment. However, as a series of recent problematic incidents about AI ethics/ethicists demonstrate, this orientation is insufficient. Before proceeding to evaluate other professions, AI ethicists should critically evaluate their own; yet, such an evaluation should be more explicitly and systematically undertaken in the literature. I argue that these insufficiencies could be mitigated by developing a research agenda for a feminist metaethics of AI. Contrary to traditional metaethics, which reflects on the nature of morality and moral judgements in a non-normative way, feminist metaethics expands its scope to ask not only what ethics is but also what our engagement with it should be like. Applying this perspective to the context of AI, I suggest that a feminist metaethics of AI would examine: (i) the continuity between theory and action in AI ethics; (ii) the real-life effects of AI ethics; (iii) the role and profile of those involved in AI ethics; and (iv) the effects of AI on power relations through methods that pay attention to context, emotions and narrative.
Image Classification using Combination of Topological Features and Neural Networks
Lima, Mariana Dória Prata, Giraldi, Gilson Antonio, Junior, Gastão Florêncio Miranda
In this work we use the persistent homology method, a technique in topological data analysis (TDA), to extract essential topological features from the data space and combine them with deep learning features for classification tasks. In TDA, the concepts of complexes and filtration are building blocks. Firstly, a filtration is constructed from some complex. Then, persistent homology classes are computed, and their evolution along the filtration is visualized through the persistence diagram. Additionally, we applied vectorization techniques to the persistence diagram to make this topological information compatible with machine learning algorithms. This was carried out with the aim of classifying images from multiple classes in the MNIST dataset. Our approach inserts topological features into deep learning approaches composed by single and two-streams neural networks architectures based on a multi-layer perceptron (MLP) and a convolutional neral network (CNN) taylored for multi-class classification in the MNIST dataset. In our analysis, we evaluated the obtained results and compared them with the outcomes achieved through the baselines that are available in the TensorFlow library. The main conclusion is that topological information may increase neural network accuracy in multi-class classification tasks with the price of computational complexity of persistent homology calculation. Up to the best of our knowledge, it is the first work that combines deep learning features and the combination of topological features for multi-class classification tasks.
Word Definitions from Large Language Models
Dictionary definitions are historically the arbitrator of what words mean, but this primacy has come under threat by recent progress in NLP, including word embeddings and generative models like ChatGPT. We present an exploratory study of the degree of alignment between word definitions from classical dictionaries and these newer computational artifacts. Specifically, we compare definitions from three published dictionaries to those generated from variants of ChatGPT. We show that (i) definitions from different traditional dictionaries exhibit more surface form similarity than do model-generated definitions, (ii) that the ChatGPT definitions are highly accurate, comparable to traditional dictionaries, and (iii) ChatGPT-based embedding definitions retain their accuracy even on low frequency words, much better than GloVE and FastText word embeddings.
Is it indeed bigger better? The comprehensive study of claim detection LMs applied for disinformation tackling
Hyben, Martin, Kula, Sebastian, Srba, Ivan, Moro, Robert, Simko, Jakub
This study compares the performance of (1) fine-tuned models and (2) extremely large language models on the task of check-worthy claim detection. For the purpose of the comparison we composed a multilingual and multi-topical dataset comprising texts of various sources and styles. Building on this, we performed a benchmark analysis to determine the most general multilingual and multi-topical claim detector. We chose three state-of-the-art models in the check-worthy claim detection task and fine-tuned them. Furthermore, we selected three state-of-the-art extremely large language models without any fine-tuning. We made modifications to the models to adapt them for multilingual settings and through extensive experimentation and evaluation. We assessed the performance of all the models in terms of accuracy, recall, and F1-score in in-domain and cross-domain scenarios. Our results demonstrate that despite the technological progress in the area of natural language processing, the models fine-tuned for the task of check-worthy claim detection still outperform the zero-shot approaches in a cross-domain settings.