Goto

Collaborating Authors

 Large Language Model


AdaCoT: Rethinking Cross-Lingual Factual Reasoning through Adaptive Chain-of-Thought

arXiv.org Artificial Intelligence

Large language models (LLMs) have shown impressive multilingual capabilities through pretraining on diverse corpora. While these models show strong reasoning abilities, their performance varies significantly across languages due to uneven training data distribution. Existing approaches using machine translation, and extensive multilingual pretraining and cross-lingual tuning face scalability challenges and often fail to capture nuanced reasoning processes across languages. In this paper, we introduce AdaCoT (Adaptive Chain-of-Thought), a framework that enhances multilingual reasoning by dynamically routing thought processes through intermediary "thinking languages" before generating target-language responses. AdaCoT leverages a language-agnostic core and incorporates an adaptive, reward-based mechanism for selecting optimal reasoning pathways without requiring additional pretraining. Our comprehensive evaluation across multiple benchmarks demonstrates substantial improvements in both factual reasoning quality and cross-lingual consistency, with particularly strong performance gains in low-resource language settings. The results suggest that adaptive reasoning paths can effectively bridge the performance gap between high and low-resource languages while maintaining cultural and linguistic nuances.


Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs

arXiv.org Artificial Intelligence

Alignment in large language models (LLMs) is used to enforce guidelines such as safety. Yet, alignment fails in the face of jailbreak attacks that modify inputs to induce unsafe outputs. In this paper, we present and evaluate a method to assess the robustness of LLM alignment. We observe that alignment embeds a safety classifier in the target model that is responsible for deciding between refusal and compliance. We seek to extract an approximation of this classifier, called a surrogate classifier, from the LLM. We develop an algorithm for identifying candidate classifiers from subsets of the LLM model. We evaluate the degree to which the candidate classifiers approximate the model's embedded classifier in benign (F1 score) and adversarial (using surrogates in a white-box attack) settings. Our evaluation shows that the best candidates achieve accurate agreement (an F1 score above 80%) using as little as 20% of the model architecture. Further, we find attacks mounted on the surrogate models can be transferred with high accuracy. For example, a surrogate using only 50% of the Llama 2 model achieved an attack success rate (ASR) of 70%, a substantial improvement over attacking the LLM directly, where we only observed a 22% ASR. These results show that extracting surrogate classifiers is a viable (and highly effective) means for modeling (and therein addressing) the vulnerability of aligned models to jailbreaking attacks.


Reviews: Dual Adversarial Semantics-Consistent Network for Generalized Zero-Shot Learning

Neural Information Processing Systems

The primary contribution of the paper is the dual-GAN structure with semantics-consistency and visual-consistency loss. The paper has novel components (although it is close to [7], see below). The paper shows that their model performs better than the existing methods on GZSL benchmark datasets. To better understand the individual contribution of these losses, the paper gives an ablation study in Table 3. However, it should also include results of ablation study on CUB and SUN datasets in the main paper.


Reviews: Dual Adversarial Semantics-Consistent Network for Generalized Zero-Shot Learning

Neural Information Processing Systems

The paper received all accept recommendations and the AC agrees with the recommendation. The authors are requested to revise the paper with the additional clarifications wrt to [7], results from the new experiments on FLO, and discussion points.


Review for NeurIPS paper: Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design

Neural Information Processing Systems

Additional Feedback: Could you add a specific example/ problem that would be easily solved by defining it as a UED? I think it would help the paper in general. The agent for the Lava environment, that replaces walls with dangerous lava, is trained from generated maps with lava instead of walls? Additionally, the paper needs further proof reading, some minor mistakes I found: *Line 511: I wouldn't start a proof section saying "it would be nice to know that..." that is too informal *Line 512: "their" should be "its" *Line 40: Section?, i.e., referenced section is missing the number *Line 138: the function T M shouldn't be defined on S M? *Line 171: This sentence needs further explanation *Line 207: based on twice *Line 209: Figure? The Broader Impact section, specially the first paragraph is too speculative, automating jobs or automated weapons are general problems of the AI field, it should focus more on the impact of this specific work.


Review for NeurIPS paper: Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design

Neural Information Processing Systems

This paper pursues a significant line of enquiry regarding an important topic: automatic, unsupervised environment design. The paper makes algorithmic, theoretical, and empirical contributions. While the reviewers had some concerns about the clarity of the theory and the adequacy of the empirical results, these have been well addressed in the rebuttal. The authors are strongly urged to incorporate all the reviewers' feedback in the final version.


Review for NeurIPS paper: Investigating Gender Bias in Language Models Using Causal Mediation Analysis

Neural Information Processing Systems

The paper studies the problem of bias in neural models where the proposed solution is based on causal mediation analysis. The focus of the paper is on pre-trained transformer language models, GPT-2. The proposed method of using mediation analysis for analyzing attention heads and neurons through interventions is novel and interesting, and can be generalized to other types of biases. The paper is well-written, and experiments are thorough.


It's only 30 to learn how to automate your job with AI

PCWorld

TL;DR: Learn how to use AI to simplify your job with the ChatGPT and Automation E-Degree on sale for 29.99 (reg. For many, AI tools kind of came out of nowhere. It can be overwhelming, especially if your industry is rapidly changing, but you don't have to get swept away in the AI sprint. Instead, check out the ChatGPT and Automation E-Degree for beginner-friendly guidance to help you use AI tools to streamline your workload. This course has 12 lectures and about 25 hours of content that you can access anywhere at any time.


This AI is so good, it's like getting ChatGPT-5 early

PCWorld

TL;DR: Get a 1minAI lifetime subscription for 99.99 before codes sell out--less than 50 are left in stock. Rumor has it that ChatGPT-5 will be released this year and be more powerful than ever before, but why wait? An even better tool combines all of the top AI models into one dashboard, so you get a taste of everything--GPT, Claude, Gemini, Llama, and more--without any recurring fees. Instead of hopping between multiple platforms, 1minAI has templates for whatever you need in one place. Need to generate an AI image?


ESGSenticNet: A Neurosymbolic Knowledge Base for Corporate Sustainability Analysis

arXiv.org Artificial Intelligence

Evaluating corporate sustainability performance is essential to drive sustainable business practices, amid the need for a more sustainable economy. However, this is hindered by the complexity and volume of corporate sustainability data (i.e. sustainability disclosures), not least by the effectiveness of the NLP tools used to analyse them. To this end, we identify three primary challenges - immateriality, complexity, and subjectivity, that exacerbate the difficulty of extracting insights from sustainability disclosures. To address these issues, we introduce ESGSenticNet, a publicly available knowledge base for sustainability analysis. ESGSenticNet is constructed from a neurosymbolic framework that integrates specialised concept parsing, GPT-4o inference, and semi-supervised label propagation, together with a hierarchical taxonomy. This approach culminates in a structured knowledge base of 44k knowledge triplets - ('halve carbon emission', supports, 'emissions control'), for effective sustainability analysis. Experiments indicate that ESGSenticNet, when deployed as a lexical method, more effectively captures relevant and actionable sustainability information from sustainability disclosures compared to state of the art baselines. Besides capturing a high number of unique ESG topic terms, ESGSenticNet outperforms baselines on the ESG relatedness and ESG action orientation of these terms by 26% and 31% respectively. These metrics describe the extent to which topic terms are related to ESG, and depict an action toward ESG. Moreover, when deployed as a lexical method, ESGSenticNet does not require any training, possessing a key advantage in its simplicity for non-technical stakeholders.