Goto

Collaborating Authors

 Large Language Model


Bridging Local Details and Global Context in Text-Attributed Graphs

arXiv.org Artificial Intelligence

Representation learning on text-attributed graphs (TAGs) is vital for real-world applications, as they combine semantic textual and contextual structural information. Research in this field generally consist of two main perspectives: local-level encoding and global-level aggregating, respectively refer to textual node information unification (e.g., using Language Models) and structure-augmented modeling (e.g., using Graph Neural Networks). Most existing works focus on combining different information levels but overlook the interconnections, i.e., the contextual textual information among nodes, which provides semantic insights to bridge local and global levels. In this paper, we propose GraphBridge, a multi-granularity integration framework that bridges local and global perspectives by leveraging contextual textual information, enhancing fine-grained understanding of TAGs. Besides, to tackle scalability and efficiency challenges, we introduce a graphaware token reduction module. Extensive experiments across various models and datasets show that our method achieves state-of-theart performance, while our graph-aware token reduction module significantly enhances efficiency and solves scalability issues.


Rethinking Negative Instances for Generative Named Entity Recognition

arXiv.org Artificial Intelligence

Large Language Models (LLMs) have demonstrated impressive capabilities for generalizing in unseen tasks. In the Named Entity Recognition (NER) task, recent advancements have seen the remarkable improvement of LLMs in a broad range of entity domains via instruction tuning, by adopting entity-centric schema. In this work, we explore the potential enhancement of the existing methods by incorporating negative instances into training. Our experiments reveal that negative instances contribute to remarkable improvements by (1) introducing contextual information, and (2) clearly delineating label boundaries. Furthermore, we introduce an efficient longest common subsequence (LCS) matching algorithm, which is tailored to transform unstructured predictions into structured entities. By integrating these components, we present GNER, a Generative NER system that shows improved zero-shot performance across unseen entity domains. Our comprehensive evaluation illustrates our system's superiority, surpassing state-of-the-art (SoTA) methods by 9 $F_1$ score in zero-shot evaluation.


Beyond Natural Language: LLMs Leveraging Alternative Formats for Enhanced Reasoning and Communication

arXiv.org Artificial Intelligence

Natural language (NL) has long been the predominant format for human cognition and communication, and by extension, has been similarly pivotal in the development and application of Large Language Models (LLMs). Yet, besides NL, LLMs have seen various non-NL formats during pre-training, such as code and logical expression. NL's status as the optimal format for LLMs, particularly in single-LLM reasoning and multi-agent communication, has not been thoroughly examined. In this work, we challenge the default use of NL by exploring the utility of non-NL formats in these contexts. We show that allowing LLMs to autonomously select the most suitable format before reasoning or communicating leads to a 3.3 to 5.7\% improvement in reasoning efficiency for different LLMs, and up to a 72.7\% reduction in token usage in multi-agent communication, all while maintaining communicative effectiveness. Our comprehensive analysis further reveals that LLMs can devise a format from limited task instructions and that the devised format is effectively transferable across different LLMs. Intriguingly, the structured communication format decided by LLMs exhibits notable parallels with established agent communication languages, suggesting a natural evolution towards efficient, structured communication in agent communication. Our code is released at \url{https://github.com/thunlp/AutoForm}.


Translation Equivariant Transformer Neural Processes

arXiv.org Machine Learning

The effectiveness of neural processes (NPs) in modelling posterior prediction maps -- the mapping from data to posterior predictive distributions -- has significantly improved since their inception. This improvement can be attributed to two principal factors: (1) advancements in the architecture of permutation invariant set functions, which are intrinsic to all NPs; and (2) leveraging symmetries present in the true posterior predictive map, which are problem dependent. Transformers are a notable development in permutation invariant set functions, and their utility within NPs has been demonstrated through the family of models we refer to as TNPs. Despite significant interest in TNPs, little attention has been given to incorporating symmetries. Notably, the posterior prediction maps for data that are stationary -- a common assumption in spatio-temporal modelling -- exhibit translation equivariance. In this paper, we introduce of a new family of translation equivariant TNPs that incorporate translation equivariance. Through an extensive range of experiments on synthetic and real-world spatio-temporal data, we demonstrate the effectiveness of TE-TNPs relative to their non-translation-equivariant counterparts and other NP baselines.


BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models

arXiv.org Machine Learning

Large Language Models (LLMs) often suffer from overconfidence during inference, particularly when adapted to downstream domain-specific tasks with limited data. Previous work addresses this issue by employing approximate Bayesian estimation after the LLMs are trained, enabling them to quantify uncertainty. However, such post-training approaches' performance is severely limited by the parameters learned during training. In this paper, we go beyond post-training Bayesianization and propose Bayesian Low-Rank Adaptation by Backpropagation (BLoB), an algorithm that continuously and jointly adjusts both the mean and covariance of LLM parameters throughout the whole fine-tuning process. Our empirical results verify the effectiveness of BLoB in terms of generalization and uncertainty estimation, when evaluated on both in-distribution and out-of-distribution data.


Imagination Augmented Generation: Learning to Imagine Richer Context for Question Answering over Large Language Models

arXiv.org Artificial Intelligence

Retrieval-Augmented-Generation and Gener-ation-Augmented-Generation have been proposed to enhance the knowledge required for question answering over Large Language Models (LLMs). However, the former relies on external resources, and both require incorporating explicit documents into the context, which increases execution costs and susceptibility to noise data. Recent works indicate that LLMs have modeled rich knowledge, albeit not effectively triggered or awakened. Inspired by this, we propose a novel knowledge-augmented framework, Imagination-Augmented-Generation (IAG), which simulates the human capacity to compensate for knowledge deficits while answering questions solely through imagination, thereby awakening relevant knowledge in LLMs without relying on external resources. Guided by IAG, we propose an imagine richer context method for question answering (IMcQA). IMcQA consists of two modules: explicit imagination, which generates a short dummy document by learning from long context compression, and implicit imagination, which creates flexible adapters by distilling from a teacher model with a long context. Experimental results on three datasets demonstrate that IMcQA exhibits significant advantages in both open-domain and closed-book settings, as well as in out-of-distribution generalization. Our code will be available at https://github.com/Xnhyacinth/IAG.


Could brain-like computers be a 'competition killer'?

BBC News

Commercial applications envisaged fall into two main categories. One, which is where SpiNNcloud is focused, is in providing a more energy efficient and higher performance platform for AI applications โ€“ including image and video analysis, speech recognition and the large-language models that power chatbots such as ChatGPT. Another is in "edge computing" applications โ€“ where data is processed not in the cloud, but in real time on connected devices, but which operate on power constraints. Autonomous vehicles, robots, cell phones and wearable technology could all benefit. Long regarded as a main stumbling block to the advance of neuromorphic computing generally is developing the software needed for the chips to run.


OpenAI-Backed Nonprofits Have Gone Back on Their Transparency Pledges

WIRED

A Sam Altmanโ€“funded nonprofit studying the effects of giving monthly checks of up to 1,000 to lower-income households in the US espouses transparency in its operations. "We aim to share data, findings, and insights widely," OpenResearch says on its website, which describes its work as a "public good." But like at least two other Altman-linked organizations--OpenAI and UBI Charitable--OpenResearch has decided to withhold information about its finances and governance. In several years of filings to US tax authorities since their founding, each of the organizations has answered a question about their voluntary disclosure of financial statements, governing documents, and conflict-of-interest policies by stating that the public can review them upon request. It remains unclear whether anyone took them up on the offer in those years.


What happened when 20 comedians got AI to write their routines

MIT Technology Review

Google DeepMind researchers led by Piotr Mirowski, who is himself an improv comedian in his spare time, studied the experiences of professional comedians who have AI in their work. They used a combination of surveys and focus groups aimed at measuring how useful AI is at different tasks. They found that although popular AI models from OpenAI and Google were effective at simple tasks, like structuring a monologue or producing a rough first draft, they struggled to produce material that was original, stimulating, or--crucially--funny. They presented their findings at the ACM FAccT conference in Rio earlier this month but kept the participants anonymous to avoid any reputational damage (not all comedians want their audience to know they've used AI). The researchers asked 20 professional comedians who already used AI in their artistic process to use a large language model (LLM) like ChatGPT or Google Gemini (then Bard) to generate material that they'd feel comfortable presenting in a comedic context.


Why Microsoft, OpenAI and Nvidia are facing anti-monopoly probes

Al Jazeera

The United States Department of Justice and the Federal Trade Commission (FTC) have reportedly reached a deal on how they will pursue an antitrust investigation into tech giants Microsoft, Nvidia, and Open AI. The companies are all major players in generative AI: OpenAI is the nonprofit startup behind ChatGPT, the blockbuster AI-powered chatbot. Microsoft, the world's largest company by market capitalisation, has invested more than 13bn in OpenAI and holds a 49 percent stake in the company's for-profit subsidiary. Chipmaker Nvidia is a global leader in graphic processing units (GPU), a key piece of hardware needed in AI. The company recently hit a 3 trillion valuation, surpassing Apple to become the world's second-largest company.