Goto

Collaborating Authors

 Large Language Model


Wisdom of the Crowds in Forecasting: Forecast Summarization for Supporting Future Event Prediction

arXiv.org Artificial Intelligence

Future Event Prediction (FEP) is an essential activity whose demand and application range across multiple domains. While traditional methods like simulations, predictive and time-series forecasting have demonstrated promising outcomes, their application in forecasting complex events is not entirely reliable due to the inability of numerical data to accurately capture the semantic information related to events. One forecasting way is to gather and aggregate collective opinions on the future to make predictions as cumulative perspectives carry the potential to help estimating the likelihood of upcoming events. In this work, we organize the existing research and frameworks that aim to support future event prediction based on crowd wisdom through aggregating individual forecasts. We discuss the challenges involved, available datasets, as well as the scope of improvement and future research directions for this task. We also introduce a novel data model to represent individual forecast statements.


Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies

arXiv.org Artificial Intelligence

Large language models (LLMs) can produce erroneous responses that sound fluent and convincing, raising the risk that users will rely on these responses as if they were correct. Mitigating such overreliance is a key challenge. Through a think-aloud study in which participants use an LLM-infused application to answer objective questions, we identify several features of LLM responses that shape users' reliance: explanations (supporting details for answers), inconsistencies in explanations, and sources. Through a large-scale, pre-registered, controlled experiment (N=308), we isolate and study the effects of these features on users' reliance, accuracy, and other measures. We find that the presence of explanations increases reliance on both correct and incorrect responses. However, we observe less reliance on incorrect responses when sources are provided or when explanations exhibit inconsistencies. We discuss the implications of these findings for fostering appropriate reliance on LLMs.


ClipRover: Zero-shot Vision-Language Exploration and Target Discovery by Mobile Robots

arXiv.org Artificial Intelligence

Vision-language navigation (VLN) has emerged as a promising paradigm, enabling mobile robots to perform zero-shot inference and execute tasks without specific pre-programming. However, current systems often separate map exploration and path planning, with exploration relying on inefficient algorithms due to limited (partially observed) environmental information. In this paper, we present a novel navigation pipeline named ''ClipRover'' for simultaneous exploration and target discovery in unknown environments, leveraging the capabilities of a vision-language model named CLIP. Our approach requires only monocular vision and operates without any prior map or knowledge about the target. For comprehensive evaluations, we design the functional prototype of a UGV (unmanned ground vehicle) system named ''Rover Master'', a customized platform for general-purpose VLN tasks. We integrate and deploy the ClipRover pipeline on Rover Master to evaluate its throughput, obstacle avoidance capability, and trajectory performance across various real-world scenarios. Experimental results demonstrate that ClipRover consistently outperforms traditional map traversal algorithms and achieves performance comparable to path-planning methods that depend on prior map and target knowledge. Notably, ClipRover offers real-time active navigation without requiring pre-captured candidate images or pre-built node graphs, addressing key limitations of existing VLN pipelines.


Reviews: Unified Language Model Pre-training for Natural Language Understanding and Generation

Neural Information Processing Systems

This paper provides a method to pretrain a single Transformer architecture on three objectives: (i) unidirectional language model (e.g. This unified architecture circumvents the shortcoming of both models like BERT (which can condition on bidirectional context, but harder to use for downstream tasks that involve generation due to bidirectionality) and GPT-2 (easy to apply for generation tasks since it works left-to-right, but bidirectional encoders have been known to work much better than unidirectional ones in sequence-to-sequence models), and thereby combines the best of both worlds. This is done using a simple masking scheme that restricts which words the model can pay attention to, depending on which objective function is used (e.g. if using a unidirectional, left-to-right objective, then all tokens to the right of the target word are masked out). Experiments on text summarisation (CNN/DailyMail and Gigaword), question answering (SQuAD, CoQA extractive, and CoQA abstractive), question generation, and GLUE indicate that the proposed pretraining approach largely matches or surpasses the current state of the art. Their masking approach crucially enables pretraining the two key ingredients of sequence-to-sequence models with a single model: (i) a bidirectional encoder, and (ii) a unidirectional decoder.


Review for NeurIPS paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Neural Information Processing Systems

Summary and Contributions: This paper proposes a retrieval augmented seq2seq model for question answering and related knowledge-intensive NLP tasks. The model is combination of a pre-trained BART and a dense passage retriever via joint probabilistic model. Two specific formulations, referred to as RAG-Sequence and RAG-Token, are proposed to let the model select relevant document(s) to generate answers. Experiments are conducted on a range of tasks including open-domain question answering and fact verification, showing that the RAG model achieve state-of-the-art or competitive performances. The design of the model share some similarity with REALM model, which is also a retrieval augmented encoder-only model.


Review for NeurIPS paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Neural Information Processing Systems

This work proposed a system that uses the retrieval results of query to aid the generation of answers. The idea is generally natural and has been explored by quite some authors in various ways. The paper is clearly written, and I enjoyed reading it. I would see this work as a nice piece of work that combines several existing models in a neat although not strikingly novel or inspirational way. The major downside of the work is its novelty, but its strong empirical results and potential impact on practice are enough to support its acceptance by NeurIPS.


Sam Altman Dismisses Elon Musk's Bid to Buy OpenAI in Letter to Staff

WIRED

Sam Altman is leaving no room for doubt about his views on an Elon Musk-led bid to take control of OpenAI. In a letter to OpenAI staff Monday, the CEO put the words "bid" and "deal" in scare quotes and said the startup's board has no interest in the offer. "Our structure exists to ensure that no individual can take control of OpenAI," Altman wrote, according to two sources with knowledge of the letter. "Elon runs a competitive AI company, and his actions are not about OpenAI's mission or values." Altman has also told employees that OpenAI's board, which he sits on, has yet to receive an official offer from Musk and the other investors.


Can simplifying AI rules in Europe create competition for US and China?

Al Jazeera

Can simplifying AI rules in Europe create competition for US and China? Can simplifying AI rules in Europe create competition for US and China? Europe to cut red tape to make artificial intelligence advancements easier.Read more The Artificial Intelligence Action Summit in Paris has drawn nearly 100 world leaders and tech firms, and the consensus is that 2025 is not the year for new AI regulations. France says it is time to simplify the rules in Europe to allow AI advances – or risk being left behind. Which countries have banned DeepSeek and why? list 2 of 3 Elon Musk-led group makes 97.4bn bid for OpenAI list 3 of 3 In January, Chinese start-up DeepSeek disrupted Wall Street and Silicon Valley.


Microsoft wants to hand off much of its Army HoloLens program to Palmer Luckey's Anduril

Engadget

Microsoft's six-year-old program to make HoloLens headsets for the US Army could be getting some extra help. If the Department of Defense approves the deal, the company will expand its existing partnership with Anduril Industries, Palmer Luckey's defense startup, for the next stages of the Integrated Visual Augmentation System (IVAS) program. Microsoft, which spearheaded the program, would transition into supplying AI and cloud infrastructure. Meanwhile, Anduril would do pretty much everything else, including "oversight of production, future development of hardware and software and delivery timelines." Anduril makes a wide array of defense tech, including drone interceptors, sentry towers, comms jammers, drones and even an autonomous submarine. But given Luckey's background as the primary inventor of the Oculus Rift -- and, by extension, the modern consumer XR industry -- the IVAS program could perhaps be the defense tech startup's most natural fit.


Language Is Not All You Need: Aligning Perception with Language Models

Neural Information Processing Systems

A big convergence of language, multimodal perception, action, and world modeling is a key step toward artificial general intelligence. In this work, we introduce KOSMOS-1, a Multimodal Large Language Model (MLLM) that can perceive general modalities, learn in context (i.e., few-shot), and follow instructions (i.e., zero-shot). Specifically, we train KOSMOS-1 from scratch on web-scale multi-modal corpora, including arbitrarily interleaved text and images, image-caption pairs, and text data. We evaluate various settings, including zero-shot, few-shot, and multimodal chain-of-thought prompting, on a wide range of tasks without any gradient updates or finetuning. Experimental results show that KOSMOS-1 achieves impressive performance on (i) language understanding, generation, and even OCR-free NLP (directly fed with document images), (ii) perception-language tasks, including multimodal dialogue, image captioning, visual question answering, and (iii) vision tasks, such as image recognition with descriptions (specifying classification via text instructions). We also show that MLLMs can benefit from cross-modal transfer, i.e., transfer knowledge from language to multimodal, and from multimodal to language.