Goto

Collaborating Authors

 Large Language Model


See or Recall: A Sanity Check for the Role of Vision in Solving Visualization Question Answer Tasks with Multimodal LLMs

arXiv.org Artificial Intelligence

Recent developments in multimodal large language models (MLLM) have equipped language models to reason about vision and language jointly. This permits MLLMs to both perceive and answer questions about data visualization across a variety of designs and tasks. Applying MLLMs to a broad range of visualization tasks requires us to properly evaluate their capabilities, and the most common way to conduct evaluation is through measuring a model's visualization reasoning capability, analogous to how we would evaluate human understanding of visualizations (e.g., visualization literacy). However, we found that in the context of visualization question answering (VisQA), how an MLLM perceives and reasons about visualizations can be fundamentally different from how humans approach the same problem. During the evaluation, even without visualization, the model could correctly answer a substantial portion of the visualization test questions, regardless of whether any selection options were provided. We hypothesize that the vast amount of knowledge encoded in the language model permits factual recall that supersedes the need to seek information from the visual signal. It raises concerns that the current VisQA evaluation may not fully capture the models' visualization reasoning capabilities. To address this, we propose a comprehensive sanity check framework that integrates a rule-based decision tree and a sanity check table to disentangle the effects of "seeing" (visual processing) and "recall" (reliance on prior knowledge). This validates VisQA datasets for evaluation, highlighting where models are truly "seeing", positively or negatively affected by the factual recall, or relying on inductive biases for question answering. Our study underscores the need for careful consideration in designing future visualization understanding studies when utilizing MLLMs.


Regional Tiny Stories: Using Small Models to Compare Language Learning and Tokenizer Performance

arXiv.org Artificial Intelligence

The 2023 TinyStories study developed an English dataset that allows Small Language Models (SLMs) with 1-10 million parameters to produce coherent outputs matching those of LLMs. Our research expands this framework by creating translated as well as synthetically generated datasets in Indian languages. Using this new dataset, we demonstrate that SLMs efficiently process regional languages with significantly fewer parameters than LLMs, and additionally offer a complementary framework for "inference-based evaluation" of tokenization strategies and linguistic complexity. Our analysis reveals that language-specific tokenizers outperform general-purpose ones for Indian languages. Empirical validations, supported by information-theoretic and morphological analyses, provide insights into the superior performance of Hindi models over Marathi and Bengali. The study uncovers distinct cross-linguistic patterns: Bengali emphasizes creativity, Hindi excels in context understanding and grammar with model scaling, and Marathi requires larger models to capture its unique linguistic features. Optimal parameter allocation varies, with Hindi benefiting more from wider architectures and Bengali favoring a balanced approach. We also show that quality synthetic datasets outperform translated content for training SLMs by 15-30 % . These findings advance both the practical application of SLMs to underserved languages and our theoretical understanding of neural language development.


VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation

arXiv.org Artificial Intelligence

Speech large language models (LLMs) have emerged as a prominent research focus in speech processing. We introduce VocalNet-1B and VocalNet-8B, a series of high-performance, low-latency speech LLMs enabled by a scalable and model-agnostic training framework designed for real-time voice interaction. Central to our contribution is the first application of multi-token prediction (MTP) to speech LLMs. This approach represents a paradigm shift from standard next-token prediction (NTP), offering simultaneous improvements in generation speed and quality. Informed by analysis of MTP's effect on speech generation and experimental comparisons, we designed a straightforward and highly effective MTP implementation. Experiments demonstrate that VocalNet performs on par with mainstream Omni LLMs even with limited training data, and significantly surpasses existing open-source speech LLMs. To foster reproducibility and community advancement, all model weights, inference code, training data, and framework implementations have been made publicly available at https://github.com/SJTU-OmniAgent/VocalNet


Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training

arXiv.org Artificial Intelligence

Large language models (LLMs) exhibit remarkable multilingual capabilities despite the extreme language imbalance in the pre-training data. In this paper, we closely examine the reasons behind this phenomenon, focusing on the pre-training corpus. We find that the existence of code-switching, alternating between different languages within a context, is key to multilingual capabilities. We conduct an analysis to investigate code-switching in the pre-training corpus, examining its presence and categorizing it into four types within two quadrants. We then assess its impact on multilingual performance. These types of code-switching data are unbalanced in proportions and demonstrate different effects on facilitating language transfer. To better explore the power of code-switching for language alignment during pre-training, we investigate the strategy of synthetic code-switching. We continuously scale up the synthetic code-switching data and observe remarkable improvements in both benchmarks and representation space. Extensive experiments indicate that incorporating synthetic code-switching data enables better language alignment and generalizes well to high, medium, and low-resource languages with pre-training corpora of varying qualities.


OpenAI says it would buy Chrome if Google is forced to sell

Engadget

Google is under the microscope following a court ruling last year that it has a monopoly over online search, but the future of its vast suite of digital services is still uncertain at this stage. Last month, the Justice Department suggested that Google would need to sell off the Chrome browser; if the tech giant does make that move, there's already at least one interested buyer. Bloomberg reports that Nick Turley, head of ChatGPT, spoke at a hearing today about the Google monopoly situation and was asked whether OpenAI would be interested in acquiring Chrome. "Yes, we would, as would many other parties," he said. Users can currently use the ChatGPT AI assistant in Chrome through a plugin, but Turley said there could be deeper integrations if OpenAI owned the browser.


The Great AI Lock-In Has Begun

The Atlantic - Technology

There are really two OpenAIs. One is the creator of world-bending machines--the start-up that unleashed ChatGPT and in turn the generative-AI boom, surging toward an unrecognizable future with the rest of the tech industry in tow. This is the OpenAI that promises to eventually bring about "superintelligent" programs that exceed humanity's capabilities. The other OpenAI is simply a business. This is the company that is reportedly working on a social network and considering an expansion into hardware; it is the company that offers user-experience updates to ChatGPT, such as an "image library" feature announced last week and the new ability to "reference" past chats to provide personalized responses.


OpenAI's newest AI models hallucinate way more, for reasons unknown

PCWorld

Last week, OpenAI released its new o3 and o4-mini reasoning models, which perform significantly better than their o1 and o3-mini predecessors and have new capabilities like "thinking with images" and agentically combining AI tools for more complex results. This is unusual as newer models tend to hallucinate less as the underlying AI tech improves. In the realm of LLMs and reasoning AIs, a "hallucination" occurs when the model makes up information that sounds convincing but has no bearing in truth. In other words, when you ask questions to ChatGPT, it may respond with an answer that's patently false or incorrect. OpenAI's in-house benchmark PersonQA--which is used to measure the factual accuracy of its AI models when talking about people--found that o3 hallucinated in 33 percent of responses while o4-mini did even worse at 48 percent.


Can We Build AI That Does Not Harm Queer People?

Communications of the ACM

AI safety is a contentious topic. While some prominent figures of the AI community have argued that destructive general artificial intelligence (AI) is on the horizon, others derided their warning as a marketing stunt to sell large language models (LLMs). "If the call for'AI safety' is couched in terms of protecting humanity from rogue AIs, it very conveniently displaces accountability away from the corporations scaling harm in the name of profits," tweeted Emily Bender, a professor of computational linguistics at the University of Washington. Focusing on potential future harm from ever more powerful AI systems distracts from harm that is already happening today. Most of us do not set out to make software that is actively harmful.


'What I Think about When I Type about Talking': Reflections on Text-Entry Acceleration Interfaces

Communications of the ACM

Today's text-entry tools offer a plethora of interface technologies to support users in a variety of situations and with a range of different input methods and devices.16 Recent hardware developments have enabled remarkable innovations, such as virtual keyboards that allow users to type in thin air, or to use their body as a surface for text entry. Similarly, advances in machine learning and natural language processing have enabled high-quality text generation for various purposes, such as summarizing, expanding, and co-authoring. As these technologies rapidly develop, there has been a rush to incorporate them into existing systems, often with little thought for the interactivity problems this may cause. The use of large language models (LLMs) to speed up text generation and improve prediction or completion models is becoming increasingly commonplace, with enormous theoretical efficiency savings;29 however, the implementation of these LLMs into text-entry interfaces is crucial to realizing their potential.


The Washington Post partners with OpenAI to bring its content to ChatGPT

Engadget

The Washington Post is partnering with OpenAI to bring its reporting to ChatGPT. The two organizations did not disclose the financial terms of the agreement, but the deal will see ChatGPT display summaries, quotes and links to articles from The Post when users prompt the chatbot to search the web. "We're all in on meeting our audiences where they are," said Peter Elkins-Williams, head of global partnerships at The Post. "Ensuring ChatGPT users have our impactful reporting at their fingertips builds on our commitment to provide access where, how and when our audiences want it." The Post is no stranger to generative AI. In November, the publisher began using the technology to offer article summaries.