Goto

Collaborating Authors

 Deep Learning


Application of Reduced-Order Models for Temporal Multiscale Representations in the Prediction of Dynamical Systems

arXiv.org Artificial Intelligence

Modeling and predicting the dynamics of complex multiscale systems remains a significant challenge due to their inherent nonlinearities and sensitivity to initial conditions, as well as limitations of traditional machine learning methods that fail to capture high frequency behaviours. To overcome these difficulties, we propose three approaches for multiscale learning. The first leverages the Partition of Unity (PU) method, integrated with neural networks, to decompose the dynamics into local components and directly predict both macro- and micro-scale behaviors. The second applies the Singular Value Decomposition (SVD) to extract dominant modes that explicitly separate macro- and micro-scale dynamics. Since full access to the data matrix is rarely available in practice, we further employ a Sparse High-Order SVD to reconstruct multiscale dynamics from limited measurements. Together, these approaches ensure that both coarse and fine dynamics are accurately captured, making the framework effective for real-world applications involving complex, multi-scale phenomena and adaptable to higher-dimensional systems with incomplete observations, by providing an approximation and interpretation in all time scales present in the phenomena under study.


Benchmarking On-Device Machine Learning on Apple Silicon with MLX

arXiv.org Artificial Intelligence

The recent widespread adoption of Large Language Models (LLMs) and machine learning in general has sparked research interest in exploring the possibilities of deploying these models on smaller devices such as laptops and mobile phones. This creates a need for frameworks and approaches that are capable of taking advantage of on-device hardware. The MLX framework was created to address this need. It is a framework optimized for machine learning (ML) computations on Apple silicon devices, facilitating easier research, experimentation, and prototyping. This paper presents a performance evaluation of MLX, focusing on inference latency of transformer models. We compare the performance of different transformer architecture implementations in MLX with their Pytorch counterparts. For this research we create a framework called MLX-transformers which includes different transformer implementations in MLX and downloads the model checkpoints in pytorch and converts it to the MLX format. By leveraging the advanced architecture and capabilities of Apple Silicon, MLX-Transformers enables seamless execution of transformer models directly sourced from Hugging Face, eliminating the need for checkpoint conversion often required when porting models between frameworks. Our study benchmarks different transformer models on two Apple Silicon macbook devices against an NVIDIA CUDA GPU. Specifically, we compare the inference latency performance of models with the same parameter sizes and checkpoints. We evaluate the performance of BERT, RoBERTa, and XLM-RoBERTa models, with the intention of extending future work to include models of different modalities, thus providing a more comprehensive assessment of MLX's capabilities. The results highlight MLX's potential in enabling efficient and more accessible on-device ML applications within Apple's ecosystem.


Misinformation Detection using Large Language Models with Explainability

arXiv.org Artificial Intelligence

The COVID Fake News dataset is a collection of mostly COVID-19 pandemic-specific news headlines and brief claims. The data is representative of the combination of proven factual statements and much misleading or outright false information widespread on digital platforms during the pandemic. The data set was then preprocessed and split into training (8,160 samples) and testing (2,041 samples) categories in a balanced portion so that both real and fake labels could be checked robustly. The dataset used to check whether the pipeline can be applied to other domains rather than the pandemic area is the FakeNewsNet GossipCop. This dataset lies in the domain of entertainment and celebrity news and it is one of the prominent areas where gossip, rumors, fabricated stories are prevalent. Approximately 10,000 samples were used to train, and 2,500 samples were used to test. In the present dataset, the labels distinguish the news objects as Real or Fake by fact-checking them with regards to the original GossipCop platform. The two datasets were combined, standardized, and stratified to ensure the balanced classes in the samples during training and validation. Such prudent training has the benefit of enabling these models to improve in identifying subtle signs in language that may be contained in actual and made-up claims that can be used in enhancing the pipeline to perform better in practical misinformation detection applications.


Improving Topic Modeling of Social Media Short Texts with Rephrasing: A Case Study of COVID-19 Related Tweets

arXiv.org Artificial Intelligence

Social media platforms such as Twitter (now X) provide rich data for analyzing public discourse, especially during crises such as the COVID-19 pandemic. However, the brevity, informality, and noise of social media short texts often hinder the effectiveness of traditional topic modeling, producing incoherent or redundant topics that are often difficult to interpret. To address these challenges, we have developed \emph{TM-Rephrase}, a model-agnostic framework that leverages large language models (LLMs) to rephrase raw tweets into more standardized and formal language prior to topic modeling. Using a dataset of 25,027 COVID-19-related Twitter posts, we investigate the effects of two rephrasing strategies, general- and colloquial-to-formal-rephrasing, on multiple topic modeling methods. Results demonstrate that \emph{TM-Rephrase} improves three metrics measuring topic modeling performance (i.e., topic coherence, topic uniqueness, and topic diversity) while reducing topic redundancy of most topic modeling algorithms, with the colloquial-to-formal strategy yielding the greatest performance gains and especially for the Latent Dirichlet Allocation (LDA) algorithm. This study contributes to a model-agnostic approach to enhancing topic modeling in public health related social media analysis, with broad implications for improved understanding of public discourse in health crisis as well as other important domains.


DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code

arXiv.org Artificial Intelligence

The prevalence of Large Language Models (LLMs) for generating multilingual text and source code has only increased the imperative for machine-generated content detectors to be accurate and efficient across domains. Current detectors, predominantly utilizing zero-shot methods, such as Fast DetectGPT or GPTZero, either incur high computational cost or lack sufficient accuracy, often with a trade-off between the two, leaving room for further improvement. To address these gaps, we propose the fine-tuning of encoder-only Small Language Models (SLMs), in particular, the pre-trained models of RoBERTA and CodeBERTa using specialized datasets on source code and other natural language to prove that for the task of binary classification, SLMs outperform LLMs by a huge margin whilst using a fraction of compute. Our encoders achieve AUROC $= 0.97$ to $0.99$ and macro-F1 $0.89$ to $0.94$ while reducing latency by $8$-$12\times$ and peak VRAM by $3$-$5\times$ at $512$-token inputs. Under cross-generator shifts and adversarial transformations (paraphrase, back-translation; code formatting/renaming), performance retains $\geq 92%$ of clean AUROC. We release training and evaluation scripts with seeds and configs; a reproducibility checklist is also included.


Transformer-Based Low-Resource Language Translation: A Study on Standard Bengali to Sylheti

arXiv.org Artificial Intelligence

WORK Although the findings highlight the effectiveness of fine - tuned transformer models for Bengali - Sylheti translation, several limitations remain. The dataset size (5,002 parallel sentences) restricts the models' capacity to generalize across diverse syntactic structures, stylistic variations, and domain - specific expressions. In addition, orthographic inconsistencies in Sylheti introduce noise, leading to training instability, particularly in models like mBART - 50. Another limitation is the reliance on automatic evaluation metrics such as BLEU and chrF, which may not fully capture the linguistic richness or cultural nuance of Sylheti. Future research should therefore focus on expanding the datas et through community - driven contributions and data augmentation strategies. Incorporating orthographic normalization could improve consistency and reduce variability during training. Hybrid approaches that combine the strengths of pre - trained LLMs with fin e - tuned NMT models may also enhance translation robustness in low - resource settings. Finally, incorporating human evaluation will provide a more comprehensive assessment of translation adequacy, fluency, and cultural alignment.


When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs

arXiv.org Artificial Intelligence

Despite widespread deployment of Large Language Models, systematic evaluation of instruction-following capabilities remains challenging. While comprehensive benchmarks exist, focused assessments that quickly diagnose specific instruction adherence patterns are valuable. As newer models may be trained on existing benchmarks, novel evaluation approaches are needed to assess genuine capabilities rather than memorized performance. This paper presents a streamlined evaluation framework using twenty carefully designed prompts to assess LLM instruction-following across diverse task categories. We demonstrate this framework through a large-scale empirical study conducted on October 14, 2025, testing 256 verified working models from 331 available via OpenRouter. To ensure methodological rigor and prevent selection bias, we first verified each model's basic functionality before inclusion. Unlike large-scale benchmarks requiring extensive computational resources, our approach offers a practical diagnostic tool researchers and practitioners can readily apply. Our methodology builds upon verifiable instructions while introducing a compact test suite balancing comprehensiveness with efficiency. Each prompt targets distinct aspects of instruction following, including format compliance, content constraints, logical sequencing, and multi-step task execution. We evaluate models from major providers (OpenAI, Anthropic, Google, Meta, Mistral) and emerging implementations (Qwen, DeepSeek, community models), providing comparative performance analysis. Our findings reveal consistent failure modes and identify specific instruction types posing particular challenges. This work contributes both a practical evaluation tool and one of the most comprehensive empirical analyses of instruction-following capabilities across the contemporary LLM landscape.


Contextual Augmentation for Entity Linking using Large Language Models

arXiv.org Artificial Intelligence

Entity Linking involves detecting and linking entity mentions in natural language texts to a knowledge graph. Traditional methods use a two-step process with separate models for entity recognition and disambiguation, which can be computationally intensive and less effective. We propose a fine-tuned model that jointly integrates entity recognition and disambiguation in a unified framework. Furthermore, our approach leverages large language models to enrich the context of entity mentions, yielding better performance in entity disambiguation. We evaluated our approach on benchmark datasets and compared with several baselines. The evaluation results show that our approach achieves state-of-the-art performance on out-of-domain datasets.


Towards Better Health Conversations: The Benefits of Context-seeking

arXiv.org Artificial Intelligence

Navigating health questions can be daunting in the modern information landscape. Large language models (LLMs) may provide tailored, accessible information, but also risk being inaccurate, biased or misleading. We present insights from 4 mixed-methods studies (total N=163), examining how people interact with LLMs for their own health questions. Qualitative studies revealed the importance of context-seeking in conversational AIs to elicit specific details a person may not volunteer or know to share. Context-seeking by LLMs was valued by participants, even if it meant deferring an answer for several turns. Incorporating these insights, we developed a "Wayfinding AI" to proactively solicit context. In a randomized, blinded study, participants rated the Wayfinding AI as more helpful, relevant, and tailored to their concerns compared to a baseline AI. These results demonstrate the strong impact of proactive context-seeking on conversational dynamics, and suggest design patterns for conversational AI to help navigate health topics.


LLM Bazaar: A Service Design for Supporting Collaborative Learning with an LLM-Powered Multi-Party Collaboration Infrastructure

arXiv.org Artificial Intelligence

Providing technological support for collaborative and discussion-based learning has long been a focus in CSCL research (Gweon et al., 2006; Kollar et al., 2006; Kumar et al., 2007; Rosé and Ferschke, 2016, Naik et al., 2024). Open - source architectures like Bazaar (Adamson et al., 2014) have enabled implementation of a plethora of dynamic support interventions, even for face - to -face collaboration through multi - modal sensing (Wang et al., 2020), which can be used in a portable fashion for nearly anytime-anywhere collaboration support (Vitiello et al., 2023). Past studies highlight the benefits of interactive and context-sensitive support in group learning (Kumar et al., 2007; Kumar and Rose, 2010). While static scaffolding like fixed prompts (Vogel et al., 2021) and scripted roles (Fischer et al., 2013) have been effective, contextualized interventions within specific conversational contexts (Ai et al., 2010; Cui et al., 2009) or support for student role taking (Gweon; et al., 2007) have also shown positive outcomes. Past studies incorporating dynamic support agents in collaborative learning activities (Kumar et al., 2007; Kumar and Rosé, 2010; Rosé and Ferschke, 2016) have shown the effectiveness of discussion-based learning integrated with conversational support using dialog agents. Finally Sankaranarayanan and colleagues (Sankaranarayanan et al., 2022a; Sankaranarayanan et al., 2022b) have shown the effectiveness of reflection-based learning for collaborative software development, showing that shifting students' focus more towards reflection than actual coding can increase conceptual learning without harming the ability to write code. The contribution of this design paper is the introduction of capabilities from Large Language Models (LLMs) (Vaswani, 2017) to enable new forms of collaborative support agents. While recent studies demonstrate that this new generation of support agents can be effective learning support, the new contribution of this paper is an extension to a publicly available and open-source plat form to easily integrate LLM agents developed in the broader CSCL community in order to facilitate needed research to answer questions about how best to use new AI capabilities to support collaborative learning effectively. We provide code for the LLMbazaar extension, the illustrative instructional example described below, and instructions for obtaining support for using this resource, available on GitHub (Bazaar, 2025).