Goto

Collaborating Authors

 Energy


Deficient Excitation in Parameter Learning

arXiv.org Artificial Intelligence

This paper investigates parameter learning problems under deficient excitation (DE). The DE condition is a rank-deficient, and therefore, a more general evolution of the well-known persistent excitation condition. Under the DE condition, a proposed online algorithm is able to calculate the identifiable and non-identifiable subspaces, and finally give an optimal parameter estimate in the sense of least squares. In particular, the learning error within the identifiable subspace exponentially converges to zero in the noise-free case, even without persistent excitation. The DE condition also provides a new perspective for solving distributed parameter learning problems, where the challenge is posed by local regressors that are often insufficiently excited. To improve knowledge of the unknown parameters, a cooperative learning protocol is proposed for a group of estimators that collect measured information under complementary DE conditions. This protocol allows each local estimator to operate locally in its identifiable subspace, and reach a consensus with neighbours in its non-identifiable subspace. As a result, the task of estimating unknown parameters can be achieved in a distributed way using cooperative local estimators. Application examples in system identification are given to demonstrate the effectiveness of the theoretical results developed in this paper.


AI persuading AI vs AI persuading Humans: LLMs' Differential Effectiveness in Promoting Pro-Environmental Behavior

arXiv.org Artificial Intelligence

Pro-environmental behavior (PEB) is vital to combat climate change, yet turning awareness into intention and action remains elusive. We explore large language models (LLMs) as tools to promote PEB, comparing their impact across 3,200 participants: real humans (n=1,200), simulated humans based on actual participant data (n=1,200), and fully synthetic personas (n=1,200). All three participant groups faced personalized or standard chatbots, or static statements, employing four persuasion strategies (moral foundations, future self-continuity, action orientation, or "freestyle" chosen by the LLM). Results reveal a "synthetic persuasion paradox": synthetic and simulated agents significantly affect their post-intervention PEB stance, while human responses barely shift. Simulated participants better approximate human trends but still overestimate effects. This disconnect underscores LLM's potential for pre-evaluating PEB interventions but warns of its limits in predicting real-world behavior. We call for refined synthetic modeling and sustained and extended human trials to align conversational AI's promise with tangible sustainability outcomes.


Building Machine Learning Challenges for Anomaly Detection in Science

arXiv.org Artificial Intelligence

Scientific discoveries are often made by finding a pattern or object that was not predicted by the known rules of science. Oftentimes, these anomalous events or objects that do not conform to the norms are an indication that the rules of science governing the data are incomplete, and something new needs to be present to explain these unexpected outliers. The challenge of finding anomalies can be confounding since it requires codifying a complete knowledge of the known scientific behaviors and then projecting these known behaviors on the data to look for deviations. When utilizing machine learning, this presents a particular challenge since we require that the model not only understands scientific data perfectly but also recognizes when the data is inconsistent and out of the scope of its trained behavior. In this paper, we present three datasets aimed at developing machine learning-based anomaly detection for disparate scientific domains covering astrophysics, genomics, and polar science. We present the different datasets along with a scheme to make machine learning challenges around the three datasets findable, accessible, interoperable, and reusable (FAIR). Furthermore, we present an approach that generalizes to future machine learning challenges, enabling the possibility of large, more compute-intensive challenges that can ultimately lead to scientific discovery.


Holistically Evaluating the Environmental Impact of Creating Language Models

arXiv.org Artificial Intelligence

As the performance of artificial intelligence systems has dramatically increased, so too has the environmental impact of creating these systems. While many model developers release estimates of the power consumption and carbon emissions from the final training runs for their latest models, there is comparatively little transparency into the impact of model development, hardware manufacturing, and total water usage throughout. In this work, we estimate the real-world environmental impact of developing a series of language models, ranging from 20 million to 13 billion active parameters, trained on up to 5.6 trillion tokens each. When accounting for hardware manufacturing, model development, and our final training runs, we find that our series of models released 493 metric tons of carbon emissions, equivalent to powering about 98 homes in the United States for one year, and consumed 2.769 million liters of water, equivalent to about 24.5 years of water usage by a person in the United States, even though our data center is extremely water-efficient. We measure and report the environmental impact of our model development; to the best of our knowledge we are the first to do so for LLMs, and we find that model development, the impact of which is generally not disclosed by most model developers, amounted to 50% of that of training. By looking at detailed time series data for power consumption, we also find that power usage throughout training is not consistent, fluctuating between 15% and 85% of our hardware's maximum power draw, with negative implications for grid-scale planning as demand continues to grow. We close with a discussion on the continued difficulty of estimating the environmental impact of AI systems, and key takeaways for model developers and the public at large. In recent years, the field of artificial intelligence has progressed at an unprecedented pace, driven in large part by the development and deployment of large language and multimodal models.


Fault Localization and State Estimation of Power Grid under Parallel Cyber-Physical Attacks

arXiv.org Artificial Intelligence

--Parallel cyber-physical attacks (PCPA) refer to those attacks on power grids by disturbing/cutting off physical transmission lines and meanwhile blocking transmission of measurement data to dwarf or delay the system protection and recovery actions. Such fierce hostile attacks impose critical threats to the modern power grids when there is a fusion of power grids and telecommunication technologies. In this paper, we investigate the fault diagnosis problem of faulty transmission lines under a broader spectrum of PCPA for a linearized (or DC) power flow model. The physical attack mechanism of PCPA includes not only disconnection but also admittance value modification on transmission lines, for example, by invading distributed flexible AC transmission system (D-F ACTS). T o tackle the problem, we first recover the information of voltage phase angles within the attacked area. Using the information of voltage phase angle and power injection of buses, a graph attention network-based fault localization (GA T -FL) algorithm is proposed to find the locations of the physical attacks. By capitalizing on the feature extraction capability of the GA T on graph data, the fault localization algorithm outperforms the existing results when under cyber attacks, e.g., denial of service (DoS) attacks. A line state identification algorithm is then developed to identify the states of the transmission lines within the attacked area. Specifically, the algorithm restores the power injection of buses within the attacked area and then identities the state of all the transmission lines within the attacked area by solving a linear programming (LP) problem. Experimental simulations are conducted on IEEE 30/118 bus standard test cases to demonstrate the effectiveness of the proposed fault diagnosis algorithms. N recent years, smart grids [1] have experienced rapid developments, driven by the the needs for more effective control of power systems and more efficient utilization of renewable and non-renewable energy resources. Junhao Ren, Kai Zhao and Gaoxi Xiao are with the School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798.


Integrated Computation and Communication with Fiber-optic Transmissions

arXiv.org Artificial Intelligence

Abstract: Fiber - optic transmission systems are leveraged not only as high - speed communication channels but also as nonlinear kernel functions for machine learning computations, enabling the seamless integration of computational intelligence and communication . Over the past few decades, the field of communication has undergone remarkable transformations, driven by advancements in network architecture and transmission technologies. Simultaneously, the emergence of machine learning (ML) has revolutionized various industries, p aving the way for intelligent communication networks [1] . As outlined in the IMT - 2030 framework [2], the future of communication systems lies in the seamless integration of ML with communication technologies, a devel opment that is expected to redefine the capabilities of these networks. This integration is particularly critical for applications like semantic communication, which demand a unified approach to bridging the physical and cyber domains.


Aerial Infrared Health Monitoring of Solar Photovoltaic Farms at Scale

arXiv.org Artificial Intelligence

Solar photovoltaic (PV) farms represent a major source of global renewable energy generation, yet their true operational efficiency often remains unknown at scale. In this paper, we present a comprehensive, data-driven framework for large-scale airborne infrared inspection of North American solar installations. Leveraging high-resolution thermal imagery, we construct and curate a geographically diverse dataset encompassing thousands of PV sites, enabling machine learning-based detection and localization of defects that are not detectable in the visible spectrum. Our pipeline integrates advanced image processing, georeferencing, and airborne thermal infrared anomaly detection to provide rigorous estimates of performance losses. We highlight practical considerations in aerial data collection, annotation methodologies, and model deployment across a wide range of environmental and operational conditions. Our work delivers new insights into the reliability of large-scale solar assets and serves as a foundation for ongoing research on performance trends, predictive maintenance, and scalable analytics in the renewable energy sector.


Correlation to Causation: A Causal Deep Learning Framework for Arctic Sea Ice Prediction

arXiv.org Artificial Intelligence

Building upon the previously introduced MVGC and PCMCI+ algorithms, we applied these methods to identify key causal variables of Arctic sea ice dynamics. For both daily and monthly datasets, MVGC identified all variables except Sea Surface T emperature (SST) as causal features. This result underscores the broad influence of atmospheric and oceanic variables on Arctic sea ice. PCMCI+, known for its robustness in handling high-dimensional and autocorrelated time series data, provided a more refined identification of causal features. For the daily dataset, PCMCI+ highlighted longwave radiation, snowfall, sea surface salinity (SSS), surface pressure, and SIE itself as the primary causal factors. For the monthly dataset, the identified causal features were longwave radiation, SST, and SIE . These results suggest temporal and spatial differences in the causal relationships influencing SIE dynamics across daily and monthly timescales. Figure 4 shows the causal graphs generated by PCMCI+ for daily and monthly datasets, highlighting the direct causal influences of key variables on Arctic SIE. The identified features guided the selection of input variables for the GRU-LSTM model, ensuring that the model leveraged causally significant information for prediction.


CrowdSelect: Synthetic Instruction Data Selection with Multi-LLM Wisdom

arXiv.org Artificial Intelligence

Distilling advanced Large Language Models' instruction-following capabilities into smaller models using a selected subset has become a mainstream approach in model training. While existing synthetic instruction data selection strategies rely mainly on single-dimensional signals (i.e., reward scores, model perplexity), they fail to capture the complexity of instruction-following across diverse fields. Therefore, we investigate more diverse signals to capture comprehensive instruction-response pair characteristics and propose three foundational metrics that leverage Multi-LLM wisdom, informed by (1) diverse LLM responses and (2) reward model assessment. Building upon base metrics, we propose CrowdSelect, an integrated metric incorporating a clustering-based approach to maintain response diversity. Our comprehensive experiments demonstrate that our foundation metrics consistently improve performance across 4 base models on MT-bench and Arena-Hard. CrowdSelect, efficiently incorporating all metrics, achieves state-of-the-art performance in both Full and LoRA fine-tuning, showing improvements of 4.81% on Arena-Hard and 11.1% on MT-bench with Llama-3.2-3b-instruct. We hope our findings will bring valuable insights for future research in this direction. Code are available at https://github.com/listentm/crowdselect.


Deep Reinforcement Learning-Based User Association in Hybrid LiFi/WiFi Indoor Networks

arXiv.org Artificial Intelligence

--Hybrid light fidelity (LiFi) and wireless fidelity (WiFi) indoor networks has been envisioned as a promising technology to alleviate radio frequency spectrum crunch to accommodate the ever-increasing data rate demand in indoor scenarios. The hybrid LiFi/WiFi indoor networks can leverage the advantages of fast data transmission from LiFi and wider coverage of WiFi, thus complementing well with each other and further improving the network performance compared with the standalone networks. However, to leverage the co-existence, several challenges should be addressed, including but not limited to user association, mobility support, and efficient resource allocation. Therefore, the objective of the paper is to design a new user-access point association algorithm to maximize the sum throughput of the hybrid networks. We first mathematically formulate the sum data rate maximization problem by determining the AP selection for each user in indoor networks with consideration of user mobility and practical capacity limitations, which is a nonconvex binary integer programming problem. T o solve this problem, we then propose a sequential-proximal policy optimization (S-PPO) based deep reinforcement learning method. Extensive simulations are conducted to evaluate the proposed method by comparing it with exhaustive search (ES), signal strength strategy (SSS), and trust region policy optimization (TRPO) methods. Comprehensive simulation results demonstrate that our solution algorithm can outperform SSS by about 32.25% of the sum throughput and 19.09% of the fairness on average, and outperform TRPO by about 10.34% and 10.23%, respectively. Over the past few years, the usage of the internet has been continuously increasing. According to the latest data, people spend an average of 6 hours and 58 minutes daily on screens connected to the internet [1]. Moreover, an increasing number of applications require high-speed support, such as video calls, VR gaming, streaming media, and so on. However, we are facing a global digital divide, i.e., internet speeds in urban areas are often much faster than in rural areas, due to the generally less developed internet infrastructure in rural locations. Visible light communication (VLC), where light-emitting diodes (LEDs) can be used to transmit data by optical spectrum, has been envisioned as a promising solution for last-mile access because of its high bandwidth, enhanced security, electromagnetic interference-free nature, and easy integration with existing infrastructure [2]-[7].