Energy
Wind speed super-resolution and validation: from ERA5 to CERRA via diffusion models
Merizzi, Fabio, Asperti, Andrea, Colamonaco, Stefano
The Copernicus Regional Reanalysis for Europe, CERRA, is a high-resolution regional reanalysis dataset for the European domain. In recent years it has shown significant utility across various climate-related tasks, ranging from forecasting and climate change research to renewable energy prediction, resource management, air quality risk assessment, and the forecasting of rare events, among others. Unfortunately, the availability of CERRA is lagging two years behind the current date, due to constraints in acquiring the requisite external data and the intensive computational demands inherent in its generation. As a solution, this paper introduces a novel method using diffusion models to approximate CERRA downscaling in a data-driven manner, without additional informations. By leveraging the lower resolution ERA5 dataset, which provides boundary conditions for CERRA, we approach this as a super-resolution task. Focusing on wind speed around Italy, our model, trained on existing CERRA data, shows promising results, closely mirroring original CERRA data. Validation with in-situ observations further confirms the model's accuracy in approximating ground measurements.
Efficient Large Language Models: A Survey
Wan, Zhongwei, Wang, Xin, Liu, Che, Alam, Samiul, Zheng, Yu, Liu, Jiachen, Qu, Zhongnan, Yan, Shen, Zhu, Yi, Zhang, Quanlu, Chowdhury, Mosharaf, Zhang, Mi
Large Language Models (LLMs) have demonstrated remarkable capabilities in important tasks such as natural language understanding, language generation, and complex reasoning and have the potential to make a substantial impact on our society. Such capabilities, however, come with the considerable resources they demand, highlighting the strong need to develop effective techniques for addressing their efficiency challenges.In this survey, we provide a systematic and comprehensive review of efficient LLMs research. We organize the literature in a taxonomy consisting of three main categories, covering distinct yet interconnected efficient LLMs topics from model-centric, data-centric, and framework-centric perspective, respectively. We have also created a GitHub repository where we compile the papers featured in this survey at https://github.com/AIoT-MLSys-Lab/Efficient-LLMs-Survey, and will actively maintain this repository and incorporate new research as it emerges. We hope our survey can serve as a valuable resource to help researchers and practitioners gain a systematic understanding of the research developments in efficient LLMs and inspire them to contribute to this important and exciting field.
Using Large Language Models to Generate, Validate, and Apply User Intent Taxonomies
Shah, Chirag, White, Ryen W., Andersen, Reid, Buscher, Georg, Counts, Scott, Das, Sarkar Snigdha Sarathi, Montazer, Ali, Manivannan, Sathish, Neville, Jennifer, Ni, Xiaochuan, Rangan, Nagu, Safavi, Tara, Suri, Siddharth, Wan, Mengting, Wang, Leijie, Yang, Longqi
Log data can reveal valuable information about how users interact with Web search services, what they want, and how satisfied they are. However, analyzing user intents in log data is not easy, especially for emerging forms of Web search such as AI-driven chat. To understand user intents from log data, we need a way to label them with meaningful categories that capture their diversity and dynamics. Existing methods rely on manual or machine-learned labeling, which are either expensive or inflexible for large and dynamic datasets. We propose a novel solution using large language models (LLMs), which can generate rich and relevant concepts, descriptions, and examples for user intents. However, using LLMs to generate a user intent taxonomy and apply it for log analysis can be problematic for two main reasons: (1) such a taxonomy is not externally validated; and (2) there may be an undesirable feedback loop. To address this, we propose a new methodology with human experts and assessors to verify the quality of the LLM-generated taxonomy. We also present an end-to-end pipeline that uses an LLM with human-in-the-loop to produce, refine, and apply labels for user intent analysis in log data. We demonstrate its effectiveness by uncovering new insights into user intents from search and chat logs from the Microsoft Bing commercial search engine. The proposed work's novelty stems from the method for generating purpose-driven user intent taxonomies with strong validation. This method not only helps remove methodological and practical bottlenecks from intent-focused research, but also provides a new framework for generating, validating, and applying other kinds of taxonomies in a scalable and adaptable way with minimal human effort.
Combining Deep Learning and Street View Imagery to Map Smallholder Crop Types
Soler, Jordi Laguarta, Friedel, Thomas, Wang, Sherrie
Accurate crop type maps are an essential source of information for monitoring yield progress at scale, projecting global crop production, and planning effective policies. To date, however, crop type maps remain challenging to create in low and middle-income countries due to a lack of ground truth labels for training machine learning models. Field surveys are the gold standard in terms of accuracy but require an often-prohibitively large amount of time, money, and statistical capacity. In recent years, street-level imagery, such as Google Street View, KartaView, and Mapillary, has become available around the world. Such imagery contains rich information about crop types grown at particular locations and times. In this work, we develop an automated system to generate crop type ground references using deep learning and Google Street View imagery. The method efficiently curates a set of street view images containing crop fields, trains a model to predict crop type by utilizing weakly-labelled images from disparate out-of-domain sources, and combines predicted labels with remote sensing time series to create a wall-to-wall crop type map. We show that, in Thailand, the resulting country-wide map of rice, cassava, maize, and sugarcane achieves an accuracy of 93%. We publicly release the first-ever crop type map for all of Thailand 2022 at 10m-resolution with no gaps. To our knowledge, this is the first time a 10m-resolution, multi-crop map has been created for any smallholder country. As the availability of roadside imagery expands, our pipeline provides a way to map crop types at scale around the globe, especially in underserved smallholder regions.
North Korea now using AI in nuclear program: report
A group of scientists from across the U.S. claim to have created the first artificial intelligence capable of generating AI without human supervision. North Korea has been developing artificial intelligence across various sectors, including in military technology and programs that safeguard nuclear reactors, which could create international threats, according to a new report. The authoritarian regime has used AI to develop wargame simulations and has collaborated with Chinese tech researchers, according to a report by 38 North, a publication for policy and technical analysis of North Korean affairs. The AI advancements and foreign collaboration could lead to sanction violations and leaked information, the report stated. North Korea has been rapidly developing artificial intelligence for a myriad of civilian and military uses, according to a new report.
Low-carbon milk to AI irrigation: tech startups powering Latin America's green revolution
Leo Prieto's passion for nature started during his childhood by the sea. "I was obsessed with what was under the surface. I'd anchor myself to a rock with my snorkel, and I was fascinated by all the little animals doing things that go unnoticed." His teenage years coincided with the arrival of the internet in Chile, where he became a web pioneer, launching and selling several startups. Inevitably, his interests in the environment, the internet and business merged, driven by the feeling that technological advances should not be wasted.
AIhub monthly digest: January 2024 – closed-loop robot planning, crowdsourced clustering, and trustworthiness in GPT models
We start 2024 with a packed monthly digest, where you can catch up with any AIhub stories you may have missed, peruse the latest news, recap recent events, and more. This month, we continue our coverage of NeurIPS, meet the first interviewee in our AAAI Doctoral Consortium series, and find out how to build AI openly. The AAAI/SIGAI Doctoral Consortium provides an opportunity for a group of PhD students to discuss and explore their research interests and career objectives in an interdisciplinary workshop together with a panel of established researchers. Over the course of the next few months, we'll be meeting the participants and finding out more about their work, PhD life, and their future research plans. In the first interview of the series, Changhoon Kim told us about his research on enhancing the reliability of image generative AI.
Enhancing Low-Order Discontinuous Galerkin Methods with Neural Ordinary Differential Equations for Compressible Navier--Stokes Equations
Kang, Shinhoo, Constantinescu, Emil M.
However, it is still challenging to solve practical problems such as blood flows, atmospheric and ocean currents, wildfires, and wind turbines because of their multiscale nature. Resolving all scales is computationally infeasible. As a result, physical modeling is typically carried out on a coarse grid using the appropriate subgrid-scale (SGS) models. For example, large eddy simulation (LES) resolves large-scale turbulent motion on a grid, which carries most of the flow energy, while modeling the small scales that have relatively little influence on the mean flow [8]. The goal of SGS models is to capture the effect of the small-scale structures that cannot be resolved in the grid on the resolved scales and to guarantee numerical stability [9]. The static Smagorinsky model [10] and dynamic Smagorinsky model [11], which predict the dissipation of SGS energy, are the most widely used SGS models for turbulence. In high-order DG methods, both Collis [12] and Sengupta et al. [13] successfully used the static Smagorinksy model and dynamic Smagorinsky model, respectively. However, these Smagorinsky models perform poorly for certain flows [14, 15] because they are based on the assumption that eddy viscosity is always purely dissipative and thus are unable to account for energy flow from small scales to large scales (backscatter) [16, 11]. Alternatively, numerical dissipation can be used for modeling unresolved scales as an implicit SGS model.
FDR-Controlled Portfolio Optimization for Sparse Financial Index Tracking
Machkour, Jasin, Palomar, Daniel P., Muma, Michael
In high-dimensional data analysis, such as financial index tracking or biomedical applications, it is crucial to select the few relevant variables while maintaining control over the false discovery rate (FDR). In these applications, strong dependencies often exist among the variables (e.g., stock returns), which can undermine the FDR control property of existing methods like the model-X knockoff method or the T-Rex selector. To address this issue, we have expanded the T-Rex framework to accommodate overlapping groups of highly correlated variables. This is achieved by integrating a nearest neighbors penalization mechanism into the framework, which provably controls the FDR at the user-defined target level. A real-world example of sparse index tracking demonstrates the proposed method's ability to accurately track the S&P 500 index over the past 20 years based on a small number of stocks. An open-source implementation is provided within the R package TRexSelector on CRAN.
Enhancing Efficiency and Robustness in Support Vector Regression with HawkEye Loss
Akhtar, Mushir, Tanveer, M., Arshad, Mohd.
Support vector regression (SVR) has garnered significant popularity over the past two decades owing to its wide range of applications across various fields. Despite its versatility, SVR encounters challenges when confronted with outliers and noise, primarily due to the use of the $\varepsilon$-insensitive loss function. To address this limitation, SVR with bounded loss functions has emerged as an appealing alternative, offering enhanced generalization performance and robustness. Notably, recent developments focus on designing bounded loss functions with smooth characteristics, facilitating the adoption of gradient-based optimization algorithms. However, it's crucial to highlight that these bounded and smooth loss functions do not possess an insensitive zone. In this paper, we address the aforementioned constraints by introducing a novel symmetric loss function named the HawkEye loss function. It is worth noting that the HawkEye loss function stands out as the first loss function in SVR literature to be bounded, smooth, and simultaneously possess an insensitive zone. Leveraging this breakthrough, we integrate the HawkEye loss function into the least squares framework of SVR and yield a new fast and robust model termed HE-LSSVR. The optimization problem inherent to HE-LSSVR is addressed by harnessing the adaptive moment estimation (Adam) algorithm, known for its adaptive learning rate and efficacy in handling large-scale problems. To our knowledge, this is the first time Adam has been employed to solve an SVR problem. To empirically validate the proposed HE-LSSVR model, we evaluate it on UCI, synthetic, and time series datasets. The experimental outcomes unequivocally reveal the superiority of the HE-LSSVR model both in terms of its remarkable generalization performance and its efficiency in training time.