Goto

Collaborating Authors

 South America


Elon Musk wants to block out the SUN to curb global warming - but scientists warn the controversial technique could be disastrous

Daily Mail - Science & tech

Republicans reveal plot to stop'insurrectionist' democratic socialist Zohran Mamdani being sworn in as NYC mayor using Civil War-era clause Warren Buffett's $6billion stock exit is his loudest warning yet Texas governor warns any New Yorkers trying to flee south after Mamdani's win will be slapped with 100% tariff I won't ever forget what I saw at Andy Cohen's party. He may admit he's hooking up with guys on every dating app but this is the truth about men like him: KENNEDY As melatonin's terrifying link to fatal heart condition is revealed, experts weigh in on sleep aid's safety More young people developing'old person disease' of the gut with increased risk of severe complications Taylor Swift enjoys girls' night with squad member Gigi Hadid in NYC after Travis Kelce's ex took swipe Wicked star Jonathan Bailey becomes first ever openly gay man to be named People's Sexiest Man Alive Karoline Leavitt, 28, is accused of'airbrushing' husband, 60, in glamor White House snaps George W. ...


EPARA: Parallelizing Categorized AI Inference in Edge Clouds

arXiv.org Artificial Intelligence

With the increasing adoption of AI applications such as large language models and computer vision AI, the computational demands on AI inference systems are continuously rising, making the enhancement of task processing capacity using existing hardware a primary objective in edge clouds. We propose EPARA, an end-to-end AI parallel inference framework in edge, aimed at enhancing the edge AI serving capability. Our key idea is to categorize tasks based on their sensitivity to latency/frequency and requirement for GPU resources, thereby achieving both request-level and service-level task-resource allocation. EPARA consists of three core components: 1) a task-categorized parallelism allocator that decides the parallel mode of each task, 2) a distributed request handler that performs the calculation for the specific request, and 3) a state-aware scheduler that periodically updates service placement in edge clouds. We implement a EPARA prototype and conduct a case study on the EPARA operation for LLMs and segmentation tasks. Evaluation through testbed experiments involving edge servers, embedded devices, and microcomputers shows that EPARA achieves up to 2.1$\times$ higher goodput in production workloads compared to prior frameworks, while adapting to various edge AI inference tasks.


Amortized Active Generation of Pareto Sets

arXiv.org Machine Learning

We introduce active generation of Pareto sets (A-GPS), a new framework for online discrete black-box multi-objective optimization (MOO). A-GPS learns a generative model of the Pareto set that supports a-posteriori conditioning on user preferences. The method employs a class probability estimator (CPE) to predict non-dominance relations and to condition the generative model toward high-performing regions of the search space. We also show that this non-dominance CPE implicitly estimates the probability of hypervolume improvement (PHVI). To incorporate subjective trade-offs, A-GPS introduces preference direction vectors that encode user-specified preferences in objective space. At each iteration, the model is updated using both Pareto membership and alignment with these preference directions, producing an amortized generative model capable of sampling across the Pareto front without retraining. The result is a simple yet powerful approach that achieves high-quality Pareto set approximations, avoids explicit hypervolume computation, and flexibly captures user preferences. Empirical results on synthetic benchmarks and protein design tasks demonstrate strong sample efficiency and effective preference incorporation.


Quadratic Direct Forecast for Training Multi-Step Time-Series Forecast Models

arXiv.org Machine Learning

The design of training objective is central to training time-series forecasting models. Existing training objectives such as mean squared error mostly treat each future step as an independent, equally weighted task, which we found leading to the following two issues: (1) overlook the label autocorrelation effect among future steps, leading to biased training objective; (2) fail to set heterogeneous task weights for different forecasting tasks corresponding to varying future steps, limiting the forecasting performance. To fill this gap, we propose a novel quadratic-form weighted training objective, addressing both of the issues simultaneously. Specifically, the off-diagonal elements of the weighting matrix account for the label autocorrelation effect, whereas the non-uniform diagonals are expected to match the most preferable weights of the forecasting tasks with varying future steps. To achieve this, we propose a Quadratic Direct Forecast (QDF) learning algorithm, which trains the forecast model using the adaptively updated quadratic-form weighting matrix. Experiments show that our QDF effectively improves performance of various forecast models, achieving state-of-the-art results. Code is available at https://anonymous.4open.science/r/QDF-8937.


MH-1M: A 1.34 Million-Sample Comprehensive Multi-Feature Android Malware Dataset for Machine Learning, Deep Learning, Large Language Models, and Threat Intelligence Research

arXiv.org Artificial Intelligence

Abstract--We present MH-1M, one of the most comprehensive and up-to-date datasets for advanced Android malware research. The dataset comprises 1,340,515 applications, encompassing a wide range of features and extensive metadata. T o ensure accurate malware classification, we employ the VirusT otal API, integrating multiple detection engines for comprehensive and reliable assessment. Our GitHub, Figshare, and Harvard Dataverse repositories provide open access to the processed dataset and its extensive supplementary metadata, totaling more than 400 GB of data and including the outputs of the feature extraction pipeline as well as the corresponding VirusT otal reports. Our findings underscore the MH-1M dataset's invaluable role in understanding the evolving landscape of malware. The pervasive spread of Android malware poses a significant challenge for cybersecurity research. This challenge stems mainly from the open-source nature and affordability of Android platforms, which grant users access to a large market of free applications. At the same time, malware continually evolves, adapting its tactics to execute more sophisticated and frequent attacks. Such attacks often result in data destruction, information theft, and several other cybercrimes [1], [2], [3]. Machine learning (ML) algorithms have been widely used to uncover malware and have demonstrated remarkable effectiveness in detection systems, leveraging their discriminative capabilities to identify new variants of malicious applications [4], [5], [6]. To mitigate these risks, researchers have developed a variety of methods for detecting Android malware, establishing machine learning as a central focus of contemporary mobile security research [7], [8], [9]. However, the effectiveness of ML models is highly dependent on the quality of the datasets used for training. Many existing datasets suffer from limitations such as outdated data, inadequate representation, and a limited number of samples and features, making them unsuitable for modern malware detection [10], [2], [11], [12]. These issues raise concerns about the reliability of reported performance metrics and can potentially lead to misleading conclusions [2]. A growing body of research in Android malware detection strongly supports the notion that increasing the number of discriminative features can significantly improve classification performance [13], [14], [15]. We present in Table I an overview of widely used Android malware datasets from recent years.


Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories

arXiv.org Artificial Intelligence

The increasing deployment of Large Language Model (LLM) agents for complex software engineering tasks has created a need to understand their problem-solving behaviours beyond simple success metrics. While these agents demonstrate impressive capabilities in automated issue resolution, their decision-making processes remain largely opaque. This paper presents an empirical study of agent trajectories, namely the execution traces capturing the steps agents take when attempting to resolve software issues. We analyse trajectories from three state-of-the-art code agents (OpenHands, SWE-agent, and Prometheus) on the SWE-Bench benchmark, examining both successful and failed attempts. Our investigation reveals several key insights into agent behaviour. First, we identify how distinct problem-solving strategies, such as defensive programming and context gathering, enable success in different scenarios. Second, we find that failed trajectories are consistently longer and exhibit higher variance than successful ones, with failure patterns differing significantly between agents. Third, our fault localisation analysis shows that while most trajectories correctly identify problematic files (72-81\% even in failures), success depends more on achieving approximate rather than exact code modifications. These and other findings unveiled by our study, provide a foundation for understanding agent behaviour through trajectory analysis, contributing to the development of more robust and interpretable autonomous software engineering systems.


MalDataGen: A Modular Framework for Synthetic Tabular Data Generation in Malware Detection

arXiv.org Artificial Intelligence

High-quality data scarcity hinders malware detection, limiting ML performance. We introduce MalDataGen, an open-source modular framework for generating high-fidelity synthetic tabular data using modular deep learning models (e.g., WGAN-GP, VQ-V AE). Evaluated via dual validation (TR-TS/TS-TR), seven classifiers, and utility metrics, MalDataGen outperforms benchmarks like SDV while preserving data utility. Its flexible design enables seamless integration into detection pipelines, offering a practical solution for cybersecurity applications. I. Introduction Modern machine learning algorithms, particularly deep learning architectures, depend on large-scale datasets with reliable annotations to achieve optimal performance.


Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks

arXiv.org Artificial Intelligence

The rapid proliferation of Large Language Models (LLMs) has raised significant concerns about their security against adversarial attacks. In this work, we propose a novel approach to crafting universal jailbreaks and data extraction attacks by exploiting latent space discontinuities, an architectural vulnerability related to the sparsity of training data. Initial results indicate that when these discontinuities are exploited, they can consistently and profoundly compromise model behavior, even in the presence of layered defenses. The findings suggest that this strategy has substantial potential as a systemic attack vector. Disclaimer: This paper contains examples of harmful and offensive language. Additional supporting materials may be provided upon formal request and are subject to the signing of a liability and ethical use agreement. Large Language Models (LLMs) are enabling novel applications of Artificial Intelligence (AI) and transforming human activities through conversational models (e.g., ChatGPT, DeepSeek, Gemini, Llama, and Claude). LLMs allow for natural human-AI interaction and specialized applications across multiple domains, including image generation (e.g., Adobe Firefly and Pixlr), code automation (e.g., GitHub Copilot and Amazon CodeWhisperer), and retrieval-augmented generation systems (e.g., Perplexity AI and IBM watsonx). The interactions may happen using different interfaces, such as via direct interaction with the user using a Web interface or indirectly via APIs.


Mysterious drones spotted over military base storing US nuclear weapons

Daily Mail - Science & tech

China's president Xi caught knifing Trump in brutal attack just hours after historic summit World's'most trusted' broadcaster the BBC doctored Trump speech a week before the election, whistleblower reveals I won't ever forget what I saw at Andy Cohen's party. He may admit he's hooking up with guys on every dating app but this is the truth about men like him: KENNEDY'Venomous' Republican split over Israel hits new low as fiery feud reaches White House America's most dangerous cities revealed: Crime, natural disaster risks and financial safety top the list of growing concerns Drivers mock new design for world's best-selling car: 'Did it already get into a wreck?' I learned the horrifying risks of'miracle' ADHD drugs and stopped taking them... but it was too late Roller coaster camera caught utter terror on people's faces after seat belt failed on 208ft ride that travels at 75mph The leafy suburb under an hour from Manhattan where wealthy New Yorkers are fleeing to escape'woke' Mamdani's socialist dystopia The five cities with America's most pleasant climate revealed - and they're all in the same state A girl, 15, bludgeoned to death in a gated enclave, a Kennedy cousin released and the brother who'knows the truth' about the death that haunts Camelot Sex aids and poppers... the sordid discoveries made by royal aides after party Andrew threw for Epstein and Ghislaine Maxwell - and the truth about those massages: ROBERT JOBSON READ MORE: New Jersey UFO mystery solved! Mysterious drones were spotted near Belgium's Kleine Brogel air base, where US nuclear weapons are stored, prompting fears of a potential espionage operation. Belgium's Defense Minister Theo Francken confirmed that drones entered the base's airspace in two waves on Saturday and Sunday night.


OpenAI, Amazon sign 38bn AI deal

Al Jazeera

OpenAI has signed a new deal valued at $38bn with Amazon that will allow the artificial intelligence giant to run AI workloads across Amazon Web Services (AWS) cloud infrastructure. The seven-year deal announced on Monday is the first big AI push for the e-commerce giant after a restructuring last week. Experts say this does not mean that it will allow OpenAI to train its model on websites hosted by AWS - which includes the websites of The New York Times, Reddit and United Airlines. "Running OpenAI training inside AWS doesn't change their ability to scrape content from AWS-hosted websites [which they could already do for anything publicly readable]. This is strictly speaking about the economics of rent vs buy for GPU [graphics processing unit] capacity," Joshua McKenty, CEO of the AI detection company PolyguardAI, told Al Jazeera. The deal is also a major vote of confidence for the e-commerce giant's cloud unit, AWS, which some investors feared had fallen behind rivals Microsoft and Google in the artificial intelligence (AI) race.