readiness
Russia kills eight people in Ukraine, attacks two vessels in Black Sea
Is the war entering a new phase? Russian forces have launched widespread missile, drone, and artillery attacks across Ukraine, causing civilian casualties while hitting regional infrastructure and commercial shipping in the Black Sea. Ukraine's Air Force on Saturday reported facing a huge overnight barrage involving 174 drones from across Russia and Ukraine's annexed Crimean Peninsula, alongside two Zircon antiship missiles launched from Russia's Kursk region. "Enemy [Zircon antiship] missiles did not reach the targets," it said. Meanwhile, Russia's Ministry of Defence claimed its forces used Geran-4 Seeker loitering drones to strike two commercial vessels in the Black Sea - a Ukrainian dry cargo ship en route to Odesa and a docked tanker in the port of Odesa - asserting that both ships were carrying military supplies. In the Kharkiv region, Governor Oleh Syniehubov confirmed Russian attacks spanning 14 settlements killed five people and injured 15 others, causing heavy destruction in towns like Balakliya and Lypkuvativka.
Businesses finally seeing AI ROI, but 62% can't handle the storage demands
Businesses finally seeing AI ROI, but 62% can't handle the storage demands A Seagate study finds that 99% of IT leaders expect AI to drive increased data storage needs, but only 38% are prepared to meet them, revealing a significant readiness gap. Kayla Solino is a ZDNET Editor based in New York City and New Jersey. A new study finds 62% of organizations are ill-prepared to tackle surging storage needs. Organizations must lean in to AI's growing data demands and infrastructure readiness. "Sustainable scaling" may be the key to optimizing AI growth, according to Seagate's study.
SoftBank buys data center investment firm DigitalBridge
SoftBank Group aims to capitalize on soaring demand for the computing capacity that underpins artificial intelligence applications. SoftBank Group agreed to buy private equity firm DigitalBridge Group for about $3 billion in cash, part of the Japanese conglomerate's push to invest in data centers and other digital infrastructure fueling the artificial intelligence boom. SoftBank will pay $16 per share for New York-listed DigitalBridge, the companies said in statement Monday, confirming an earlier Bloomberg News report. The offer -- valued at $4 billion, including debt -- is a 65% premium to DigitalBridge's closing share price on Dec. 4, the last trading day before talks between the two companies were reported. SoftBank's billionaire founder Masayoshi Son aims to capitalize on soaring demand for digital infrastructure, driven by the AI boom.
Beyond Prototyping: Autonomous, Enterprise-Grade Frontend Development from Pixel to Production via a Specialized Multi-Agent Framework
Ganesaraja, Ramprasath, N, Swathika, AP, Saravanan, Rathinasamy, Kamalkumar, Amancharla, Chetana, Das, Rahul, Panse, Sahil Dilip, Batwe, Aditya, Vijayan, Dileep, Ashok, Veena, P, Thanushree A, Rao, Kausthubh J, Olivero, Alden, Roshan, null, Manthena, Rajeshwar Reddy, A, Asmitha Yuga Sre, Tripathi, Harsh, Selvaraj, Suganya, Chin, Vito, Bhaskar, Kasthuri Rangan, Bhaskar, Kasthuri Rangan, R, Venkatraman, Vijayakumar, Sajit
We present AI4UI, a framework of autonomous front-end development agents purpose-built to meet the rigorous requirements of enterprise-grade application delivery. Unlike general-purpose code assistants designed for rapid prototyping, AI4UI focuses on production readiness delivering secure, scalable, compliant, and maintainable UI code integrated seamlessly into enterprise workflows. AI4UI operates with targeted human-in-the-loop involvement: at the design stage, developers embed a Gen-AI-friendly grammar into Figma prototypes to encode requirements for precise interpretation; and at the post processing stage, domain experts refine outputs for nuanced design adjustments, domain-specific optimizations, and compliance needs. Between these stages, AI4UI runs fully autonomously, converting designs into engineering-ready UI code. Technical contributions include a Figma grammar for autonomous interpretation, domain-aware knowledge graphs, a secure abstract/package code integration strategy, expertise driven architecture templates, and a change-oriented workflow coordinated by specialized agent roles. In large-scale benchmarks against industry baselines and leading competitor systems, AI4UI achieved 97.24% platform compatibility, 87.10% compilation success, 86.98% security compliance, 78.00% feature implementation success, 73.50% code-review quality, and 73.36% UI/UX consistency. In blind preference studies with 200 expert evaluators, AI4UI emerged as one of the leaders demonstrating strong competitive standing among leading solutions. Operating asynchronously, AI4UI generates thousands of validated UI screens in weeks rather than months, compressing delivery timeline
Lost in the Pipeline: How Well Do Large Language Models Handle Data Preparation?
Spreafico, Matteo, Tassini, Ludovica, Sancricca, Camilla, Cappiello, Cinzia
Large language models have recently demonstrated their exceptional capabilities in supporting and automating various tasks. Among the tasks worth exploring for testing large language model capabilities, we considered data preparation, a critical yet often labor-intensive step in data-driven processes. This paper investigates whether large language models can effectively support users in selecting and automating data preparation tasks. To this aim, we considered both general-purpose and fine-tuned tabular large language models. We prompted these models with poor-quality datasets and measured their ability to perform tasks such as data profiling and cleaning. We also compare the support provided by large language models with that offered by traditional data preparation tools. To evaluate the capabilities of large language models, we developed a custom-designed quality model that has been validated through a user study to gain insights into practitioners' expectations.
UK lacks plan to defend itself from invasion, MPs warn
The UK lacks a plan to defend itself from military attack, a committee of MPs has warned. In a highly critical report, the defence committee says the UK is over-reliant on US resources and that preparations to defend itself and overseas territories in the event of attack are nowhere near where they need to be. The committee's chair, Labour MP Tan Dhesi, said: Putin's brutal invasion of Ukraine, unrelenting disinformation campaigns, and repeated incursions into European airspace mean that we cannot afford to bury our heads in the sand. It comes as the Ministry of Defence (MoD) identified parts of the country where six or more new munitions factories could be built. In June, Defence Secretary John Healey announced plans to move the UK to war-fighting readiness, including £1.5bn to support the construction of new munitions factories, which will be built by private contractors.
On the Soundness and Consistency of LLM Agents for Executing Test Cases Written in Natural Language
Salva, Sébastien, Taguelmimt, Redha
The use of natural language (NL) test cases for validating graphical user interface (GUI) applications is emerging as a promising direction to manually written executable test scripts, which are costly to develop and difficult to maintain. Recent advances in large language models (LLMs) have opened the possibility of the direct execution of NL test cases by LLM agents. This paper investigates this direction, focusing on the impact on NL test case unsoundness and on test case execution consistency. NL test cases are inherently unsound, as they may yield false failures due to ambiguous instructions or unpredictable agent behaviour. Furthermore, repeated executions of the same NL test case may lead to inconsistent outcomes, undermining test reliability. To address these challenges, we propose an algorithm for executing NL test cases with guardrail mechanisms and specialised agents that dynamically verify the correct execution of each test step. We introduce measures to evaluate the capabilities of LLMs in test execution and one measure to quantify execution consistency. We propose a definition of weak unsoundness to characterise contexts in which NL test case execution remains acceptable, with respect to the industrial quality levels Six Sigma. Our experimental evaluation with eight publicly available LLMs, ranging from 3B to 70B parameters, demonstrates both the potential and current limitations of current LLM agents for GUI testing. Our experiments show that Meta Llama 3.1 70B demonstrates acceptable capabilities in NL test case execution with high execution consistency (above the level 3-sigma). We provide prototype tools, test suites, and results.
HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
Wang, Xin, Dang, Ting, Zhang, Xinyu, Kostakos, Vassilis, Witbrock, Michael J., Jia, Hong
Mobile and wearable healthcare monitoring play a vital role in facilitating timely interventions, managing chronic health conditions, and ultimately improving individuals' quality of life. Previous studies on large language models (LLMs) have highlighted their impressive generalization abilities and effectiveness in healthcare prediction tasks. However, most LLM-based healthcare solutions are cloud-based, which raises significant privacy concerns and results in increased memory usage and latency. To address these challenges, there is growing interest in compact models, Small Language Models (SLMs), which are lightweight and designed to run locally and efficiently on mobile and wearable devices. Nevertheless, how well these models perform in healthcare prediction remains largely unexplored. We systematically evaluated SLMs on health prediction tasks using zero-shot, few-shot, and instruction fine-tuning approaches, and deployed the best performing fine-tuned SLMs on mobile devices to evaluate their real-world efficiency and predictive performance in practical healthcare scenarios. Our results show that SLMs can achieve performance comparable to LLMs while offering substantial gains in efficiency and privacy. However, challenges remain, particularly in handling class imbalance and few-shot scenarios. These findings highlight SLMs, though imperfect in their current form, as a promising solution for next-generation, privacy-preserving healthcare monitoring.