caption
Asbestos killed my mum in her 40s – was her school to blame?
Asbestos killed my mum in her 40s - was her school to blame? To play this video you need to enable JavaScript in your browser. Figure caption, 'Mum died young from asbestos - now my family want justice' When Caroline Bryan was enjoying high school with friends in the early 90s, she would never have imagined replaying those memories in her final weeks, while searching for answers about her terminal condition. The mum was 46 when she died from mesothelioma, an incurable cancer linked to asbestos exposure. As she came to terms with her difficult diagnosis, Caroline's belief was the only place she could have encountered asbestos was during her time at school decades earlier.
Zelensky has 'questions to answer' about corruption in his government, sacked minister tells BBC
Zelensky has'questions to answer' about corruption in his government, sacked minister tells BBC To play this video you need to enable JavaScript in your browser. Ukrainian President Volodymyr Zelensky has questions to answer about corruption in his government, his sacked defence minister Mykhailo Fedorov has told the BBC. Fedorov, a popular minister who modernised the armed forces and pioneered drone warfare, was fired last month after reportedly falling out with the head of Ukraine's armed forces. He told the BBC a money laundering investigation announced last week raised questions about Zelensky's awareness of potential corruption in his inner circle. Fedorov also reiterated that democracy should not be the hostage of Putin - but stopped short of declaring his own political ambitions days after he posted a video calling for elections in Ukraine. Supporters have protested on Kyiv's streets weekly since he was sacked, calling on the president to reinstate him.
Russia says at least seven killed in largest Ukrainian attack of 2026
To play this video you need to enable JavaScript in your browser. At least seven people have been killed in Russia after what officials called the largest-scale attack launched by Ukraine this year. Some 822 drones were used - 600 aiming for Moscow, according to its regional governor who said the Russian capital had suffered one of the most massive drone attacks in recent memory. A warehouse belonging to Russian online retailer Wildberries was hit. Russian attacks on Ukraine also killed at least seven people and injured more than 39.
'Like a bus seat back in the day' - new Wales kit leaves fans divided
'Like a bus seat back in the day' - new Wales kit leaves fans divided Wales football fans have been left divided over the country's new home kit, with some comparing the pattern on the red shirt to that seen on bus seats. The new shirt was revealed on Thursday by the Football Association of Wales (FAW), which said it was inspired by the rich tradition of Welsh double cloth weaving. It is the first Wales shirt produced by sportswear brand Sudu, after the FAW parted ways with Adidas after 13 years. FAW said the reaction to football kits was subjective, adding it remained incredibly proud of the creative process behind the new Cymru home kit. The association said the kit rejected reused designs, and that Sudu had worked with the association and fans to understand the DNA of the nation.
Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
Video Large Language Models (Video-LLMs) excel at understanding videos incontext, provided they have full access to the video when answering queries. However, these models face challenges in streaming scenarios where hour-long videos must be processed online, and questions need timely responses. In this work, we propose a training-free approach compatible with standard Video-LLMs, leveraging three key concepts: 1) LLM-informed selection of visual tokens to identify those that the LLM has attended to and contributed to its understanding of each short clip. Our attention-based selection allows us to discard up to 95% of unimportant visual tokens with minimal performance loss; 2) Recurrent processing of past selected tokens to generate temporally coherent understanding of each processed clip; 3) Caption-based question answering for lightweight and accurate responses. Our method achieves state-of-the-art performance on streaming video benchmarks, striking a balance between efficiency and effectiveness.
Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
Benchmarks for out-of-distribution (OOD) generalization frequently show a strong positive correlation between in-distribution (ID) and OOD accuracy across models, termed "accuracy-on-the-line." This pattern is often taken to imply that spurious correlations--correlations that improve ID but reduce OOD performance--are rare in practice. We find that this positive correlation is often an artifact of aggregating heterogeneous OOD examples. Using a simple gradient-based method, OODSelect, we identify semantically coherent OOD subsets where accuracy on the line does not hold. Across widely used distribution shift benchmarks, the OODSelect uncovers subsets, sometimes up to over half of the standard OOD set, where higher ID accuracy predicts lower OOD accuracy. Our findings indicate that aggregate metrics can obscure important failure modes of OOD robustness. We release code and the identified subsets to facilitate further research.
LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
Despite significant progress in Video Large Language Models (Video-LLMs) for offline video understanding, existing online Video-LLMs typically struggle to simultaneously process continuous frame-by-frame inputs and determine optimal response timing, often compromising real-time responsiveness and narrative coherence. To address these limitations, we introduce LiveStar, a pioneering live streaming assistant that achieves always-on proactive responses through adaptive streaming decoding. Specifically, LiveStar incorporates: (1) a training strategy enabling incremental video-language alignment for variable-length video streams, preserving temporal consistency across dynamically evolving frame sequences; (2) a response-silence decoding framework that determines optimal proactive response timing via a single forward pass verification; (3) memory-aware acceleration via peak-end memory compression for online inference on 10+ minute videos, combined with streaming key-value cache to achieve 1.53 faster inference. We also construct an OmniStar dataset, a comprehensive dataset for training and benchmarking that encompasses 15 diverse real-world scenarios and 5 evaluation tasks for online video understanding. Extensive experiments across three benchmarks demonstrate LiveStar's state-of-the-art performance, achieving an average 19.5% improvement in semantic correctness with 18.1% reduced timing difference compared to existing online Video-LLMs, while improving FPS by 12.0% across all five OmniStar tasks.
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of wellannotated long videos has left the training of hour-long Video-LMMs underexplored. To close this gap, we present VideoMarathon, a large-scale hour-long video instruction-following dataset. This dataset includes around 9,700 hours of long videos sourced from diverse domains, ranging from 3 to 60 minutes per video. Specifically, it contains 3.3M high-quality QA pairs, spanning six fundamental topics: temporality, spatiality, object, action, scene, and event. Compared to existing video instruction datasets, VideoMarathon significantly extends training video durations up to 1 hour, and supports 22 diverse tasks requiring both short-and long-term video comprehension. Building on VideoMarathon, we propose Hour-LLaVA, a powerful and efficient Video-LMM for hour-scale video-language modeling. It enables hour-long video training and inference at 1-FPS sampling by leveraging a memory augmentation module, which adaptively integrates question-relevant and spatiotemporally informative semantics from the cached full video context. In our experiments, Hour-LLaVA achieves the best performance on multiple representative long video-language benchmarks, demonstrating the high quality of the VideoMarathon dataset and the superiority of the Hour-LLaVA model.
JavisGPT: AUnified Multi-modal LLM for Sounding-Video Comprehension and Generation
This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for joint audio-video (JAV) comprehension and generation. JavisGPT has a concise encoderLLM-decoder fusion and synchron architecture, y-aware which learnable has a queries SyncFusion to bridge module a pretrained for spatio-temporal JAV-DiT generator audio-video . This design enables temporally coherent video-audio understanding and generation from multimodal instructions. We design an effective three-stage training pipeline consisting of multimodal pretraining, audio-video fine-tuning, and large-scale instruction-tuning, to progressively build multimodal comprehension and generation from existing vision-language models. For instruction tuning, we construct JavisInst-Omni, a high-quality instruction dataset with over 200K GPT and generation -4o-curated scenarios.