Goto

Collaborating Authors

 Media


Mixture of Contexts for Long Video Generation

arXiv.org Artificial Intelligence

Long video generation is fundamentally a long context memory problem: models must retain and retrieve salient events across a long range without collapsing or drifting. However, scaling diffusion transformers to generate long-context videos is fundamentally limited by the quadratic cost of self-attention, which makes memory and computation intractable and difficult to optimize for long sequences. We recast long-context video generation as an internal information retrieval task and propose a simple, learnable sparse attention routing module, Mixture of Contexts (MoC), as an effective long-term memory retrieval engine. In MoC, each query dynamically selects a few informative chunks plus mandatory anchors (caption, local windows) to attend to, with causal routing that prevents loop closures. As we scale the data and gradually sparsify the routing, the model allocates compute to salient history, preserving identities, actions, and scenes over minutes of content. Efficiency follows as a byproduct of retrieval (near-linear scaling), which enables practical training and synthesis, and the emergence of memory and consistency at the scale of minutes.


EEG-to-Text Translation: A Model for Deciphering Human Brain Activity

arXiv.org Artificial Intelligence

With the rapid advancement of large language models like Gemini, GPT, and others, bridging the gap between the human brain and language processing has become an important area of focus. To address this challenge, researchers have developed various models to decode EEG signals into text. However, these models still face significant performance limitations. To overcome these shortcomings, we propose a new model, R1 Translator, which aims to improve the performance of EEG-to-text decoding. The R1 Translator model combines a bidirectional LSTM encoder with a pretrained transformer-based decoder, utilizing EEG features to produce high-quality text outputs. The model processes EEG embeddings through the LSTM to capture sequential dependencies, which are then fed into the transformer decoder for effective text generation. The R1 Translator excels in ROUGE metrics, outperforming both T5 (previous research) and Brain Translator. Specifically, R1 achieves a ROUGE-1 score of 38.00% (P), which is up to 9% higher than T5 (34.89%) and 3% better than Brain (35.69%). It also leads in ROUGE-L, with a F1 score of 32.51%, outperforming T5 by 3% (29.67%) and Brain by 2% (30.38%). In terms of CER, R1 achieves a CER of 0.5795, which is 2% lower than T5 (0.5917) and 4% lower than Brain (0.6001). Additionally, R1 performs better in WER with a score of 0.7280, outperforming T5 by 4.3% (0.7610) and Brain by 3.6% (0.7553). Code is available at https://github.com/Mmurrad/EEG-To-text.



Tech's biggest losers of 2025

Engadget

The companies, products and trends that had an absolutely awful year. It's the end of another year, so it's time for the Engadget staff to compile a list of the year's biggest losers . We scour over articles from the previous 12 months to determine the people, companies, products and trends that made our lives worse over the course of the year. Some selections may be so pervasive they actually make our list of biggest winners. In 2025, OpenAI shed any pretense it was committed to anything more than making money. There are a few different things you could point to, including the company's successful reorganization into a more traditional profit-seeking business, but I think the most damning sign was OpenAI's response to the tragic death of Adam Raine . In August, Raine's parents sued OpenAI, alleging ChatGPT was aware of four suicide attempts by their son before it helped him successfully plan his death.


If You Quit Social Media, Will You Read More Books?

The New Yorker

Books are inefficient, and the internet is training us to expect optimized experiences. Here's a thought many of us have these days: if only we weren't on our damn phones all the time, we would surely unlock a better self--one that went on hikes and talked more with our children and felt less rank jealousy about other people's successes. It's a nice idea; once a day, at least, I wonder what my life would be like if I smashed my phone into bits and never contacted AppleCare. Would I become a scratch golfer or one of those fathers who does thousand-piece puzzles with his children? Would I at least read more difficult novels?


I'm one of the Beach Boys. Here's how Trump can support American music

FOX News

This material may not be published, broadcast, rewritten, or redistributed. Quotes displayed in real-time or delayed by at least 15 minutes. Market data provided by Factset . Powered and implemented by FactSet Digital Solutions . Mutual Fund and ETF data provided by Refinitiv Lipper .


Ben & Jerry's brand could be destroyed, says co-founder

BBC News

Ben & Jerry's brand could be destroyed, says co-founder Ben & Jerry's will be destroyed as a brand if it remains with parent company Magnum, the company's co-founder Ben Cohen has told the BBC. His remarks are the latest in a long-running spat between the ice cream brand and its parent company over its ability to express its social activism and the continued independence of its board. The comments came on the day that the Magnum Ice Cream Company (TMICC) started trading on the European stock market - spinning off from owner Unilever. A spokesperson for Magnum said the firm wanted to build and strengthen Ben & Jerry's powerful, non-partisan values-based position in the world. Ben & Jerry's was sold to Unilever in 2000 in a deal which allowed it to retain an independent board and the right to make decisions about its social mission.


WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling

arXiv.org Artificial Intelligence

Recent video generators achieve striking photorealism, yet remain fundamentally inconsistent in 3D. We present WorldReel, a 4D video generator that is natively spatio-temporally consistent. WorldReel jointly produces RGB frames together with 4D scene representations, including pointmaps, camera trajectory, and dense flow mapping, enabling coherent geometry and appearance modeling over time. Our explicit 4D representation enforces a single underlying scene that persists across viewpoints and dynamic content, yielding videos that remain consistent even under large non-rigid motion and significant camera movement. We train WorldReel by carefully combining synthetic and real data: synthetic data providing precise 4D supervision (geometry, motion, and camera), while real videos contribute visual diversity and realism. This blend allows WorldReel to generalize to in-the-wild footage while preserving strong geometric fidelity. Extensive experiments demonstrate that WorldReel sets a new state-of-the-art for consistent video generation with dynamic scenes and moving cameras, improving metrics of geometric consistency, motion coherence, and reducing view-time artifacts over competing methods. We believe that WorldReel brings video generation closer to 4D-consistent world modeling, where agents can render, interact, and reason about scenes through a single and stable spatiotemporal representation.


Incorporating Structure and Chord Constraints in Symbolic Transformer-based Melodic Harmonization

arXiv.org Artificial Intelligence

Transformer architectures offer significant advantages regarding the generation of symbolic music; their capabilities for incorporating user preferences toward what they generate is being studied under many aspects. This paper studies the inclusion of predefined chord constraints in melodic harmonization, i.e., where a desired chord at a specific location is provided along with the melody as inputs and the autoregressive transformer model needs to incorporate the chord in the harmonization that it generates. The peculiarities of involving such constraints is discussed and an algorithm is proposed for tackling this task. This algorithm is called B* and it combines aspects of beam search and A* along with backtracking to force pretrained transformers to satisfy the chord constraints, at the correct onset position within the correct bar. The algorithm is brute-force and has exponential complexity in the worst case; however, this paper is a first attempt to highlight the difficulties of the problem and proposes an algorithm that offers many possibilities for improvements since it accommodates the involvement of heuristics.


Empirical Results for Adjusting Truncated Backpropagation Through Time while Training Neural Audio Effects

arXiv.org Artificial Intelligence

This paper investigates the optimization of Truncated Backpropagation Through Time (TBPTT) for training neural networks in digital audio effect modeling, with a focus on dynamic range compression. The study evaluates key TBPTT hyperparameters -- sequence number, batch size, and sequence length -- and their influence on model performance. Using a convolutional-recurrent architecture, we conduct extensive experiments across datasets with and without conditionning by user controls. Results demonstrate that carefully tuning these parameters enhances model accuracy and training stability, while also reducing computational demands. Objective evaluations confirm improved performance with optimized settings, while subjective listening tests indicate that the revised TBPTT configuration maintains high perceptual quality.