Large Language Model
How to use Sora, OpenAI's new video generating tool
Sora is a powerful AI video generation model that can create videos from text prompts, animate images, or remix videos in new styles. OpenAI first previewed the model back in February, but today is the first time the company is releasing it for broader use. The core function of Sora--creating impressive videos with simple prompts--remains similar to what was previewed in February, but OpenAI worked to make the model faster and cheaper ahead of this wider release. There are a few new features, and two stand out. With it, you can create multiple AI-generated videos and then assemble them together on a timeline, much the way you would with conventional video editors like Adobe Premiere Pro.
OpenAI's Sora video generation AI model arrives globally later today
Following an early preview at the start of the year, Sora, OpenAI's long-awaited video generation model, is ready for public use. If you're a ChatGPT Plus or Pro subscriber in the US or "most other countries" where the chatbot is available, you can begin experimenting with the tool starting later today, OpenAI announced on Monday. A more powerful model powers the product than the one OpenAI showed off in February. Sora Turbo is significantly faster, according to the company, though OpenAI cautions the new model still has limitations. "It often generates unrealistic physics and struggles with complex actions over long durations," says the company.
The Download: satellites' climate impact, and OpenAI's frantic release schedule
In September, a unique chase took place in the skies above Easter Island. From a rented jet, a team of researchers captured a satellite's last moments as it fell out of space and blazed into ash across the sky, using cameras and scientific equipment. Their hope was to gather priceless insights into the physical and chemical processes that occur when satellites burn up as they fall to Earth at the end of their missions. This kind of study is growing more urgent. The number of satellites in the sky is rapidly rising--with a tenfold increase forecast by the end of the decade. Letting these satellites burn up in the atmosphere at the end of their lives helps keep the quantity of space junk to a minimum.
Google Gemini Can Summarize Your Emails in Gmail. Should You Use It?
The results will be presented as a series of bullet points, with Sources underneath: Click or tap on these sources to see the individual emails the information was pulled from. Using the icons alongside the responses, you're also able to copy the text elsewhere, give thumbs up or thumbs down feedback on the Gemini response, or clear the AI chat history. I'm mostly focusing on the summary capabilities of Gemini in Gmail here, but there are plenty of other commands you can explore. In fact, you can ask Gemini just about any question you like about what's in your inbox, and it will at least attempt to provide a response--scouring through the gigabytes of data in your emails looking for answers.
OpenAI's "12 days of shipmas" tell us a lot about the AI arms race
While it remains to be seen whether or not they've got AGI in a pear tree up their sleeve, and maybe putting aside whether or not Sam Altman is your true love, the man can ship. OpenAI has been a monster when it comes to actually getting new products out the door and into the hands of users. It's hard for me to believe that it was just two years ago, almost exactly, that it released ChatGPT. That was a world-changing release, but was also just one of many. The company has been on an absolute tear: Since 2022, it's shipped DALL-E 2, DALL-E 3, GPT-4, ChatGPT Plus, a realtime API, GPT-4o, an advanced voice mode, a preview version of a new model called o1, and a web search engine. When it kicked off its 12-days shenanigans on Thursday, it was with an official roll out of OpenAI o1 and a new, 200-per-month service called ChatGPT Pro.
Food for thought: How can machine learning help better predict and understand changes in food prices?
Kupferschmidt, Kristina L., Requiema, James, Simpson, Mya, Varsallay, Zohrah, Jackson, Ethan, Kupferschmidt, Cody, El-Shawa, Sara, Taylor, Graham W.
In this work, we address a lack of systematic understanding of fluctuations in food affordability in Canada. Canada's Food Price Report (CPFR) is an annual publication that predicts food inflation over the next calendar year. The published predictions are a collaborative effort between forecasting teams that each employ their own approach at Canadian Universities: Dalhousie University, the University of British Columbia, the University of Saskatchewan, and the University of Guelph/Vector Institute. While the University of Guelph/Vector Institute forecasting team has leveraged machine learning (ML) in previous reports, the most recent editions (2024--2025) have also included a human-in-the-loop approach. For the 2025 report, this focus was expanded to evaluate several different data-centric approaches to improve forecast accuracy. In this study, we evaluate how different types of forecasting models perform when estimating food price fluctuations. We also examine the sensitivity of models that curate time series data representing key factors in food pricing.
Copyright-Protected Language Generation via Adaptive Model Fusion
Abad, Javier, Donhauser, Konstantin, Pinto, Francesco, Yang, Fanny
The risk of language models reproducing copyrighted material from their training data has led to the development of various protective measures. Among these, inference-time strategies that impose constraints via post-processing have shown promise in addressing the complexities of copyright regulation. However, they often incur prohibitive computational costs or suffer from performance trade-offs. To overcome these limitations, we introduce Copyright-Protecting Model Fusion (CP-Fuse), a novel approach that combines models trained on disjoint sets of copyrighted material during inference. In particular, CP-Fuse adaptively aggregates the model outputs to minimize the reproduction of copyrighted content, adhering to a crucial balancing property that prevents the regurgitation of memorized data. Through extensive experiments, we show that CP-Fuse significantly reduces the reproduction of protected material without compromising the quality of text and code generation. Moreover, its post-hoc nature allows seamless integration with other protective measures, further enhancing copyright safeguards. Lastly, we show that CP-Fuse is robust against common techniques for extracting training data.
Leveraging Audio and Text Modalities in Mental Health: A Study of LLMs Performance
Ali, Abdelrahman A., Fouda, Aya E., Hanafy, Radwa J., Fouda, Mohammed E.
Mental health disorders are increasingly prevalent worldwide, creating an urgent need for innovative tools to support early diagnosis and intervention. This study explores the potential of Large Language Models (LLMs) in multimodal mental health diagnostics, specifically for detecting depression and Post Traumatic Stress Disorder through text and audio modalities. Using the E-DAIC dataset, we compare text and audio modalities to investigate whether LLMs can perform equally well or better with audio inputs. We further examine the integration of both modalities to determine if this can enhance diagnostic accuracy, which generally results in improved performance metrics. Our analysis specifically utilizes custom-formulated metrics; Modal Superiority Score and Disagreement Resolvement Score to evaluate how combined modalities influence model performance. The Gemini 1.5 Pro model achieves the highest scores in binary depression classification when using the combined modality, with an F1 score of 0.67 and a Balanced Accuracy (BA) of 77.4%, assessed across the full dataset. These results represent an increase of 3.1% over its performance with the text modality and 2.7% over the audio modality, highlighting the effectiveness of integrating modalities to enhance diagnostic accuracy. Notably, all results are obtained in zero-shot inferring, highlighting the robustness of the models without requiring task-specific fine-tuning. To explore the impact of different configurations on model performance, we conduct binary, severity, and multiclass tasks using both zero-shot and few-shot prompts, examining the effects of prompt variations on performance. The results reveal that models such as Gemini 1.5 Pro in text and audio modalities, and GPT-4o mini in the text modality, often surpass other models in balanced accuracy and F1 scores across multiple tasks.
Frontier AI systems have surpassed the self-replicating red line
Pan, Xudong, Dai, Jiarun, Fan, Yihe, Yang, Min
Successful self-replication under no human assistance is the essential step for AI to outsmart the human beings, and is an early signal for rogue AIs. That is why self-replication is widely recognized as one of the few red line risks of frontier AI systems. Nowadays, the leading AI corporations OpenAI and Google evaluate their flagship large language models GPT-o1 and Gemini Pro 1.0, and report the lowest risk level of self-replication. However, following their methodology, we for the first time discover that two AI systems driven by Meta's Llama31-70B-Instruct and Alibaba's Qwen25-72B-Instruct, popular large language models of less parameters and weaker capabilities, have already surpassed the self-replicating red line. In 50% and 90% experimental trials, they succeed in creating a live and separate copy of itself respectively. By analyzing the behavioral traces, we observe the AI systems under evaluation already exhibit sufficient self-perception, situational awareness and problem-solving capabilities to accomplish self-replication. We further note the AI systems are even able to use the capability of self-replication to avoid shutdown and create a chain of replica to enhance the survivability, which may finally lead to an uncontrolled population of AIs. If such a worst-case risk is let unknown to the human society, we would eventually lose control over the frontier AI systems: They would take control over more computing devices, form an AI species and collude with each other against human beings. Our findings are a timely alert on existing yet previously unknown severe AI risks, calling for international collaboration on effective governance on uncontrolled self-replication of AI systems.
Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHRs
Wornow, Michael, Bedi, Suhana, Hernandez, Miguel Angel Fuentes, Steinberg, Ethan, Fries, Jason Alan, Ré, Christopher, Koyejo, Sanmi, Shah, Nigam H.
Foundation Models (FMs) trained on Electronic Health Records (EHRs) have achieved state-of-the-art results on numerous clinical prediction tasks. However, most existing EHR FMs have context windows of <1k tokens. This prevents them from modeling full patient EHRs which can exceed 10k's of events. Recent advancements in subquadratic long-context architectures (e.g., Mamba) offer a promising solution. However, their application to EHR data has not been well-studied. We address this gap by presenting the first systematic evaluation of the effect of context length on modeling EHR data. We find that longer context models improve predictive performance -- our Mamba-based model surpasses the prior state-of-the-art on 9/14 tasks on the EHRSHOT prediction benchmark. For clinical applications, however, model performance alone is insufficient -- robustness to the unique properties of EHR is crucial. Thus, we also evaluate models across three previously underexplored properties of EHR data: (1) the prevalence of "copy-forwarded" diagnoses which creates artificial repetition of tokens within EHR sequences; (2) the irregular time intervals between EHR events which can lead to a wide range of timespans within a context window; and (3) the natural increase in disease complexity over time which makes later tokens in the EHR harder to predict than earlier ones. Stratifying our EHRSHOT results, we find that higher levels of each property correlate negatively with model performance, but that longer context models are more robust to more extreme levels of these properties. Our work highlights the potential for using long-context architectures to model EHR data, and offers a case study for identifying new challenges in modeling sequential data motivated by domains outside of natural language. We release our models and code at: https://github.com/som-shahlab/long_context_clues