Large Language Model
JRadiEvo: A Japanese Radiology Report Generation Model Enhanced by Evolutionary Optimization of Model Merging
Baba, Kaito, Yagi, Ryota, Takahashi, Junichiro, Kishikawa, Risa, Kodera, Satoshi
With the rapid advancement of large language models (LLMs), foundational models (FMs) have seen significant advancements. Healthcare is one of the most crucial application areas for these FMs, given the significant time and effort required for physicians to analyze large volumes of patient data. Recent efforts have focused on adapting multimodal FMs to the medical domain through techniques like instruction-tuning, leading to the development of medical foundation models (MFMs). However, these approaches typically require large amounts of training data to effectively adapt models to the medical field. Moreover, most existing models are trained on English datasets, limiting their practicality in non-English-speaking regions where healthcare professionals and patients are not always fluent in English. The need for translation introduces additional costs and inefficiencies. To address these challenges, we propose a \textbf{J}apanese \textbf{Radi}ology report generation model enhanced by \textbf{Evo}lutionary optimization of model merging (JRadiEvo). This is the first attempt to extend a non-medical vision-language foundation model to the medical domain through evolutionary optimization of model merging. We successfully created a model that generates accurate Japanese reports from X-ray images using only 50 translated samples from publicly available data. This model, developed with highly efficient use of limited data, outperformed leading models from recent research trained on much larger datasets. Additionally, with only 8 billion parameters, this relatively compact foundation model can be deployed locally within hospitals, making it a practical solution for environments where APIs and other external services cannot be used due to strict privacy and security requirements.
How to Get Started on Bluesky
The social media app Bluesky just reached the top of the free download charts for Apple's app store in the United States, making it--for the moment, at least--more popular than Meta's Threads and OpenAI's ChatGPT. The decentralized social media platform has received a fresh influx of users over the last week as more X users sour on Elon Musk's political ambitions and abandon his social media platform for an alternative. Over a million new users have joined Bluesky since the 2024 US presidential election on November 5, which was shaped by Musk's influence. First launched in 2019 as a project within Twitter, Bluesky gained independence from the company before Musk's acquisition and its subsequent name change. Bluesky also captured the attention of some ex-tweeters back in 2023, when new users were only able to sign up through an invite system.
The First Entirely AI-Generated Video Game Is Insanely Weird and Fun
Minecraft remains remarkably popular a decade or so after it was first released, thanks to a unique mix of quirky gameplay and open world building possibilities. A knock-off called Oasis, released last month, captures much of the original game's flavor with a remarkable and weird twist. The entire game is generated not by a game engine and hand-coded rules, but by an AI model that dreams up each frame. Oasis was built by an Israeli AI startup called Decart in collaboration with Etched, a company that designs custom silicon, to demonstrate the potential of hardware optimized to power transformer-based AI algorithms. Oasis uses a transformer AI model, similar to the one that powers a large language model--only trained, apparently, on endless examples of people playing Minecraft, to dream up each new video frame in response to the previous one and to user input like clicks or mouse moves.
Fox News AI Newsletter: AI developers discover 'Donald Trump neuron', expert says
Kurt'CyberGuy' Knutsson on President-elect Trump's plan to deregulate cryptocurrency and A.I. in his second administration. 'DONALD TRUMP NEURON': Artificial intelligence recognizes images and the name of President-elect Donald Trump so much that the phenomenon is referred to as a "Donald Trump neuron," expert Chris Olah says. MUSK PETITION: An artificial intelligence (AI) advocacy group is urging President-elect Trump to make billionaire entrepreneur Elon Musk a special adviser to the White House focused on AI. INDIA - 2024/05/17: In this photo illustration, the OpenAI logo is seen displayed on a mobile phone screen with ChatGPT logo in the background. HELP FROM SILICON VALLEY: OpenAI has assembled a "blueprint" for artificial intelligence infrastructure that the company hopes will be considered by the incoming Trump administration and Congress – suggesting that the plan will help the United States maintain its lead in the field over competitors like China.
How to have Microsoft Copilot recap a Teams meeting when you're late
Running late to a Microsoft Teams meeting can land you seriously out of sync with your colleagues and the latest updates from work. A coworker is resigning, you're about to get a promotion, the company is being liquidated -- these are just some of the real-life bombshells that could have been dropped while you were waiting for your morning coffee order to arrive. You could spend the whole rest of the meeting interrupting your colleagues to find out what juicy tidbits you missed, but doing that isn't going to win you any fans. A better option is to just use Copilot to recap what happened in your absence. For Copilot recap to work in Teams you'll need to have a current license for Microsoft 365 and Copilot.
OpenAI touts AI infrastructure 'blueprint' to outcompete China, bolster economy under incoming Trump admin
Kurt'CyberGuy' Knutsson on President-elect Trump's plan to deregulate cryptocurrency and A.I. in his second administration. OpenAI has assembled a "blueprint" for artificial intelligence (AI) infrastructure that the company hopes will be considered by the incoming Trump administration and Congress – suggesting that the plan will help the United States maintain its lead in the field over competitors like China. The company's Vice President of Global Affairs, Chris Lehane, announced the "Infrastructure Blueprint for the U.S." on Wednesday during an event hosted by the Center for Strategic and International Studies (CSIS). The company says AI's potential presents an "unmissable opportunity to revitalize the American Dream and reindustrialize the US." "Investments to extend the current U.S. lead in AI will yield tens of thousands of skilled-trade and other jobs, growth in productivity and GDP; a modernized grid including power generated by nuclear energy; a state-of-the-art network of semiconductor manufacturing facilities; and a new generation of AI-powered businesses and entrepreneurship," OpenAI claims. In this photo illustration, the OpenAI logo is seen displayed on a mobile phone screen with ChatGPT logo in the background.
Data movement limits to frontier model training
Erdil, Ege, Schneider-Joseph, David
We present a theoretical model of distributed training, and use it to analyze how far dense and sparse training runs can be scaled. FLOP, two orders of magnitude above the largest training run to date, suggesting the arrival of fundamental barriers to scaling in three years given recent rates of growth. FLOP is infeasible even at low utilization. However, more aggressive batch size scaling and/or shorter and fatter model shapes, if achievable, have the potential to permit much larger training runs. An interactive version of our model will shortly be accessible here. In this work, we address unexamined fundamental questions about limits to scaling in the future: Q1 Given present-day algorithms, GPUs, and interconnects, what is the biggest training run that can be performed within a fixed duration, before intra-and inter-GPU data movement starts to seriously worsen utilization or even render it impossible? Q2 How far might this limit be extended, and what algorithmic or hardware progress can achieve that? Answering these questions empirically would require millions of GPUs and large-scale engineering efforts, so we instead approach them theoretically. In doing so, we develop a simulator that can find optimal training run configurations accounting for the factors that we identify as fundamental. We focus on GPUs, but our theoretical model and findings are broadly applicable to other accelerators, and even groups of accelerators. A2 Improved hardware interconnects may buy no more than two orders of magnitude in training run size, assuming technology anything like the current paradigm. Beyond that, the critical innovations must come from machine learning algorithms: The key challenge is transforming two serial dependencies -- between batches and between layers -- into opportunities for parallelism, by making batch sizes bigger (perhaps enabled by sparsity) and models wider and shallower. Achieving these goals may be quite difficult in practice. However, with innovations in scaling (such as techniques to enable much larger batch sizes) or dramatic increases in network bandwidth coupled with a 10 reduction in interand intra-GPU latency, training runs can be at least a few orders of magnitude larger (right).
VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation
Wen, Youpeng, Lin, Junfan, Zhu, Yi, Han, Jianhua, Xu, Hang, Zhao, Shen, Liang, Xiaodan
Recent advancements utilizing large-scale video data for learning video generation models demonstrate significant potential in understanding complex physical dynamics. It suggests the feasibility of leveraging diverse robot trajectory data to develop a unified, dynamics-aware model to enhance robot manipulation. However, given the relatively small amount of available robot data, directly fitting data without considering the relationship between visual observations and actions could lead to suboptimal data utilization. To this end, we propose VidMan (Video Diffusion for Robot Manipulation), a novel framework that employs a two-stage training mechanism inspired by dual-process theory from neuroscience to enhance stability and improve data utilization efficiency. Specifically, in the first stage, VidMan is pre-trained on the Open X-Embodiment dataset (OXE) for predicting future visual trajectories in a video denoising diffusion manner, enabling the model to develop a long horizontal awareness of the environment's dynamics. In the second stage, a flexible yet effective layer-wise self-attention adapter is introduced to transform VidMan into an efficient inverse dynamics model that predicts action modulated by the implicit dynamics knowledge via parameter sharing. Our VidMan framework outperforms state-of-the-art baseline model GR-1 on the CALVIN benchmark, achieving a 11.7% relative improvement, and demonstrates over 9% precision gains on the OXE small-scale dataset. These results provide compelling evidence that world models can significantly enhance the precision of robot action prediction. Codes and models will be public.
Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
Ouyang, Rongxin, Jaidka, Kokil, Mukerjee, Subhayan, Cui, Guangyu
The prevalence of multi-modal content on social media complicates automated moderation strategies. This calls for an enhancement in multi-modal classification and a deeper understanding of understated meanings in images and memes. Although previous efforts have aimed at improving model performance through fine-tuning, few have explored an end-to-end optimization pipeline that accounts for modalities, prompting, labeling, and fine-tuning. In this study, we propose an end-to-end conceptual framework for model optimization in complex tasks. Experiments support the efficacy of this traditional yet novel framework, achieving the highest accuracy and AUROC. Ablation experiments demonstrate that isolated optimizations are not ineffective on their own.
DROJ: A Prompt-Driven Attack against Large Language Models
Large Language Models (LLMs) have demonstrated exceptional capabilities across various natural language processing tasks. Due to their training on internet-sourced datasets, LLMs can sometimes generate objectionable content, necessitating extensive alignment with human feedback to avoid such outputs. Despite massive alignment efforts, LLMs remain susceptible to adversarial jailbreak attacks, which usually are manipulated prompts designed to circumvent safety mechanisms and elicit harmful responses. Here, we introduce a novel approach, Directed Rrepresentation Optimization Jailbreak (DROJ), which optimizes jailbreak prompts at the embedding level to shift the hidden representations of harmful queries towards directions that are more likely to elicit affirmative responses from the model. Our evaluations on LLaMA-2-7b-chat model show that DROJ achieves a 100\% keyword-based Attack Success Rate (ASR), effectively preventing direct refusals. However, the model occasionally produces repetitive and non-informative responses. To mitigate this, we introduce a helpfulness system prompt that enhances the utility of the model's responses. Our code is available at https://github.com/Leon-Leyang/LLM-Safeguard.