Africa
Brain Tumor Segmentation (BraTS) Challenge 2024: Meningioma Radiotherapy Planning Automated Segmentation
LaBella, Dominic, Schumacher, Katherine, Mix, Michael, Leu, Kevin, McBurney-Lin, Shan, Nedelec, Pierre, Villanueva-Meyer, Javier, Shapey, Jonathan, Vercauteren, Tom, Chia, Kazumi, Al-Salihi, Omar, Leu, Justin, Halasz, Lia, Velichko, Yury, Wang, Chunhao, Kirkpatrick, John, Floyd, Scott, Reitman, Zachary J., Mullikin, Trey, Bagci, Ulas, Sachdev, Sean, Hattangadi-Gluth, Jona A., Seibert, Tyler, Farid, Nikdokht, Puett, Connor, Pease, Matthew W., Shiue, Kevin, Anwar, Syed Muhammad, Faghani, Shahriar, Haider, Muhammad Ammar, Warman, Pranav, Albrecht, Jake, Jakab, András, Moassefi, Mana, Chung, Verena, Aristizabal, Alejandro, Karargyris, Alexandros, Kassem, Hasan, Pati, Sarthak, Sheller, Micah, Huang, Christina, Coley, Aaron, Ghanta, Siddharth, Schneider, Alex, Sharp, Conrad, Saluja, Rachit, Kofler, Florian, Lohmann, Philipp, Vollmuth, Phillipp, Gagnon, Louis, Adewole, Maruf, Li, Hongwei Bran, Kazerooni, Anahita Fathi, Tahon, Nourel Hoda, Anazodo, Udunna, Moawad, Ahmed W., Menze, Bjoern, Linguraru, Marius George, Aboian, Mariam, Wiestler, Benedikt, Baid, Ujjwal, Conte, Gian-Marco, Rauschecker, Andreas M. T., Nada, Ayman, Abayazeed, Aly H., Huang, Raymond, de Verdier, Maria Correia, Rudie, Jeffrey D., Bakas, Spyridon, Calabrese, Evan
The 2024 Brain Tumor Segmentation Meningioma Radiotherapy (BraTS-MEN-RT) challenge aims to advance automated segmentation algorithms using the largest known multi-institutional dataset of radiotherapy planning brain MRIs with expert-annotated target labels for patients with intact or post-operative meningioma that underwent either conventional external beam radiotherapy or stereotactic radiosurgery. Each case includes a defaced 3D post-contrast T1-weighted radiotherapy planning MRI in its native acquisition space, accompanied by a single-label "target volume" representing the gross tumor volume (GTV) and any at-risk post-operative site. Target volume annotations adhere to established radiotherapy planning protocols, ensuring consistency across cases and institutions. For pre-operative meningiomas, the target volume encompasses the entire GTV and associated nodular dural tail, while for post-operative cases, it includes at-risk resection cavity margins as determined by the treating institution. Case annotations were reviewed and approved by expert neuroradiologists and radiation oncologists. Participating teams will develop, containerize, and evaluate automated segmentation models using this comprehensive dataset. Model performance will be assessed using the lesion-wise Dice Similarity Coefficient and the 95% Hausdorff distance. The top-performing teams will be recognized at the Medical Image Computing and Computer Assisted Intervention Conference in October 2024. BraTS-MEN-RT is expected to significantly advance automated radiotherapy planning by enabling precise tumor segmentation and facilitating tailored treatment, ultimately improving patient outcomes.
Mutation-Bias Learning in Games
Bauer, Johann, West, Sheldon, Alonso, Eduardo, Broom, Mark
We present two variants of a multi-agent reinforcement learning algorithm based on evolutionary game theoretic considerations. The intentional simplicity of one variant enables us to prove results on its relationship to a system of ordinary differential equations of replicator-mutator dynamics type, allowing us to present proofs on the algorithm's convergence conditions in various settings via its ODE counterpart. The more complicated variant enables comparisons to Q-learning based algorithms. We compare both variants experimentally to WoLF-PHC and frequency-adjusted Q-learning on a range of settings, illustrating cases of increasing dimensionality where our variants preserve convergence in contrast to more complicated algorithms. The availability of analytic results provides a degree of transferability of results as compared to purely empirical case studies, illustrating the general utility of a dynamical systems perspective on multi-agent reinforcement learning when addressing questions of convergence and reliable generalisation.
Position: Towards Implicit Prompt For Text-To-Image Models
Yang, Yue, Lin, Yuqi, Liu, Hong, Shao, Wenqi, Chen, Runjian, Shang, Hailong, Wang, Yu, Qiao, Yu, Zhang, Kaipeng, Luo, Ping
Recent text-to-image (T2I) models have had great success, and many benchmarks have been proposed to evaluate their performance and safety. However, they only consider explicit prompts while neglecting implicit prompts (hint at a target without explicitly mentioning it). These prompts may get rid of safety constraints and pose potential threats to the applications of these models. This position paper highlights the current state of T2I models toward implicit prompts. We present a benchmark named ImplicitBench and conduct an investigation on the performance and impacts of implicit prompts with popular T2I models. Specifically, we design and collect more than 2,000 implicit prompts of three aspects: General Symbols, Celebrity Privacy, and Not-Safe-For-Work (NSFW) Issues, and evaluate six well-known T2I models' capabilities under these implicit prompts. Experiment results show that (1) T2I models are able to accurately create various target symbols indicated by implicit prompts; (2) Implicit prompts bring potential risks of privacy leakage for T2I models. (3) Constraints of NSFW in most of the evaluated T2I models can be bypassed with implicit prompts. We call for increased attention to the potential and risks of implicit prompts in the T2I community and further investigation into the capabilities and impacts of implicit prompts, advocating for a balanced approach that harnesses their benefits while mitigating their risks.
A Unified Temporal Knowledge Graph Reasoning Model Towards Interpolation and Extrapolation
Chen, Kai, Wang, Ye, Li, Yitong, Li, Aiping, Yu, Han, Song, Xin
Temporal knowledge graph (TKG) reasoning has two settings: interpolation reasoning and extrapolation reasoning. Both of them draw plenty of research interest and have great significance. Methods of the former de-emphasize the temporal correlations among facts sequences, while methods of the latter require strict chronological order of knowledge and ignore inferring clues provided by missing facts of the past. These limit the practicability of TKG applications as almost all of the existing TKG reasoning methods are designed specifically to address either one setting. To this end, this paper proposes an original Temporal PAth-based Reasoning (TPAR) model for both the interpolation and extrapolation reasoning. TPAR performs a neural-driven symbolic reasoning fashion that is robust to ambiguous and noisy temporal data and with fine interpretability as well. Comprehensive experiments show that TPAR outperforms SOTA methods on the link prediction task for both the interpolation and the extrapolation settings. A novel pipeline experimental setting is designed to evaluate the performances of SOTA combinations and the proposed TPAR towards interpolation and extrapolation reasoning. More diverse experiments are conducted to show the robustness and interpretability of TPAR.
Why are Visually-Grounded Language Models Bad at Image Classification?
Zhang, Yuhui, Unell, Alyssa, Wang, Xiaohan, Ghosh, Dhruba, Su, Yuchang, Schmidt, Ludwig, Yeung-Levy, Serena
Image classification is one of the most fundamental capabilities of machine vision intelligence. In this work, we revisit the image classification task using visually-grounded language models (VLMs) such as GPT-4V and LLaVA. We find that existing proprietary and public VLMs, despite often using CLIP as a vision encoder and having many more parameters, significantly underperform CLIP on standard image classification benchmarks like ImageNet. To understand the reason, we explore several hypotheses concerning the inference algorithms, training objectives, and data processing in VLMs. Our analysis reveals that the primary cause is data-related: critical information for image classification is encoded in the VLM's latent space but can only be effectively decoded with enough training data. Specifically, there is a strong correlation between the frequency of class exposure during VLM training and instruction-tuning and the VLM's performance in those classes; when trained with sufficient data, VLMs can match the accuracy of state-of-the-art classification models. Based on these findings, we enhance a VLM by integrating classification-focused datasets into its training, and demonstrate that the enhanced classification performance of the VLM transfers to its general capabilities, resulting in an improvement of 11.8% on the newly collected ImageWikiQA dataset.
Bridging Mini-Batch and Asymptotic Analysis in Contrastive Learning: From InfoNCE to Kernel-Based Losses
Koromilas, Panagiotis, Bouritsas, Giorgos, Giannakopoulos, Theodoros, Nicolaou, Mihalis, Panagakis, Yannis
What do different contrastive learning (CL) losses actually optimize for? Although multiple CL methods have demonstrated remarkable representation learning capabilities, the differences in their inner workings remain largely opaque. In this work, we analyse several CL families and prove that, under certain conditions, they admit the same minimisers when optimizing either their batch-level objectives or their expectations asymptotically. In both cases, an intimate connection with the hyperspherical energy minimisation (HEM) problem resurfaces. Drawing inspiration from this, we introduce a novel CL objective, coined Decoupled Hyperspherical Energy Loss (DHEL). DHEL simplifies the problem by decoupling the target hyperspherical energy from the alignment of positive examples while preserving the same theoretical guarantees. Going one step further, we show the same results hold for another relevant CL family, namely kernel contrastive learning (KCL), with the additional advantage of the expected loss being independent of batch size, thus identifying the minimisers in the non-asymptotic regime. Empirical results demonstrate improved downstream performance and robustness across combinations of different batch sizes and hyperparameters and reduced dimensionality collapse, on several computer vision datasets.
Signal-Plus-Noise Decomposition of Nonlinear Spiked Random Matrix Models
Moniri, Behrad, Hassani, Hamed
In this paper, we study a nonlinear spiked random matrix model where a nonlinear function is applied element-wise to a noise matrix perturbed by a rank-one signal. We establish a signal-plus-noise decomposition for this model and identify precise phase transitions in the structure of the signal components at critical thresholds of signal strength. To demonstrate the applicability of this decomposition, we then utilize it to study new phenomena in the problems of signed signal recovery in nonlinear models and community detection in transformed stochastic block models. Finally, we validate our results through a series of numerical simulations.
Bias in Motion: Theoretical Insights into the Dynamics of Bias in SGD Training
Jain, Anchit, Nobahari, Rozhin, Baratin, Aristide, Mannelli, Stefano Sarao
Over the past decade, the problem of assessing the fairness of classifiers has garnered significant attention, revealing that machine learning (ML) systems not only reproduce existing biases in the data but also tend to amplify them [1, 2, 3]. Given the complexity of the ML pipeline, isolating and characterising the key drivers of this amplification is challenging. Recent studies have begun to disentangle the contributions from architectural design choices, including overparameterisation [4], model complexity, activation functions [5, 6], learning protocols [7, 8], post-processing practices such as pruning [9], and intrinsic aspects of the data like its geometrical properties [10]. Theoretical results in this area (e.g., [4, 10]) are mostly based on asymptotic analysis, leaving the transient learning regime poorly understood. Due to limitations on computational resources, a trained ML system may operate far from the asymptotic regime and hence existing results may not always apply. Insights from class imbalance literature [7, 6] indicate that classifiers converge faster for classes with more data, but how this applies to fairness, where datasets might be balanced by label but imbalanced by demographics, remains unclear.
Tensor Methods in High Dimensional Data Analysis: Opportunities and Challenges
Auddy, Arnab, Xia, Dong, Yuan, Ming
Large amount of multidimensional data represented by multiway arrays or tensors are prevalent in modern applications across various fields such as chemometrics, genomics, physics, psychology, and signal processing. The structural complexity of such data provides vast new opportunities for modeling and analysis, but efficiently extracting information content from them, both statistically and computationally, presents unique and fundamental challenges. Addressing these challenges requires an interdisciplinary approach that brings together tools and insights from statistics, optimization and numerical linear algebra among other fields. Despite these hurdles, significant progress has been made in the last decade. This review seeks to examine some of the key advancements and identify common threads among them, under eight different statistical settings.
Optimality of Approximate Message Passing Algorithms for Spiked Matrix Models with Rotationally Invariant Noise
Dudeja, Rishabh, Liu, Songbin, Ma, Junjie
We study the problem of estimating a rank one signal matrix from an observed matrix generated by corrupting the signal with additive rotationally invariant noise. We develop a new class of approximate message-passing algorithms for this problem and provide a simple and concise characterization of their dynamics in the high-dimensional limit. At each iteration, these algorithms exploit prior knowledge about the noise structure by applying a non-linear matrix denoiser to the eigenvalues of the observed matrix and prior information regarding the signal structure by applying a non-linear iterate denoiser to the previous iterates generated by the algorithm. We exploit our result on the dynamics of these algorithms to derive the optimal choices for the matrix and iterate denoisers. We show that the resulting algorithm achieves the smallest possible asymptotic estimation error among a broad class of iterative algorithms under a fixed iteration budget.