scalenet
ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
Hao, Zhiwei, Guo, Jianyuan, Shen, Li, Han, Kai, Tang, Yehui, Hu, Han, Wang, Yunhe
Recent advancements in vision transformers (ViTs) have demonstrated that larger models often achieve superior performance. However, training these models remains computationally intensive and costly. To address this challenge, we introduce ScaleNet, an efficient approach for scaling ViT models. Unlike conventional training from scratch, ScaleNet facilitates rapid model expansion with negligible increases in parameters, building on existing pretrained models. This offers a cost-effective solution for scaling up ViTs. Specifically, ScaleNet achieves model expansion by inserting additional layers into pretrained ViTs, utilizing layer-wise weight sharing to maintain parameters efficiency. Each added layer shares its parameter tensor with a corresponding layer from the pretrained model. To mitigate potential performance degradation due to shared weights, ScaleNet introduces a small set of adjustment parameters for each layer. These adjustment parameters are implemented through parallel adapter modules, ensuring that each instance of the shared parameter tensor remains distinct and optimized for its specific function. Experiments on the ImageNet-1K dataset demonstrate that ScaleNet enables efficient expansion of ViT models. With a 2$\times$ depth-scaled DeiT-Base model, ScaleNet achieves a 7.42% accuracy improvement over training from scratch while requiring only one-third of the training epochs, highlighting its efficiency in scaling ViTs. Beyond image classification, our method shows significant potential for application in downstream vision areas, as evidenced by the validation in object detection task.
ScaleNet: Scale Invariance Learning in Directed Graphs
Jiang, Qin, Wang, Chengjia, Lones, Michael, Yuan, Yingfang, Pang, Wei
Graph Neural Networks (GNNs) have advanced relational data analysis but lack invariance learning techniques common in image classification. In node classification with GNNs, it is actually the ego-graph of the center node that is classified. This research extends the scale invariance concept to node classification by drawing an analogy to image processing: just as scale invariance being used in image classification to capture multi-scale features, we propose the concept of ``scaled ego-graphs''. Scaled ego-graphs generalize traditional ego-graphs by replacing undirected single-edges with ``scaled-edges'', which are ordered sequences of multiple directed edges. We empirically assess the performance of the proposed scale invariance in graphs on seven benchmark datasets, across both homophilic and heterophilic structures. Our scale-invariance-based graph learning outperforms inception models derived from random walks by being simpler, faster, and more accurate. The scale invariance explains inception models' success on homophilic graphs and limitations on heterophilic graphs. To ensure applicability of inception model to heterophilic graphs as well, we further present ScaleNet, an architecture that leverages multi-scaled features. ScaleNet achieves state-of-the-art results on five out of seven datasets (four homophilic and one heterophilic) and matches top performance on the remaining two, demonstrating its excellent applicability. This represents a significant advance in graph learning, offering a unified framework that enhances node classification across various graph types. Our code is available at https://github.com/Qin87/ScaleNet/tree/July25.
Non-invasive Localization of the Ventricular Excitation Origin Without Patient-specific Geometries Using Deep Learning
Pilia, Nicolas, Schuler, Steffen, Rees, Maike, Moik, Gerald, Potyagaylo, Danila, Dössel, Olaf, Loewe, Axel
Ventricular tachycardia (VT) can be one cause of sudden cardiac death affecting 4.25 million persons per year worldwide. A curative treatment is catheter ablation in order to inactivate the abnormally triggering regions. To facilitate and expedite the localization during the ablation procedure, we present two novel localization techniques based on convolutional neural networks (CNNs). In contrast to existing methods, e.g. using ECG imaging, our approaches were designed to be independent of the patient-specific geometries and directly applicable to surface ECG signals, while also delivering a binary transmural position. One method outputs ranked alternative solutions. Results can be visualized either on a generic or patient geometry. The CNNs were trained on a data set containing only simulated data and evaluated both on simulated and clinical test data. On simulated data, the median test error was below 3mm. The median localization error on the clinical data was as low as 32mm. The transmural position was correctly detected in up to 82% of all clinical cases. Using the ranked alternative solutions, the top-3 median error dropped to 20mm on clinical data. These results demonstrate a proof of principle to utilize CNNs to localize the activation source without the intrinsic need of patient-specific geometrical information. Furthermore, delivering multiple solutions can help the physician to find the real activation source amongst more than one possible locations. With further optimization, these methods have a high potential to speed up clinical interventions. Consequently they could decrease procedural risk and improve VT patients' outcomes.
ScaleNet: Searching for the Model to Scale
Xie, Jiyang, Su, Xiu, You, Shan, Ma, Zhanyu, Wang, Fei, Qian, Chen
Recently, community has paid increasing attention on model scaling and contributed to developing a model family with a wide spectrum of scales. Current methods either simply resort to a one-shot NAS manner to construct a non-structural and non-scalable model family or rely on a manual yet fixed scaling strategy to scale an unnecessarily best base model. In this paper, we bridge both two components and propose ScaleNet to jointly search base model and scaling strategy so that the scaled large model can have more promising performance. Concretely, we design a super-supernet to embody models with different spectrum of sizes (e.g., FLOPs). Then, the scaling strategy can be learned interactively with the base model via a Markov chain-based evolution algorithm and generalized to develop even larger models. To obtain a decent super-supernet, we design a hierarchical sampling strategy to enhance its training sufficiency and alleviate the disturbance. Experimental results show our scaled networks enjoy significant performance superiority on various FLOPs, but with at least 2.53x reduction on search cost. Codes are available at https://github.com/luminolx/ScaleNet.