Benchmarking and optimizing organism wide single-cell RNA alignment methods

Diaz-Mejia, Juan Javier, Williams, Elias, Focsa, Octavian, Mendonca, Dylan, Singh, Swechha, Innes, Brendan, Cooper, Sam

arXiv.org Artificial Intelligence 

Many methods have been proposed for removing batch effects and aligning single-cell RNA (scRNA) datasets. However, performance is typically evaluated based on multiple parameters and few datasets, creating challenges in assessing which method is best for aligning data at scale. Here, we introduce the K-Neighbors Intersection (KNI) score, a single score that both penalizes batch effects and measures accuracy at cross-dataset cell-type label prediction alongside carefully cu-rated small (scMARK) and large (scREF) benchmarks comprising 11 and 46 human scRNA studies respectively, where we have standardized author labels. Using the KNI score, we evaluate and optimize approaches for cross-dataset single-cell RNA integration. We introduce Batch Adversarial single-cell V ariational Inference (BA-scVI), as a new variant of scVI that uses adversarial training to penalize batch-effects in the encoder and decoder, and show this approach outperforms other methods. In the resulting aligned space, we find that the granularity of cell-type groupings is conserved, supporting the notion that whole-organism cell-type maps can be created by a single model without loss of information. To build comprehensive organism-wide and inter-species maps of cell types and states, we must build integrated transcriptional atlases that combine studies and patient populations at scale Regev et al. (2017). The now large number of published scRNA studies creates an opportunity for building a largely aligned scRNA atlas that would enable standardized reference-based analysis and cross-dataset comparison Lotfollahi et al. (2024). However, the challenge in combining data from disparate scRNA studies remains due to batch effects Lotfollahi et al. (2024); L ahnemann et al. (2020); Gavish et al. (2023). While studies have looked at the alignment of batches within datasets or between a handful of datasets focused on a specific tissue type, few have looked at model alignment across studies from different tissue types and instruments, as would be required for the generation of a reference atlas. Meanwhile, those studies that have used models to align datasets across tissue types and studies have used supervised models trained on cell-type labels, such as scBERT, Celltypist, SCimilarity, and SA TURN Y ang et al. (2022); Dom ınguez Conde et al. (2022); Heimberg et al. (2024); Rosen et al. (2024).

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found