On the bias of H-scores for comparing biclusters, and how to correct it

Di Iorio, Jacopo, Chiaromonte, Francesca, Cremona, Marzia A.

arXiv.org Machine Learning 

Cheng and Churchs algorithm has 2400 citations to date, 597 since 2015, and 179 in 2018-19 alone. It was the first to be applied to gene microarray data, and it is one of the main tools available in biclustering packages (e.g., the biclust R library) as well as in gene expression data analysis packages (e.g., IRIS-EDA, Monier et al. 2019). In addition, it is widely used as a benchmark: almost all published biclustering algorithms include a comparison with it. The role of the H-score in a biclustering algorithm is to allow validation and comparisons of biclusters, which may have different numbers of rows and columns. Our findings document a bias that can distort biclustering results. We prove, both analytically and by simulation, that the average H-score increases with the number of rows/columns in a bicluster - even in the ideal (and simplest) case of a single bicluster generated by an additive model plus a white noise. This biases the H-score, and hence all H-score based algorithms, towards small biclusters. Importantly, our analytical proof provides a straightforward way to correct this bias.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found