Selective Inference for Latent Block Models

Watanabe, Chihiro, Suzuki, Taiji

arXiv.org Machine Learning 

A latent block model or an LBM [10, 7] has been widely used as a generative model of a relational data matrix, where the rows and columns represent different objects (e.g., customers and items), and its (i, j)-th element shows some relationship between objects i and j (e.g., how many times the customer i purchased item j). Until now, its effectiveness has been shown in various practical datasets, including customer-product transaction relationships [25] and gene expression data [24, 28]. In LBMs, we assume that there is an underlying block structure (i.e., a set of row and column cluster memberships) behind the observed data matrix and that each element of the matrix is generated independently from an identical distribution, given such a block structure. Particularly, a Gaussian LBM [22, 21] is useful to model a relational data matrix with real elements; this type of LBM is the focus of the current study. In a Gaussian LBM, we assume that each entry follows a Gaussian distribution, whose mean and variance are fixed constants in the same block (a formal description of Gaussian LBMs is given in Section 2.1). Besides estimating the block structure from a given observed data matrix based on an LBM, it is also important to test the validity of a model (i.e., the number of blocks) or an estimation result. Until now, several tests [2, 19, 12, 29, 27] have been proposed for determining the number of blocks in block models, such as a stochastic block model (SBM), which is a model for a square symmetric matrix (e.g., an adjacency matrix of the network structure). Among these studies, only [27]'s test can be applied to the LBM setting; however, its target is different from ours in that it is limited to the

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found