The penalized biclustering model and related algorithms

Thierry Chekouo, Alejandro Murua

Research output: Contribution to journalArticlepeer-review

8 Scopus citations

Abstract

Biclustering is the simultaneous clustering of two related dimensions, for example, of individuals and features, or genes and experimental conditions. Very few statistical models for biclustering have been proposed in the literature. Instead, most of the research has focused on algorithms to find biclusters. The models underlying them have not received much attention. Hence, very little is known about the adequacy and limitations of the models and the efficiency of the algorithms. In this work, we shed light on associated statistical models behind the algorithms. This allows us to generalize most of the known popular biclustering techniques, and to justify, and many times improve on, the algorithms used to find the biclusters. It turns out that most of the known techniques have a hidden Bayesian flavor. Therefore, we adopt a Bayesian framework to model biclustering. We propose a measure of biclustering complexity (number of biclusters and overlapping) through a penalized plaid model, and present a suitable version of the deviance information criterion to choose the number of biclusters, a problem that has not been adequately addressed yet. Our ideas are motivated by the analysis of gene expression data.

Original languageEnglish (US)
Pages (from-to)1255-1277
Number of pages23
JournalJournal of Applied Statistics
Volume42
Issue number6
DOIs
StatePublished - Jun 3 2015

Bibliographical note

Funding Information:
This research was supported by NSERC [grant number 327689-06].

Keywords

  • clustering
  • deviance information criterion
  • gene expression
  • mixture
  • model selection
  • plaid model

Fingerprint Dive into the research topics of 'The penalized biclustering model and related algorithms'. Together they form a unique fingerprint.

Cite this