Bayesian Co-clustering

Hanhuai Shan; Arindam Banerjee

doi:10.1109/ICDM.2008.91

Bayesian Co-clustering

Hanhuai Shan, Arindam Banerjee

Computer Science and Engineering

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution

146 Scopus citations

Abstract

In recent years, co-clustering has emerged as a powerful data mining tool that can analyze dyadic data connecting two entities. However, almost all existing co-clustering techniques are partitional, and allow individual rows and columns of a data matrix to belong to only one cluster. Several current applications, such as recommendation systems and market basket analysis, can substantially benefit from a mixed membership of rows and columns. In this paper, we present Bayesian co-clustering (BCC) models, that allow a mixed membership in row and column clusters. BCC maintains separate Dirichlet priors for rows and columns over the mixed membership and assumes each observation to be generated by an exponential family distribution corresponding to its row and column clusters. We propose a fast variational algorithm for inference and parameter estimation. The model is designed to naturally handle sparse matrices as the inference is done only based on the nonmissing entries. In addition to finding a co-cluster structure in observations, the model outputs a low dimensional coembedding, and accurately predicts missing values in the original matrix. We demonstrate the efficacy of the model through experiments on both simulated and real data.

Original language	English (US)
Title of host publication	Proceedings - 8th IEEE International Conference on Data Mining, ICDM 2008
Pages	530-539
Number of pages	10
DOIs	https://doi.org/10.1109/ICDM.2008.91
State	Published - Dec 1 2008
Event	8th IEEE International Conference on Data Mining, ICDM 2008 - Pisa, Italy Duration: Dec 15 2008 → Dec 19 2008

Publication series

Name	Proceedings - IEEE International Conference on Data Mining, ICDM
ISSN (Print)	1550-4786

Other

Other	8th IEEE International Conference on Data Mining, ICDM 2008
Country/Territory	Italy
City	Pisa
Period	12/15/08 → 12/19/08

Access

10.1109/ICDM.2008.91

OpenUrl availability

Full text

Cite this

@inproceedings{a9b5097c8cba4f50b8fb6369a1c40c2d,

title = "Bayesian Co-clustering",

abstract = "In recent years, co-clustering has emerged as a powerful data mining tool that can analyze dyadic data connecting two entities. However, almost all existing co-clustering techniques are partitional, and allow individual rows and columns of a data matrix to belong to only one cluster. Several current applications, such as recommendation systems and market basket analysis, can substantially benefit from a mixed membership of rows and columns. In this paper, we present Bayesian co-clustering (BCC) models, that allow a mixed membership in row and column clusters. BCC maintains separate Dirichlet priors for rows and columns over the mixed membership and assumes each observation to be generated by an exponential family distribution corresponding to its row and column clusters. We propose a fast variational algorithm for inference and parameter estimation. The model is designed to naturally handle sparse matrices as the inference is done only based on the nonmissing entries. In addition to finding a co-cluster structure in observations, the model outputs a low dimensional coembedding, and accurately predicts missing values in the original matrix. We demonstrate the efficacy of the model through experiments on both simulated and real data.",

author = "Hanhuai Shan and Arindam Banerjee",

year = "2008",

month = dec,

day = "1",

doi = "10.1109/ICDM.2008.91",

language = "English (US)",

isbn = "9780769535029",

series = "Proceedings - IEEE International Conference on Data Mining, ICDM",

pages = "530--539",

booktitle = "Proceedings - 8th IEEE International Conference on Data Mining, ICDM 2008",

note = "8th IEEE International Conference on Data Mining, ICDM 2008 ; Conference date: 15-12-2008 Through 19-12-2008",

}

TY - GEN

T1 - Bayesian Co-clustering

AU - Shan, Hanhuai

AU - Banerjee, Arindam

PY - 2008/12/1

Y1 - 2008/12/1

N2 - In recent years, co-clustering has emerged as a powerful data mining tool that can analyze dyadic data connecting two entities. However, almost all existing co-clustering techniques are partitional, and allow individual rows and columns of a data matrix to belong to only one cluster. Several current applications, such as recommendation systems and market basket analysis, can substantially benefit from a mixed membership of rows and columns. In this paper, we present Bayesian co-clustering (BCC) models, that allow a mixed membership in row and column clusters. BCC maintains separate Dirichlet priors for rows and columns over the mixed membership and assumes each observation to be generated by an exponential family distribution corresponding to its row and column clusters. We propose a fast variational algorithm for inference and parameter estimation. The model is designed to naturally handle sparse matrices as the inference is done only based on the nonmissing entries. In addition to finding a co-cluster structure in observations, the model outputs a low dimensional coembedding, and accurately predicts missing values in the original matrix. We demonstrate the efficacy of the model through experiments on both simulated and real data.

AB - In recent years, co-clustering has emerged as a powerful data mining tool that can analyze dyadic data connecting two entities. However, almost all existing co-clustering techniques are partitional, and allow individual rows and columns of a data matrix to belong to only one cluster. Several current applications, such as recommendation systems and market basket analysis, can substantially benefit from a mixed membership of rows and columns. In this paper, we present Bayesian co-clustering (BCC) models, that allow a mixed membership in row and column clusters. BCC maintains separate Dirichlet priors for rows and columns over the mixed membership and assumes each observation to be generated by an exponential family distribution corresponding to its row and column clusters. We propose a fast variational algorithm for inference and parameter estimation. The model is designed to naturally handle sparse matrices as the inference is done only based on the nonmissing entries. In addition to finding a co-cluster structure in observations, the model outputs a low dimensional coembedding, and accurately predicts missing values in the original matrix. We demonstrate the efficacy of the model through experiments on both simulated and real data.

UR - http://www.scopus.com/inward/record.url?scp=67049165560&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=67049165560&partnerID=8YFLogxK

U2 - 10.1109/ICDM.2008.91

DO - 10.1109/ICDM.2008.91

M3 - Conference contribution

AN - SCOPUS:67049165560

SN - 9780769535029

T3 - Proceedings - IEEE International Conference on Data Mining, ICDM

SP - 530

EP - 539

BT - Proceedings - 8th IEEE International Conference on Data Mining, ICDM 2008

T2 - 8th IEEE International Conference on Data Mining, ICDM 2008

Y2 - 15 December 2008 through 19 December 2008

ER -

Bayesian Co-clustering

Abstract

Publication series

Other

Access

OpenUrl availability

Other files and links

Fingerprint

Cite this