Clustering very large data sets with principal direction divisive partitioning

D. Littau; D. Boley

doi:10.1007/3-540-28349-8_4

Clustering very large data sets with principal direction divisive partitioning

D. Littau, D. Boley

Computer Science and Engineering

Research output: Chapter in Book/Report/Conference proceeding › Chapter

11 Scopus citations

Abstract

We present a method to cluster data sets too large to fit in memory, based on a Low-Memory Factored Representation (LMFR). The LMFR represents the original data in a factored form with much less memory, while preserving the individuality of each of the original samples. The scalable clustering algorithm Principal Direction Divisive Partitioning (PDDP) can use the factored form in a natural way to obtain a clustering of the original dataset. The resulting algorithm is the PieceMeal PDDP (PMPDDP) method. The scalability of PMPDDP is demonstrated with a complexity analysis and experimental results. A discussion on the practical use of this method by a casual user is provided.

Original language	English (US)
Title of host publication	Grouping Multidimensional Data
Subtitle of host publication	Recent Advances in Clustering
Publisher	Springer Berlin Heidelberg
Pages	99-126
Number of pages	28
ISBN (Print)	354028348X, 9783540283485
DOIs	https://doi.org/10.1007/3-540-28349-8_4
State	Published - 2006

Access

10.1007/3-540-28349-8_4

OpenUrl availability

Full text

Cite this

@inbook{c6c2e3229b2a4c0a9d4d51adc65789e6,

title = "Clustering very large data sets with principal direction divisive partitioning",

abstract = "We present a method to cluster data sets too large to fit in memory, based on a Low-Memory Factored Representation (LMFR). The LMFR represents the original data in a factored form with much less memory, while preserving the individuality of each of the original samples. The scalable clustering algorithm Principal Direction Divisive Partitioning (PDDP) can use the factored form in a natural way to obtain a clustering of the original dataset. The resulting algorithm is the PieceMeal PDDP (PMPDDP) method. The scalability of PMPDDP is demonstrated with a complexity analysis and experimental results. A discussion on the practical use of this method by a casual user is provided.",

author = "D. Littau and D. Boley",

year = "2006",

doi = "10.1007/3-540-28349-8_4",

language = "English (US)",

isbn = "354028348X",

pages = "99--126",

booktitle = "Grouping Multidimensional Data",

publisher = "Springer Berlin Heidelberg",

}

TY - CHAP

T1 - Clustering very large data sets with principal direction divisive partitioning

AU - Littau, D.

AU - Boley, D.

PY - 2006

Y1 - 2006

N2 - We present a method to cluster data sets too large to fit in memory, based on a Low-Memory Factored Representation (LMFR). The LMFR represents the original data in a factored form with much less memory, while preserving the individuality of each of the original samples. The scalable clustering algorithm Principal Direction Divisive Partitioning (PDDP) can use the factored form in a natural way to obtain a clustering of the original dataset. The resulting algorithm is the PieceMeal PDDP (PMPDDP) method. The scalability of PMPDDP is demonstrated with a complexity analysis and experimental results. A discussion on the practical use of this method by a casual user is provided.

AB - We present a method to cluster data sets too large to fit in memory, based on a Low-Memory Factored Representation (LMFR). The LMFR represents the original data in a factored form with much less memory, while preserving the individuality of each of the original samples. The scalable clustering algorithm Principal Direction Divisive Partitioning (PDDP) can use the factored form in a natural way to obtain a clustering of the original dataset. The resulting algorithm is the PieceMeal PDDP (PMPDDP) method. The scalability of PMPDDP is demonstrated with a complexity analysis and experimental results. A discussion on the practical use of this method by a casual user is provided.

UR - http://www.scopus.com/inward/record.url?scp=34748835594&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=34748835594&partnerID=8YFLogxK

U2 - 10.1007/3-540-28349-8_4

DO - 10.1007/3-540-28349-8_4

M3 - Chapter

AN - SCOPUS:34748835594

SN - 354028348X

SN - 9783540283485

SP - 99

EP - 126

BT - Grouping Multidimensional Data

PB - Springer Berlin Heidelberg

ER -

Clustering very large data sets with principal direction divisive partitioning

Abstract

Access

OpenUrl availability

Other files and links

Fingerprint

Cite this