Quantitative evaluation of approximate frequent pattern mining algorithms

Rohit Gupta; Gang Fang; Blayne Field; Michael Steinbach; Vipin Kumar

doi:10.1145/1401890.1401930

Quantitative evaluation of approximate frequent pattern mining algorithms

Rohit Gupta, Gang Fang, Blayne Field, Michael Steinbach, Vipin Kumar

Computer Science and Engineering

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution

49 Scopus citations

Abstract

Traditional association mining algorithms use a strict definition of support that requires every item in a frequent itemset to occur in each supporting transaction. In real-life datasets, this limits the recovery of frequent itemset patterns as they are fragmented due to random noise and other errors in the data. Hence, a number of methods have been proposed recently to discover approximate frequent itemsets in the presence of noise. These algorithms use a relaxed definition of support and additional parameters, such as row and column error thresholds to allow some degree of "error" in the discovered patterns. Though these algorithms have been shown to be successful in finding the approximate frequent itemsets, a systematic and quantitative approach to evaluate them has been lacking. In this paper, we propose a comprehensive evaluation framework to compare different approximate frequent pattern mining algorithms. The key idea is to select the optimal parameters for each algorithm on a given dataset and use the itemsets generated with these optimal parameters in order to compare different algorithms. We also propose simple variations of some of the existing algorithms by introducing an additional post-processing step. Subsequently, we have applied our proposed evaluation framework to a wide variety of synthetic datasets with varying amounts of noise and a real dataset to compare existing and our proposed variations of the approximate pattern mining algorithms. Source code and the datasets used in this study are made publicly available.

Original language	English (US)
Title of host publication	KDD 2008 - Proceedings of the 14th ACMKDD International Conference on Knowledge Discovery and Data Mining
Pages	301-309
Number of pages	9
DOIs	https://doi.org/10.1145/1401890.1401930
State	Published - 2008
Event	14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD 2008 - Las Vegas, NV, United States Duration: Aug 24 2008 → Aug 27 2008

Publication series

Name	Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining

Other

Other	14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD 2008
Country/Territory	United States
City	Las Vegas, NV
Period	8/24/08 → 8/27/08

Keywords

Approximate frequent itemsets
Association analysis
Error tolerance
Quantitative evaluation

Access

10.1145/1401890.1401930

OpenUrl availability

Full text

Cite this

Gupta, R., Fang, G., Field, B., Steinbach, M., & Kumar, V. (2008). Quantitative evaluation of approximate frequent pattern mining algorithms. In KDD 2008 - Proceedings of the 14th ACMKDD International Conference on Knowledge Discovery and Data Mining (pp. 301-309). (Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining). https://doi.org/10.1145/1401890.1401930

Quantitative evaluation of approximate frequent pattern mining algorithms. / Gupta, Rohit; Fang, Gang; Field, Blayne et al.
KDD 2008 - Proceedings of the 14th ACMKDD International Conference on Knowledge Discovery and Data Mining. 2008. p. 301-309 (Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining).

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution

Gupta, R, Fang, G, Field, B, Steinbach, M & Kumar, V 2008, Quantitative evaluation of approximate frequent pattern mining algorithms. in KDD 2008 - Proceedings of the 14th ACMKDD International Conference on Knowledge Discovery and Data Mining. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 301-309, 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD 2008, Las Vegas, NV, United States, 8/24/08. https://doi.org/10.1145/1401890.1401930

Gupta R, Fang G, Field B, Steinbach M , Kumar V. Quantitative evaluation of approximate frequent pattern mining algorithms. In KDD 2008 - Proceedings of the 14th ACMKDD International Conference on Knowledge Discovery and Data Mining. 2008. p. 301-309. (Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining). doi: 10.1145/1401890.1401930

@inproceedings{fc5591097d4d46d98b5d5d59c0cd6591,

title = "Quantitative evaluation of approximate frequent pattern mining algorithms",

abstract = "Traditional association mining algorithms use a strict definition of support that requires every item in a frequent itemset to occur in each supporting transaction. In real-life datasets, this limits the recovery of frequent itemset patterns as they are fragmented due to random noise and other errors in the data. Hence, a number of methods have been proposed recently to discover approximate frequent itemsets in the presence of noise. These algorithms use a relaxed definition of support and additional parameters, such as row and column error thresholds to allow some degree of {"}error{"} in the discovered patterns. Though these algorithms have been shown to be successful in finding the approximate frequent itemsets, a systematic and quantitative approach to evaluate them has been lacking. In this paper, we propose a comprehensive evaluation framework to compare different approximate frequent pattern mining algorithms. The key idea is to select the optimal parameters for each algorithm on a given dataset and use the itemsets generated with these optimal parameters in order to compare different algorithms. We also propose simple variations of some of the existing algorithms by introducing an additional post-processing step. Subsequently, we have applied our proposed evaluation framework to a wide variety of synthetic datasets with varying amounts of noise and a real dataset to compare existing and our proposed variations of the approximate pattern mining algorithms. Source code and the datasets used in this study are made publicly available.",

keywords = "Approximate frequent itemsets, Association analysis, Error tolerance, Quantitative evaluation",

author = "Rohit Gupta and Gang Fang and Blayne Field and Michael Steinbach and Vipin Kumar",

year = "2008",

doi = "10.1145/1401890.1401930",

language = "English (US)",

isbn = "9781605581934",

series = "Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining",

pages = "301--309",

booktitle = "KDD 2008 - Proceedings of the 14th ACMKDD International Conference on Knowledge Discovery and Data Mining",

note = "14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD 2008 ; Conference date: 24-08-2008 Through 27-08-2008",

}

TY - GEN

T1 - Quantitative evaluation of approximate frequent pattern mining algorithms

AU - Gupta, Rohit

AU - Fang, Gang

AU - Field, Blayne

AU - Steinbach, Michael

AU - Kumar, Vipin

PY - 2008

Y1 - 2008

N2 - Traditional association mining algorithms use a strict definition of support that requires every item in a frequent itemset to occur in each supporting transaction. In real-life datasets, this limits the recovery of frequent itemset patterns as they are fragmented due to random noise and other errors in the data. Hence, a number of methods have been proposed recently to discover approximate frequent itemsets in the presence of noise. These algorithms use a relaxed definition of support and additional parameters, such as row and column error thresholds to allow some degree of "error" in the discovered patterns. Though these algorithms have been shown to be successful in finding the approximate frequent itemsets, a systematic and quantitative approach to evaluate them has been lacking. In this paper, we propose a comprehensive evaluation framework to compare different approximate frequent pattern mining algorithms. The key idea is to select the optimal parameters for each algorithm on a given dataset and use the itemsets generated with these optimal parameters in order to compare different algorithms. We also propose simple variations of some of the existing algorithms by introducing an additional post-processing step. Subsequently, we have applied our proposed evaluation framework to a wide variety of synthetic datasets with varying amounts of noise and a real dataset to compare existing and our proposed variations of the approximate pattern mining algorithms. Source code and the datasets used in this study are made publicly available.

AB - Traditional association mining algorithms use a strict definition of support that requires every item in a frequent itemset to occur in each supporting transaction. In real-life datasets, this limits the recovery of frequent itemset patterns as they are fragmented due to random noise and other errors in the data. Hence, a number of methods have been proposed recently to discover approximate frequent itemsets in the presence of noise. These algorithms use a relaxed definition of support and additional parameters, such as row and column error thresholds to allow some degree of "error" in the discovered patterns. Though these algorithms have been shown to be successful in finding the approximate frequent itemsets, a systematic and quantitative approach to evaluate them has been lacking. In this paper, we propose a comprehensive evaluation framework to compare different approximate frequent pattern mining algorithms. The key idea is to select the optimal parameters for each algorithm on a given dataset and use the itemsets generated with these optimal parameters in order to compare different algorithms. We also propose simple variations of some of the existing algorithms by introducing an additional post-processing step. Subsequently, we have applied our proposed evaluation framework to a wide variety of synthetic datasets with varying amounts of noise and a real dataset to compare existing and our proposed variations of the approximate pattern mining algorithms. Source code and the datasets used in this study are made publicly available.

KW - Approximate frequent itemsets

KW - Association analysis

KW - Error tolerance

KW - Quantitative evaluation

UR - http://www.scopus.com/inward/record.url?scp=65449186692&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=65449186692&partnerID=8YFLogxK

U2 - 10.1145/1401890.1401930

DO - 10.1145/1401890.1401930

M3 - Conference contribution

AN - SCOPUS:65449186692

SN - 9781605581934

T3 - Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining

SP - 301

EP - 309

BT - KDD 2008 - Proceedings of the 14th ACMKDD International Conference on Knowledge Discovery and Data Mining

T2 - 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD 2008

Y2 - 24 August 2008 through 27 August 2008

ER -

Quantitative evaluation of approximate frequent pattern mining algorithms

Abstract

Publication series

Other

Keywords

Access

OpenUrl availability

Other files and links

Fingerprint

Cite this