Assessing data mining results via swap randomization
Explore this paper's citation graph
Summary
For some datasets the structure discovered by the data mining algorithms is expected, given the row and column margins of the datasets, while for other datasets the discovered structure conveys information that is not captured by the margin counts.
- Type
- article
- Published
- 2007-12-01
- Cited by
- 287
- References
- 40
- OpenAlex
- https://openalex.org/W1978036582
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:52305658
Keywords
Computer science, Data mining, Margin (machine learning), Cluster analysis, Swap (finance)
References
- Discovering Predictive Association Rules
- On Inverse Frequent Set Mining
- Generalized Monte Carlo significance tests
- Using association rules for product assortment decisions: a case study
- Sequential Monte Carlo p-values
- Approximate counting by dynamic programming
- Computational complexity of itemset frequency satisfiability
- Identifying non-actionable association rules
- Testing Ecological Patterns
- On the precise number of (0, 1)-matrices in U(R, R)
- Sequential Monte Carlo Methods for Statistical Analysis of Tables
- Sampling binary contingency tables with a greedy start
- Discovering significant rules
- Pruning and summarizing the discovered associations
- An Application of Markov Chain Monte Carlo to Community Ecology
- Selecting the right interestingness measure for association patterns
- Empirical bayes screening for multi-item associations
- Efficient sampling algorithm for estimating subgraph concentrations and detecting network motifs
- Exploiting a support-based upper bound of Pearson's correlation coefficient for efficiently identifying strongly correlated pairs
- Monte Carlo Sampling Methods Using Markov Chains and Their Applications
Cited by
- Measuring structural similarity in large online networks.
- Evaluating Query Result Significance in Databases via Randomizations
- Controlling False Positives in Frequent Itemsets Mining through the VC-Dimension
- Graph Generation with Prescribed Feature Constraints
- Low-Entropy Set Selection
- Finding Subgroups having Several Descriptions: Algorithms for Redescription Mining
- Randomization of real-valued matrices for assessing the significance of data mining results
- Data Mining and Analysis: Fundamental Concepts and Algorithms
- Achieving privacy-preserving distributed statistical computation
- Different flavors of randomness: comparing random graph models with fixed degree sequences
- Randomization Techniques for Graphs
- Discovering combinatorial disease biomarkers
- A people-to-people matching system using graph mining techniques
- Matrix Decomposition Methods for Data Mining : Computational Complexity and Algorithms
- A swap randomization approach for mining motion field time series over the Argentiere glacier
- Characterizing Discriminative Patterns
- Efficient search for statistically significant dependency rules in binary data
- Interactive Data Mining Considered Harmful (If Done Wrong)
- Query Significance in Databases via Randomizations
- Apriori-based algorithms for km-anonymizing trajectory data
Related papers
- Remarks on Algorithm 2, Algorithm 3, Algorithm 15, Algorithm 25 and Algorithm 26
- Remarks on Algorithm 332: Jacobi polynomials: Algorithm 344: student's t-distribution: Algorithm 351: modified Romberg quadrature: Algorithm 359: factoral analysis of variance
- Intensive Margin and Extensive Margin Adjustments of Labor Market : Turkey versus United States
- On the Necessity and Feasibility of Constructing Big Water Margin Cultural System——Taking Big Folk Water Margin by Fan Chaoyang and Zhang Qingjian as an Example
- The legal aspects of swaps : an analysis based on economic substance
- Large margin distribution machine
- Improving generalization of deep neural networks by leveraging margin distribution