Extraction of Template using Clustering from Heterogeneous Web Documents
Explore this paper's citation graph
Summary
This paper focuses on extracting templates from heterogeneous web pages by examining the problem during automatic database value extraction from different web pages, which is done without any human data input.
- Type
- article
- Published
- 2015-06-18
- Cited by
- 5
- References
- 14
- Access
- Open access
- OpenAlex
- https://openalex.org/W1950708444
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:1502513
Keywords
Computer science, Cluster analysis, Information retrieval, Extraction (chemistry), Data mining
References
- Selectively estimation for Boolean queries
- Classification Using Streaming Random Forests
- Min-wise independent permutations (extended abstract)
- Clustering Web pages based on their structure
- Page-level template detection via isotonic smoothing
- RoadRunner: Towards Automatic Data Extraction from Large Web Sites
- The volume and evolution of web page templates
- Co-clustering by block value decomposition
- TEXT: Automatic Template Extraction from Heterogeneous Web Pages
- CRD: fast co-clustering on large datasets utilizing sampling-based matrix decomposition
- Min-Wise Independent Permutations
- Information-theoretic co-clustering
- Extracting Structured Data from Web Pages
Cited by
- Identification of Bahasa Indonesia official computer terms in Indonesian government websites
- Interactive Visualization of Template Graph for Daily Clinical Notes
- What Web Template Extractor Should I Use? A Benchmarking and Comparison for Five Template Extractors
- Document clustering for knowledge synthesis and project portfolio funding decision in R&D organizations
- Optimized Template Detection and Extraction Algorithm for Web Scraping of Dynamic Web Pages
Related papers
- Clustering - What Both Theoreticians and Practitioners Are Doing Wrong
- Classifying infants in the Strange Situation with three-way mixture method of clustering.
- Data clustering: application and trends
- A Methodology for Clustering Transient Biomedical Signals by Variable
- Improved accelerating large data K-means clustering algorithm
- Using version information in architectural clustering - a case study
- The Influence on Clustering Results of Electricity Load Curves Using Different Distances
- K-means clustering with multiresolution peak detection
- Improving a Centroid-Based Clustering by Using Suitable Centroids from Another Clustering