Duplicate Record Detection: A Survey
Explore this paper's citation graph
- Type
- article
- Published
- 2007-01-01
- Cited by
- 2,262
- References
- 111
- Access
- Open access
- OpenAlex
- https://openalex.org/W3123518987
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:386036
Keywords
Computer science, Data mining, Scalability, Information retrieval, Task (project management)
References
- The Data Warehouse ETL Toolkit: Practical Techniques for Extracting, Cleaning, Conforming, and Delivering Data
- The State of Record Linkage and Current Research Problems
- Using q-grams in a DBMS for Approximate String Processing.
- Methods for Linking and Mining Massive Heterogeneous Databases
- Learning String-Edit Distance
- The double metaphone search algorithm
- Bayesian Classification (AutoClass): Theory and Results
- Learning to Understand Information on the Internet: An Example-Based Approach
- An Efficient Domain-Independent Algorithm for Detecting Approximately Duplicate Database Records
- Making large scale SVM learning practical
- Advances in kernel methods: support vector learning
- Real-world Data is Dirty: Data Cleansing and The Merge/Purge Problem
- Entity matching in heterogeneous databases: a distance-based decision model
- Hierarchically Classifying Documents Using Very Few Words
- A Hierarchical Graphical Model for Record Linkage
- The inter-database instance identification problem in integrating autonomous systems
- Maximum Entropy Markov Models for Information Extraction and Segmentation
- Record linking: the design of efficient systems for linking records into individual and family histories.
- Network Flows: Theory, Algorithms, and Applications
- Automating the approximate record-matching process
Cited by
- Data Integration over Distributed and Heterogeneous Data Endpoints
- Toward an adaptive String Similarity Measure for Matching Product Offers
- Learning-based fusion for data deduplication: A robust and automated solution
- A Framework for Identity Resolution and Merging for Multi-source Information Extraction
- A Supervised Learning and Group Linking Method for Historical Census Household Linkage
- CONE: Metrics for Automatic Evaluation of Named Entity Co-Reference Resolution
- Enterprise Data Analysis and Visualization: An Interview Study
- A Comparison and Generalization of Blocking and Windowing Algorithms for Duplicate Detection
- Data Quality support to on-the-fly data integration using Adaptive Query Processing
- Automatic key discovery for Data Linking. (Découverte des clés pour le Liage de Données)
- Fusing automatically extracted annotations for the Semantic Web
- Web-based Affiliation Matching
- Visualization and Detection of Multiple Aliases in Social Media
- A Novel Framework and Model for Data Warehouse Cleansing
- Training selection for tuning entity matching
- Supporting Decision-making in Fraud Sensitive Environments: Including Personal Data from Public Sources in Risk Analyses
- Finding Ontological Correspondences for a Domain-Independent Natural Language Dialog Agent
- EARL: an Evolutionary Algorithm for Record Linkage
- Multi-source entity resolution
- Mining the Digital Information Networks
Related papers
- Duplicate Record Detection
- A Survey: Detection of Duplicate Record
- A survey on duplicate record detection in real world data
- Duplicate Detection Of Records in Queries using Clustering
- An Introduction to Duplicate Detection
- Detection of Duplicate records by using Progressive Windowing Technique
- Efficient partial-duplicate detection based on sequence matching
- Duplicate Record Detection for Database Cleansing