Building a scalable and accurate copy detection mechanism
Explore this paper's citation graph
Summary
This paper study's the performance of various copy detection mechanisms, including the disk storage requirements, main memory requirements, response times for registration, and response time for querying, and contrast performance to the accuracy of the mechanisms (how well they detectpartial copies).
- Type
- article
- Published
- 1996-04-01
- Cited by
- 188
- References
- 16
- Access
- Open access
- OpenAlex
- https://openalex.org/W1997657677
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:13910952
Keywords
Computer science, Copying, Scalability, Web server, The Internet
References
- GLIMPSE: A Tool to Search Through Entire File Systems
- New Indices for Text: Pat Trees and Pat Arrays
- SCAM: A Copy Detection Mechanism for Digital Documents
- Duplicate Detection in Information Dissemination
- The State of Retrieval System Evaluation
- Suffix arrays: a new method for on-line string searches
- Term-Weighting Approaches in Automatic Text Retrieval
- Encryption and Secure Computer Networks
- Copy detection mechanisms for digital documents
- Copyright protection for electronic publishing over computer networks
- Plagiarism in the web
- Electronic marking and identification techniques to discourage document copying
- Document marking and identification using both line and word shifting
- Adaptive Sentence Boundary Disambiguation
Cited by
- Software Watermarking: Protective Terminology
- FastDocode: Finding Approximated Segments of N-Grams for Document Copy Detection - Lab Report for PAN at CLEF 2010
- Parallel Overlap and Similarity Detection in Semi- Structured Document Collections
- Document plagiarism detection algorithm using semantic networks
- Coderivative Document Recognition
- Parallel and Distributed Overlap Detection on the Web
- Opal: In Vivo Based Preservation Framework for Locating Lost Web Pages
- RLT-S: A Web System for Record Linkage
- Using fingerprints based on the sequence of a set of selected words in a document for 1-to-n similarity analysis
- Words and Intelligence II
- On retrieving intelligently plagiarized documents using semantic similarity
- Detecting Short Passages of Similar Text in Large Document Collections
- Creating A Corpus of Plagiarised Academic Texts
- A Copy detection Method for Malayalam Text Documents using N-grams Model
- Finding Near-Replicas of Documents and Servers on the Web
- Filtering with Approximate Predicates
- University of Sheffield - Lab Report for PAN at CLEF 2010
- Querying multiple document collections across the internet
- Diseño e Implementación de una Técnica para la Detección de Plagio en Documentos Digitales
- Report of the Ad Hoc Committee on Member Misconduct to the AIS Council
Related papers
- Performance Comparison of Web Cluster Load Balance Algorithms
- Research on web server cluster load balancing algorithm in web education system
- Agricultural information on the Internet: what is out there and how to find it
- Design and implementation of embedded Web server based on arm and Linux
- LSMAC and LSNAT: two approaches for cluster-based scalable Web servers
- Design and Research of the Web Server Based on ARM-Linux Embedded System
- Implementasi Distributed File Server pada Apache Web Server
- Scalability of a Web Server: How Does Vertical Scalability Improve the Performance of a Server?