Sparse Indexing: Large Scale, Inline Deduplication Using Sampling and Locality
Explore this paper's citation graph
Summary
Sparse indexing, a technique that uses sampling and exploits the inherent locality within backup streams to solve for large-scale backup the chunk-lookup disk bottleneck problem that inline, chunk-based deduplication schemes face, is presented.
- Type
- article
- Published
- 2009-02-24
- Cited by
- 472
- References
- 30
- OpenAlex
- https://openalex.org/W69510097
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:7858129
Keywords
Computer science, Data deduplication, Bottleneck, Search engine indexing, Backup
References
- Bigtable: a distributed storage system for structured data
- Evaluation of Efficient Archival Storage Techniques
- Avoiding the Disk Bottleneck in the Data Domain Deduplication File System
- Microsoft Exchange Server
- Alternatives for Detecting Redundancy in Storage Systems Data
- Jumbo Store: Providing Efficient Incremental Upload and Versioning for a Utility Rendering Service
- An evaluation of buffer management strategies for relational database systems
- Finding Similar Files in a Large File System
- Bigtable: A Distributed Storage System for Structured Data
- Copy detection mechanisms for digital documents
- TAPER: tiered approach for eliminating redundancy in replica synchronization
- A low-bandwidth network file system
- A universal algorithm for sequential data compression
- Finding similar files in large document repositories
- Space/time trade-offs in hash coding with allowable errors
- On the resemblance and containment of documents
- Opportunistic Use of Content Addressable Storage for Distributed File Systems
- Compactly encoding unstructured inputs with differential compression
- Fast, Inexpensive Content-Addressed Storage in Foundation
- Deep Store: an archival storage system architecture
Cited by
- Tradeoffs in Scalable Data Routing for Deduplication Clusters
- Characteristics of backup workloads in production systems
- Low-Cost Data Deduplication for Virtual Machine Backup in Cloud Storage
- Statistical characterization of storage system workloads for data deduplication and load placement in heterogeneous storage environments
- Improving restore speed for backup systems that use inline chunk-based deduplication
- CAFTL: A Content-Aware Flash Translation Layer Enhancing the Lifespan of Flash Memory based Solid State Drives
- Using multi-threads to hide deduplication I/O latency with low synchronization overhead
- Optimization Techniques To Record Deduplication
- iDedup: latency-aware, inline data deduplication for primary storage
- Building a High-performance Deduplication System
- Leveraging content properties to optimize distributed storage systems. (Exploitation du contenu pour l'optimisation du stockage distribué)
- Research on a new peer to cloud and peer model and a deduplication storage system
- Incremental parallel and distributed systems
- Courbes remplissant l'espace et leur application en traitement d'images. (Spacer-filling curves and their application in image processing)
- Memory efficient sanitization of a deduplicated storage system
- Design Tradeoffs for Data Deduplication Performance in Backup Workloads
- File recipe compression in data deduplication systems
- Primary Data Deduplication - Large Scale Study and System Design
- Improving caches in consolidated environments
- ProSy: A similarity based inline deduplication system for primary storage
Related papers
- Tradeoffs in Scalable Data Routing for Deduplication Clusters
- HYDRAstor: A Scalable Secondary Storage
- SiLo: A Similarity-Locality based Near-Exact Deduplication Scheme with Low RAM Overhead and High Throughput
- Characteristics of backup workloads in production systems
- dedupv1: Improving deduplication throughput using solid state drives (SSD)
- Fast, Inexpensive Content-Addressed Storage in Foundation
- Redundancy elimination within large collections of files
- Space/time trade-offs in hash coding with allowable errors
- A study of practical deduplication
- A low-bandwidth network file system