Towards a provenance-aware distributed filesystem
Explore this paper's citation graph
Summary
The preliminary results on a 32-node cluster show that FusionFS+SPADE is a promising prototype with negligible provenance overhead and has promise to scale to larger scales as FusionFS has been shown to scale.
- Type
- article
- Published
- 2013-10-11
- Cited by
- 22
- References
- 10
- OpenAlex
- https://openalex.org/W125125877
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:9310489
Keywords
Computer science, World Wide Web
References
- Provenance-Aware Storage Systems
- Ceph: a scalable, high-performance distributed file system
- The virtual data grid: a new model and architecture for data-intensive collaboration
- Foundations for provenance-aware systems
- Swift: Fast, Reliable, Loosely Coupled Parallel Computation
- ZHT: A Light-Weight Reliable Persistent Dynamic Scalable Zero-Hop Distributed Hash Table
- SPADE: Support for Provenance Auditing in Distributed Environments
Cited by
- Distributed data provenance for large-scale data-intensive computing
- Virtual chunks: On supporting random accesses to scientific data in compressible storage systems
- Research on big data information retrieval based on hadoop architecture
- FusionFS: Toward supporting data-intensive scientific applications on extreme-scale high-performance computing systems
- HyCache+: Towards Scalable High-Performance Caching Middleware for Parallel File Systems
- Dynamic Virtual Chunks: On Supporting Efficient Accesses to Compressed Scientific Data
- Towards Exploring Data-Intensive Scientific Applications at Extreme Scales through Systems and Simulations
- Towards cost-effective and high-performance caching middleware for distributed systems
- Weighted frequent multi partitioned itemset mining of market-basket data using MapReduce on YARN framework
- The future of scientific workflows
- BAASH: Enabling Blockchain-as-a-Service on High-Performance Computing Systems
- BAASH: Lightweight, Efficient, and Reliable Blockchain-As-A-Service for HPC Systems
- Improving the I / O Throughput for Data-Intensive Scientific Applications with Efficient Compression Mechanisms
- Exploring Data Compression in Distributed File Systems
- Supporting Large Scale Data-Intensive Computing with the FusionFS Distributed File System
- DISTRIBUTED NOSQL STORAGE FOR EXTREME-SCALE SYSTEM SERVICES IN CLOUDS AND SUPERCOMPUTERS
- BIG DATA SYSTEM INFRASTRUCTURE AT EXTREME SCALES
- Storage Support for Data-Intensive Scientific Applications on the Cloud
- A CONVERGENCE OF NOSQL STORAGE SYSTEMS FROM CLOUDS TO SUPERCOMPUTERS
- High-Performance Storage Support for Scientific Big Data Applications on the Cloud
Related papers
- The Hadoop Distributed File System
- ZHT: A Light-Weight Reliable Persistent Dynamic Scalable Zero-Hop Distributed Hash Table
- Towards high-performance and cost-effective distributed storage systems with information dispersal algorithms
- HyCache: A User-Level Caching Middleware for Distributed File Systems
- Exploring reliability of exascale systems through simulations
- Incremental Isometric Embedding of High-Dimensional Data Using Connected Neighborhood Graphs
- Swift: Fast, Reliable, Loosely Coupled Parallel Computation
- GPFS: A Shared-Disk File System for Large Computing Clusters
- Incremental Construction of Neighborhood Graphs for Nonlinear Dimensionality Reduction
- Distributed data provenance for large-scale data-intensive computing