Failure data analysis of a large-scale heterogeneous server environment
Explore this paper's citation graph
- Type
- article
- Published
- 2004-06-28
- Cited by
- 247
- References
- 21
- Access
- Open access
- OpenAlex
- https://openalex.org/W2160821994
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:10036202
Keywords
Computer science, Robustness (evolution), Workload, Server, Software deployment
References
- Failure analysis and modeling of a VAXcluster system
- A First Course on Stochastic Processes
- Improving cluster availability using workstation validation
- On the Reliability of the IBM MVS/XA Operating System
- Software fault tolerance in a clustered architecture: techniques and reliability modeling
- Analysis and implementation of software rejuvenation in cluster systems
- Terrestrial cosmic rays
- Automatic Statistical Analysis of Bivariate Nonstationary Time Series
- Measurement and modeling of computer reliability as affected by system activity
- Effect of System Workload on Operating System Reliability: A Study on IBM 3081
- Networked Windows NT system field failure data analysis
- Software Dependability in the Tandem GUARDIAN System
- Software defects and their impact on system availability-a study of field failures in operating systems
- Analysis of workload influence on dependability
- Modeling the effect of technology trends on the soft error rate of combinational logic
- A FIRST COURSE IN STOCHASTIC PROCESSES
- Impact of Correlated Failures on Dependability in a VAXcluster System
Cited by
- Using Byzantine Fault-Tolerance to Improve Dependability in Federated Cloud Computing
- Understanding and Improving the Performance Consistency of Distributed Computing Systems
- Propitious Checkpoint Intervals to Improve System Performance
- Failure analysis, modeling, and prediction for BlueGene/L
- An Online Controller Towards Self-Adaptive File System Availability and Performance
- Data Mining for Autonomic System Management: A Case Study at FIU-SCIS
- Fault tolerance: Validating a mathematical model via a case study of RAxML, an HPC community code
- A Large-scale Study of Failures in High-performance-computing Systems (CMU-PDL-05-112)
- Autonomic Failure Identification and Diagnosis for Building Dependable Cloud Computing Systems
- An Analysis of Traces from a Production MapReduce Cluster
- Multifaceted resource management on virtualized providers
- Fault Diagnosis in Enterprise Software Systems Using Discrete Monitoring Data
- A Framework for the Study of Grid Inter-Operation Mechanisms
- Opportunistic Checkpoint Intervals to Improve System Performance
- On-line Detection of Anomalies in Mission-critical Software Systems
- Holistic cloud computing environmental quantification and behavioural analysis
- Strategies for achieving dependability in parallel file systems
- Dependability evaluation of mobile distributed systems via Field Failure Data Analysis
- Design of a 50° Field-of-View Object for Head-Mounted Projective Displays and Investigation of Retro-Reflective Materials
- Measuring and Understanding Extreme-Scale Application Resilience: A Field Study of 5,000,000 HPC Application Runs
Related papers
- Fault-Tolerance in the Scope of Software-Defined Networking (SDN)
- Dependable computing depends on structured fault tolerance
- Towards an Integrated Approach to Fault Tolerance in Delta-4
- Incorporating fault tolerance tactics in software architecture patterns
- Progress in real-time fault tolerance
- Application-level fault tolerance in real-time embedded systems
- A Classification-Based Approach to Fault-Tolerance Support in Parallel Programs
- Fault-tolerance in process control : possibilities, limitations and trends