Open AccessProceedings Article
Improving Storage System Reliability with Proactive Error Prediction
Farzaneh Mahdisoltani,Ioan Stefanovici,Bianca Schroeder +2 more
- 16 May 2017
TL;DR: A range of different machine learning techniques are explored and it is shown that sector errors can be predicted ahead of time with high accuracy, even when only little training data or only training data for a different drive model is available.
read more
Abstract: This paper proposes using techniques from machine learning to make storage systems more reliable in the face of sector errors Sector errors are partial drive failures, where individual sectors on a drive become unavailable, and occur at a high rate in both hard disk drives and solid state drives The data in the affected sectors can only be recovered through external forms of redundancy (eg another drive in the same RAID), and be lost if the error is encountered while the system operates in degraded mode, eg during RAID reconstruction In this paper, we explore a range of different machine learning techniques and show that sector errors can be predicted ahead of time with high accuracy Prediction is robust, even when only little training data or only training data for a different drive model is available We also discuss a number of possible use cases for improving storage system reliability through the use of sector error predictors We evaluate one such use case in detail: We show that the mean time to detecting errors (and hence the window of vulnerability to data loss) can be greatly reduced by adapting the speed of a scrubber based on error predictions
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Anomaly Detection Model for Predicting Hard Disk Drive Failures
TL;DR: The electromechanical design of the HDD (Hard Disk Drive) renders it more susceptible to failures than other components of the computer system as discussed by the authors and the failure of HDD leads to permanent data loss.
12
DFPE: Explaining Predictive Models for Disk Failure Prediction
Yanwen Xie,Dan Feng,Fang Wang,Xuehai Tang,Jizhong Han,Xinyan Zhang +5 more
- 20 May 2019
TL;DR: A new explanation method DFPE designed for disk failure prediction is proposed to explain failure predictions made by a model and infer prediction rules learned by amodel to target and handle the hidden bias and overfitting, measures feature importances from a new perspective and enables intelligent failure handling.
10
Reliability Characterization and Failure Prediction of 3D TLC SSDs in Large-Scale Storage Systems
TL;DR: This paper investigates system-level 3D TLC SSDs to characterize reliability and sub-health status based on field Self-Monitoring, Analysis and Reporting Technology (SMART) data, and predict impending failure proactively and derives some findings for each selected attribute in predetermined categories.
7
A novel approach for predictive maintenance combining GAF encoding strategies and deep networks
Antonino Ferraro,Antonio Galli,Vincenzo Moscato,Giancarlo Sperlì +3 more
- 01 Dec 2020
TL;DR: In this article, the authors exploit GAF (Gramian Angular Field) encoding to obtain images from time series related to production systems, which can be used in pre-trained convolutional neural networks for better prediction performance and thus to create a more efficient and simple predictive maintenance scheme.
6
A Disk Failure Prediction Method Based on Active Semi-supervised Learning
Yang-Fan Zhou,Fang Wang,Dan Feng +2 more
TL;DR: ASLDP can overcome the problem of missing sample labels and data redundancy in large data centers, which are not considered and implemented in all offline learning methods for disk failure prediction to the best of the authors' knowledge.
4
References
An analysis of latent sector errors in disk drives
Lakshmi Narayanan Bairavasundaram,Garth R. Goodson,Shankar Pasupathy,Jiri Schindler +3 more
- 12 Jun 2007
TL;DR: This is the first study of such large scale the sample size is at least an order of magnitude larger than previously published studies and the first one to focus specifically on latent sector errors and their implications on the design and reliability of storage systems.
•Proceedings Article
Flash reliability in production: the expected and the unexpected
Bianca Schroeder,Raghav Lagisetty,Arif Merchant +2 more
- 22 Feb 2016
TL;DR: A large-scale field study covering many millions of drive days, ten different drive models, different flash technologies, and no evidence that higher-end SLC drives are more reliable than MLC drives within typical drive lifetimes is provided.
Improved disk-drive failure warnings
TL;DR: Improved methods are proposed for disk-drive failure prediction using the SMART internal drive attribute measurements in present drives, and the present warning-algorithm based on maximum error thresholds is replaced by distribution-free statistical hypothesis tests.
Understanding latent sector errors and how to protect against them
TL;DR: This article provides an extended statistical analysis of latent sector errors in the field, specifically from the view point of how to protect against LSEs, and includes schemes and policies that have been suggested before, but have never been evaluated on field data.
Hard Drive Failure Prediction Using Classification and Regression Trees
Jing Li,Xinpu Ji,Yuhan Jia,Bingpeng Zhu,Gang Wang,Zhongwei Li,Xiaoguang Liu +6 more
- 23 Jun 2014
TL;DR: A health degree model based on Regression Tree (RT) as well, which can give the drive a health assessment rather than a simple classification result and deal with warnings raised by the prediction model in order of their health degrees is proposed.
163