Diagnosability Security and Safety of Hybrid Dynamic and Cyber Physical Systems 1st edition by Moamar Sayed Mouchaweh ISBN 3319749617 978-3319749617 pdf download https://ebookball.com/product/diagnosability-security-and-safetyof-hybrid-dynamic-and-cyber-physical-systems-1st-edition-bymoamar-sayed-mouchaweh-isbn-3319749617-978-3319749617-16676/
Explore and download more ebooks or textbooks at ebookball.com
Get Your Digital Files Instantly: PDF, ePub, MOBI and More Quick Digital Downloads: PDF, ePub, MOBI and Other Formats
Diagnosability Security and Safety of Hybrid Dynamic and Cyber Physical Systems 1st ediiton by Moamar Sayed Mouchaweh ISBN 3319749617 ‎ 978-3319749617 https://ebookball.com/product/diagnosability-security-and-safetyof-hybrid-dynamic-and-cyber-physical-systems-1st-ediiton-bymoamar-sayed-mouchawehisbn-3319749617-aeurz-978-3319749617-16668/
Safety and Security of Cyber Physical Systems Engineering dependable Software using Principle based Development 1st Edition by Frank Furrer ISBN 9783658371821 365837182X https://ebookball.com/product/safety-and-security-of-cyberphysical-systems-engineering-dependable-software-using-principlebased-development-1st-edition-by-frank-furrerisbn-9783658371821-365837182x-20088/
Security and Resilience in Cyber Physical Systems 1st edition by Masoud Abbaszadeh 303097166X 9783030971663
https://ebookball.com/product/security-and-resilience-in-cyberphysical-systems-1st-edition-by-masoudabbaszadeh-303097166x-9783030971663-20040/
Cyber Physical Systems Architecture Security and Application 1st edition by Song Guo, Deze Zeng ISBN 3030064611 9783030064617
https://ebookball.com/product/cyber-physical-systemsarchitecture-security-and-application-1st-edition-by-song-guodeze-zeng-isbn-3030064611-9783030064617-17050/
Cyber Physical Systems Architecture Security and Application 1st Edition by Song Guo, Deze Zeng ISBN 3319925644 9783319925646
https://ebookball.com/product/cyber-physical-systemsarchitecture-security-and-application-1st-edition-by-song-guodeze-zeng-isbn-3319925644-9783319925646-15834/
Security of Cyber Physical Systems Vulnerability and Impact 1st Edition by Hadis Karimipour, Pirathayini Srikantha, Hany Farag ISBN 9783030455415 3030455416 https://ebookball.com/product/security-of-cyber-physical-systemsvulnerability-and-impact-1st-edition-by-hadis-karimipourpirathayini-srikantha-hany-faragisbn-9783030455415-3030455416-20086/
Security and Resilience of Cyber Physical Systems 1st edition by Krishan Kumar, Sunny Behal, Abhinav Bhandari, Sajal Bhatia ISBN 1032028637 9781032028637 https://ebookball.com/product/security-and-resilience-of-cyberphysical-systems-1st-edition-by-krishan-kumar-sunny-behalabhinav-bhandari-sajal-bhatiaisbn-1032028637-9781032028637-16628/
Security Engineering for Embedded and Cyber Physical Systems 1st edition by Saad Motahhir, Yassine Maleh 1000644235 9781000644234
https://ebookball.com/product/security-engineering-for-embeddedand-cyber-physical-systems-1st-edition-by-saad-motahhir-yassinemaleh-1000644235-9781000644234-20190/
Cyber Physical Systems Security 1st edition by Çetin Kaya Koç 9783319989358 3319989359
https://ebookball.com/product/cyber-physical-systemssecurity-1st-edition-by-a-etin-kayakoass-9783319989358-3319989359-16782/
Moamar Sayed-Mouchaweh Editor
Diagnosability, Security and Safety of Hybrid Dynamic and CyberPhysical Systems
Diagnosability, Security and Safety of Hybrid Dynamic and Cyber-Physical Systems
Moamar Sayed-Mouchaweh Editor
Diagnosability, Security and Safety of Hybrid Dynamic and Cyber-Physical Systems
123
Editor Moamar Sayed-Mouchaweh Institute Mines-Telecom Lille Douai Douai, France
ISBN 978-3-319-74961-7 ISBN 978-3-319-74962-4 (eBook) https://doi.org/10.1007/978-3-319-74962-4 Library of Congress Control Number: 2018934979 © Springer International Publishing AG 2018 This work is subject to copyright. All rights are reserved by the Publisher, whether the whole or part of the material is concerned, specifically the rights of translation, reprinting, reuse of illustrations, recitation, broadcasting, reproduction on microfilms or in any other physical way, and transmission or information storage and retrieval, electronic adaptation, computer software, or by similar or dissimilar methodology now known or hereafter developed. The use of general descriptive names, registered names, trademarks, service marks, etc. in this publication does not imply, even in the absence of a specific statement, that such names are exempt from the relevant protective laws and regulations and therefore free for general use. The publisher, the authors and the editors are safe to assume that the advice and information in this book are believed to be true and accurate at the date of publication. Neither the publisher nor the authors or the editors give a warranty, express or implied, with respect to the material contained herein or for any errors or omissions that may have been made. The publisher remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Printed on acid-free paper This Springer imprint is published by the registered company Springer International Publishing AG part of Springer Nature. The registered company address is: Gewerbestrasse 11, 6330 Cham, Switzerland
Preface
Cyber-physical systems (CPS) are characterized as a combination of physical (physical plant, process, network) and cyber (software, algorithm, computation) components whose operations are monitored, controlled, coordinated, and integrated by a computing and communicating core. The interaction between physical and computational components in CPS is intensive. They cover an increasing number of real life applications such as autonomous vehicles, aircrafts, smart manufacturing processes, surgical robots and human robot collaboration, smart electric grids, home appliances, air traffic control, automated farming, and implanted medical devices. The interaction between both physical and cyber components requires tools allowing analyzing and modeling both the discrete (discrete event control, communication protocols, discrete sensors/actuators, scheduling algorithms, etc.) and continuous (continuous dynamics, physics, continuous sensors/actuators, etc.) dynamics. Therefore, many CPS can be modeled as hybrid dynamic systems in order to take into account both discrete and continuous behaviors as well as the interactions between them. Many critical infrastructures, such as power generation and distribution networks, water networks and mass transportation systems, autonomous vehicles and traffic monitoring, are CPS. Such systems, including critical infrastructures, are becoming widely used and covering many aspects of our daily life. Therefore, the security, safety, and reliability of CPS is essential for the success of their implementation and operation. However, these systems are prone to major incidents resulting from cyberattacks and system failures. These incidents affect significantly their security and safety. Attacks can be represented as an interference in the communication channel between the supervisor and the system intentionally generated by intruders in order to damage the system or as the enablement, respectively disablement, of actuators events that are disabled, respectively enabled, by the supervisor. In general, intruders hide, create, or even change intentionally events that transit from one device (actuator, sensor) to another in a control communication
v
vi
Preface
channel. These attacks in a supervisory control system can lead the plant to execute event sequences entailing the system to reach unsafe or dangerous states that can damage the system. Therefore, reliable, scalable, and timely fault diagnosis is crucial in order to improve the robustness of CPS to failures. In addition, it is primordial to detect intrusions that exploit the vulnerabilities of industrial control systems in order to alter intentionally the integrity, confidentiality, and availability of CPS. These cyberattacks affect the control commands of the controller [Programmable Logic Controller (PLC)], the reports (sensors readings) coming from the plant as well as the communication between them. Moreover, a thorough understanding of the vulnerability of CPS’ components against such incidents can be incorporated in future design processes in order to better design such systems. Finally, the timely fault diagnosis can help operators to have better situation awareness and give them ample time to implement correction (maintenance) actions. However, guaranteeing the security and safety of CPS requires verifying their behavioral or safety properties either at design stage such as state reachability, diagnosability, and predictability or online such as fault detection and isolation. This is a challenging task because of the inherent interconnected and heterogeneous combination of behaviors (cyber/physical, discrete/continuous) in these systems. Indeed, fault propagation in CPS is governed not only by the behaviors of components in the physical and cyber subsystems but also by their interactions. This makes the identification of the root cause of observed anomalies and predicting the failure events a hard problem. Moreover, the increasing complexity of CPS and the security and safety requirements of their operation as well as their decentralized resource management entail a significant increase in the likelihood of failures in these systems. Finally, it is worth mentioning that computing the reachable set of states of HDS is an undecidable matter due to the infinite state space of continuous systems. This edited Springer book presents recent and advanced approaches and techniques that address the complex problem of analyzing the diagnosability property of CPS and ensuring their security and safety against faults and attacks. The CPS are modeled as hybrid dynamic systems using different model-based and data-driven approaches in different application domains (electric transmission networks, wireless communication networks, intrusions in industrial control systems, intrusions in production systems, wind farms, etc.). These approaches handle the problem of ensuring the security of CPS in presence of attacks and verifying their diagnosability in presence of different kinds of uncertainty (uncertainty related to the event occurrences, to their order of occurrence, to their value etc.). Finally, the editor is very grateful to all authors and reviewers for their very valuable contribution allowing setting another cornerstone in the research and publication history of studding the diagnosability, security, and safety of CPS modeled as hybrid dynamic systems. I would like also to acknowledge Mrs. Mary E. James for establishing the contract with Springer and supporting the editor in any
Preface
vii
organizational aspects. I hope that this volume will be a useful basis for further fruitful investigations and fresh ideas for researcher and engineers as well as a motivation and inspiration for newcomers to address the problems related to this very important and promising field of research. Douai, France
Moamar Sayed-Mouchaweh
Contents
1
Prologue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Moamar Sayed-Mouchaweh
2
Wind Turbine Fault Localization: A Practical Application of Model-Based Diagnosis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Roxane Koitz, Franz Wotawa, Johannes Lüftenegger, Christopher S. Gray, and Franz Langmayr
3
4
1
17
Fault Detection and Localization Using Modelica and Abductive Reasoning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Ingo Pill and Franz Wotawa
45
Robust Data-Driven Fault Detection in Dynamic Process Environments Using Discrete Event Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . Edwin Lughofer
73
5
Critical States Distance Filter Based Approach for Detection and Blockage of Cyberattacks in Industrial Control Systems . . . . . . . . . 117 Franck Sicard, Éric Zamai, and Jean-Marie Flaus
6
Active Diagnosis for Switched Systems Using Mealy Machine Modeling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147 Jeremy Van Gorp, Alessandro Giua, Michael Defoort, and Mohamed Djemaï
7
Secure Diagnosability of Hybrid Dynamical Systems . . . . . . . . . . . . . . . . . . 175 Gabriella Fiore, Elena De Santis, and Maria Domenica Di Benedetto
8
Diagnosis in Cyber-Physical Systems with Fault Protection Assemblies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201 Ajay Chhokra, Abhishek Dubey, Nagabhushan Mahadevan, Saqib Hasan, and Gabor Karsai
ix
x
Contents
9
Passive Diagnosis of Hidden-Mode Switched Affine Models with Detection Guarantees via Model Invalidation . . . . . . . . . . . . . . . . . . . . . 227 Farshad Harirchi, Sze Zheng Yong, and Necmiye Ozay
10
Diagnosability of Discrete Faults with Uncertain Observations . . . . . . . 253 Alban Grastien and Marina Zanella
11
Abstractions Refinement for Hybrid Systems Diagnosability Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 279 Hadi Zaatiti, Lina Ye, Philippe Dague, Jean-Pierre Gallois, and Louise Travé-Massuyès
Index . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
Chapter 1
Prologue Moamar Sayed-Mouchaweh
1.1 Cyber-Physical Systems as Hybrid Dynamic Systems Cyber-physical systems (CPS) [1] are characterized as a combination of physical (physical plant, process, network) and cyber (software, algorithm, computation) components whose operations are monitored, controlled, coordinated, and integrated by a computing and communicating core. The interaction between physical and computational components in CPS is intensive. They cover increasing number of real life applications such as autonomous vehicles, aircrafts, smart manufacturing processes, surgical robots and human robot collaboration, smart electric grids, home appliances, air traffic control, automated farming and implanted medical devices. Hybrid dynamic systems (HDS) [2] are systems in which the discrete and continuous dynamics cohabit. The discrete dynamics is described by discrete state variables while the continuous dynamics is described by continuous state variables. HDS exhibit different continuous dynamic behavior depending on the current operation mode q as follows: XP D A.q/ X C B.q/ u where X is the state vector and u is the input vector. In the case of linear systems, A(q) and B(q) are constant matrices of appropriate dimensions. The interaction between both physical and cyber components requires tools allowing analyzing and modeling both the discrete (discrete event control, communication protocols, discrete sensors/actuators, scheduling algorithms, etc.) and continuous (continuous dynamics, physics, continuous sensors/actuators, etc.)
M. Sayed-Mouchaweh ( ) Institute Mines-Telecom Lille Douai, Douai, France e-mail: moamar.sayed-mouchaweh@imt-lille-douai.fr © Springer International Publishing AG 2018 M. Sayed-Mouchaweh (ed.), Diagnosability, Security and Safety of Hybrid Dynamic and Cyber-Physical Systems, https://doi.org/10.1007/978-3-319-74962-4_1
1
2
M. Sayed-Mouchaweh
Fig. 1.1 Three-cell power converter as discretely controlled continuous system (DCCS) where capacitors C1 and C2 represent the continuous components (Cc) and switches S1 , S2 and S3 the discrete components (Dc)
dynamics. Therefore, many CPS can be modeled as hybrid dynamic systems [3, 4] in order to take into account both discrete and continuous behaviors as well as the interactions between them. There are different classes of HDS, e.g., autonomous switching systems [5], discretely controlled switching systems [2], pricewise affine systems [6], discretely controlled jumping systems [7]. Many complex systems are embedded in the sense that they consist of a physical plant with a discrete controller. Therefore, the system has several discrete changes between different configuration modes through the actions of the controller exercised on the system plant (e.g., actuators). This kind of HDS is called discretely controlled continuous or switching systems (DCCS) [7]. Piecewise affine systems [6] are another important class of HDS where complex nonlinearities are substituted by a sequence of simpler piecewise linear behaviors. The three-cellular power converter [8], depicted in Fig. 1.1, presents an example of DCCS. The continuous dynamics of the system is described by state vector X D [Vc1 Vc2 I]T , where Vc1 and Vc2 represent, respectively, the floating voltage of capacitors C1 and C2 and I represents the load current flowing from source E towards the load (R, L) through three elementary switching cells Sj , j 2 f1, 2, 3g. The latter represent the system discrete dynamics. Each discrete switch Sj has two discrete states: Sj opened or Sj closed. The control of this system has two main tasks: (1) balancing the voltages between the switches and (2) regulating the load current to a desired value. To accomplish that, the controller changes the switches’ states from opened to closed or from closed to opened by applying discrete commands
1 Prologue
3
”CSj ” or “OSj ” to each discrete switch Sj , j 2 f1, 2, 3g (see Fig. 1.1) where CSj refers to “close switch Sj ” and OSj to “open switch Sj .” Thus, the considered example is a DCCS. There are three major modeling tools widely used in the literature to model HDS. These tools are hybrid Petri nets [9], hybrid bond graphs [10], and hybrid automata [11]. Hybrid Petri nets (HPN) model HDS by combining discrete and continuous parts. HPN is formally defined by the tuple: HPN D fP; T; h; Pre; Postg where P D Pd [ Pc is a finite, not empty, set of places partitioned into a set of discrete places Pd , represented as circles, and a set of continuous places Pc , represented as double circles. T D Td [ Tc is a finite, not empty, set of discrete transitions Td and a set of continuous transitions Tc represented as double boxes. h : P \ T ! fD, Cg, called “hybrid function,” indicates for every node whether it is a discrete node (D) or a continuous node (C). Pre : Pc xT ! RC or Pre : Pd xT ! N is a function that defines an arc from a place to a transition. Post : Pi xTj ! RC or Post : Pd xT ! Nis a function that defines an arc from a transition to a place. Hybrid bond graph is a graphical description of a physical dynamic system with discontinuities. The latter represent the transitions between discrete modes. Similar to a regular bond graph, it is an energy-based technique. It is directed graphs defined by a set of summits and a set of edges. Summits represent components. The latter are: (1) passive components which transform energy into potential energy (C-components), inertia energy (L-components), and dissipated energy (Tcomponents), (2) active components that can be source of effort or source pf flow. The edges, called bonds (drawn as half arrows), represent ideal energy connections between the components. The components interconnected by the edges construct the model of the global system. This model is represented by 1 junction for components having a common flow, 0 junction for common effort and transformers and gyrators to connect different kinds of energy. In order to take into account the information during the transitions between discrete modes, hybrid bond graph is extended by adding controlled junctions (CJs). The latter allow considering the local changes in individual component modes due to discrete transitions. The CJs may be switched ON (activated) or OFF (deactivated). An activated CJ behaves like a conventional bond graph junction. Deactivated CJs turn inactive the entire incident junction and hence do not influence any part of the system. Hybrid automata are a mathematical model for HDS, which combines, in a single formalism, transitions for capturing discrete change with differential equations for capturing continuous dynamics. A hybrid automaton is a finite state machine with a finite set of continuous variables whose values are described by a set of ordinary differential equations. A hybrid automaton is defined by the tuple: G D .Q; †; X; flux; Init; ı/
4
M. Sayed-Mouchaweh
where Q is the set of states, † is the set of discrete events, X is a finite set of continuous variables describing the continuous dynamics of the system, flux : Q X ! Rn is a function characterizing the continuous dynamics of X in each state q of Q, Init D (q 2 Q, X(q), flux(q)) is the set of initial conditions and ı : Q † ! Q is the state transition function. A transition ı(q, e) D qC corresponds to a change from state q to state qC after the occurrence of discrete event e 2 †.
1.2 Diagnosability, Security, and Safety in Cyber-Physical Systems: Problem Formulation, Methods, and Challenges A fault can be defined as a non-permitted deviation of at least one characteristic property of a system or one of its components from its normal or intended behavior. Fault diagnosis is the operation of detecting faults and determining possible candidates that explain their occurrence. Online fault diagnosis is crucial to ensure safe operation of complex dynamic systems in spite of faults affecting the system behaviors. Consequences of the occurrence of faults can be severe and result in human casualties, environmentally harmful emissions, high repair costs, and economical losses caused by unexpected stops in production lines. Therefore, early detection and isolation of faults is the key to maintaining system performance, ensuring system safety, and increasing system life. Faults may manifest in different parts of the system, namely, the actuators (loss of engine power, leakage in a cylinder, etc.), the system (e.g., leakage in the tank), the sensors (e.g., reduction of the displayed value relative to the true value, or the presence of a skew or increased noise preventing proper reading), and the controller (i.e., the controller does not respond properly to its inputs sensor reading). Faults can be abrupt (e.g., the failed-on or failed-off of the pump and the stuck opened or stuck closed of the valve), intermittent or gradual (degradation of a component). Faults also may occur in a single or a multiple scenario. In the former, one fault candidate explains the observations (is responsible for the fault behavior). In the latter, several fault candidates are responsible for the fault behavior. In HDS, faults can occur as a change in the nominal values of parameters characterizing the continuous dynamics, and are called parametric faults. Faults can also occur in the form of abnormal or unpredicted mode-changing behavior and are called discrete faults. Therefore, two types of faults should be considered for HDS depending on the dynamics that is affected by faults (parametric or discrete). Discrete faults are related to faults in actuators and usually exhibit great discontinuities in system behavior, whilst parametric faults are related to tear and wear and introduce faults with much slower dynamics. For parametric faults, after the fault detection and isolation (determining the fault candidate), a fault identification phase is required in order to estimate the amplitude (e.g., the section of leakage of a tank) of the fault, its time of occurrence, its importance, etc.
1 Prologue
5 Fault types
Fault labels 1
Discrete faults
2 3
Parametric faults
4
Fault event - Fault description fs1so - 1 stuck opened fs1sc - 1 stuck closed fs2so - 2 stuck opened fs2sc - 2 stuck closed fs3so - 3 stuck opened fs3sc - 3 stuck closed fC1 – Abnormal change in the nominal values of 1 due to C1 ageing fC2 - Abnormal change in the nominal values of 2 due to C2 ageing
Fig. 1.2 Faults for the diagnosis of three-cell converters
For the example of three-cellular converters, eight faults can be considered for the diagnosis [7] as it is depicted in Fig. 1.2. Parametric faults (abnormal deviation of the nominal value of capacitors) are principally due to the effect of aging or pollution. The discrete faults (switch stuck-on or stuck-off) are more frequent and their consequences are more destructive. For instance, in open-circuit (stuck-off) failure, the system operates in degraded performance. However, unstable load may lead to further damage on the system. Therefore, the fault diagnosis of these faults is necessary to ensure the system safety and quality. The fault diagnosis task [12, 13] is generally performed by reasoning over differences between desired or expected behavior, defined by a model, and observed behavior provided by sensors. This task can be performed offline or online. Offline diagnosis assumes that the system is not operating in normal conditions but it is in a test bed, i.e., ready to be tested for possible prior failures. The test is based on inputs, e.g. commands, and outputs, e.g. sensors readings, in order to observe a difference between the resulting signals with the ones obtained in normal conditions. In online diagnosis, the system is assumed to be operational and the diagnostic module is designed in order to continuously monitor the system behavior, isolate and identify failures. Within these methods, we can distinguish between active diagnosis that uses both inputs and outputs, and passive diagnosis that uses only system outputs. The diagnosis can also be non-incremental (i.e., the diagnosis inference engine is built offline) or incremental (the diagnosis inference engine is built online in response to the observation). Diagnosability notion [13] aims at verifying if the system model is rich enough in information in order to allow the diagnosis inference engine, generally called diagnoser, to infer the occurrence of parametric and discrete faults within a bounded delay after their occurrence. The diagnosability notion was initially defined for discrete event systems (discrete faults) but it can be extended for HDS (parametric and discrete faults). In general, there are two categories of methods to build fault diagnosis inference engine allowing taking into account both the discrete and continuous dynamics in HDS as well as the interactions between them. In the first category, [14], the system’s model is an extension of the continuous model by adding the system discrete modes. The fault-free continuous behavior is defined
6
M. Sayed-Mouchaweh
Disturbances Faults Inputs
HDS
Measured outputs
Model for parameter or state/output estimation Estimated Parameters/outputs Nominal parameters/ Fault Measured Residuals detection + Residual outputs evaluation
Fault analysis
Fault isolation/ identificaion
Fig. 1.3 Internal methods for fault diagnosis
in each discrete mode by relations over observable variables. These relations are used in order to generate residuals sensitive to a certain subset of faults. A fault is diagnosed when the value of the sensitive residuals to this fault is different from zero. In the second category, the discrete model is extended or enriched by adding events generated by the abstraction of system’s continuous dynamics. There are numerous methods in the literature that are used to build the fault diagnosis inference engine in HDS. They can be divided into internal, or modelbased, and external, or data-driven, methods. The internal methods (see Fig. 1.3) use a mathematical or/and structural model to represent the relationships between measurable variables by exploiting the physical knowledge or/and experimental data about the system dynamics. They can be categorized into residual-based and setmembership [15] approaches. In residual-based approaches, the response of the mathematical model is compared to the observed values of variables in order to generate indicators used as a basis for the fault diagnosis. Generally, the model is used to estimate the system state, its output or its parameters. The difference between the system and the model responses is monitored on the basis of residual generation. Then, the trend analysis of this difference can be used to detect changing characteristics of the system resulting from a fault occurrence. Set-membership based fault diagnosis techniques are used for the detection of some specific faults. Generally, they discard models that are not compatible with observed data, in contrast to the residual-based approaches which identify the most likely model. The external methods [6, 16–19] (see Fig. 1.4) consider the system as a black box, in other words, they do not need any mathematical model to describe the system dynamical behaviors. They use exclusively a set of measurements or/and heuristic knowledge about system dynamics to build a mapping from the measurement space into a decision space. They include expert systems and machine learning and data mining techniques.
1 Prologue
7
Historical/ stream data
Data preparation
Useful data
Data preprocessing
Clean data
Data labeling Labeled data
Evaluation criteria
Model validation
Decision space
Model design
Feature space
Data analysis
Fig. 1.4 External methods for fault diagnosis
Many critical infrastructures, such as power generation and distribution networks, water networks and mass transportation systems, autonomous vehicles and traffic monitoring, etc., are CPS. Such systems, including critical infrastructures, are becoming widely used and covering many aspects of our daily life. Therefore, the security, safety, and reliability of CPS is essential for the success of their implementation and operation. However, these systems are prone to major incidents resulting of cyberattacks and system failures. These incidents affect significantly their security and safety. Attacks can be represented as an interference [20, 21] in the communication channel between the supervisor and the system intentionally generated by intruders in order to damage the system or as the enablement, respectively disablement, of actuators events that are disabled, respectively enabled, by the supervisor. In general, intruders hide, create, or even change intentionally events that transit from one device (actuator, sensor) to another in a control communication channel. These attacks in a supervisory control system can lead the plant to execute event sequences entailing the system to reach unsafe or dangerous states that can damage the system. Therefore, reliable, scalable, and timely fault and attack diagnosis is crucial in order to improve the robustness of CPS to failures and adversarial attacks. Moreover, a thorough understanding of the vulnerability of CPS’ components against such incidents can be incorporated in future design processes in order to better design such systems. Finally, the timely fault diagnosis can help operators to have better situation awareness and give them ample time to implement correction (maintenance) actions. However, guaranteeing the security and safety of CPS requires verifying their behavioral or safety properties either at design stage such as state reachability, diagnosability, and predictability or online such as fault detection and isolation. This is a challenging task because of the inherent interconnected and heterogeneous combination of behaviors (cyber/physical, discrete/continuous) in these systems. Indeed, fault propagation in CPS is not only governed by the behaviors of components in the physical and cyber sub-systems but also by their interactions. This makes the identification of the root cause of observed anomalies and predicting the failure events a hard problem. Moreover, the increasing complexity of CPS and the security and safety requirements of their operation as well as their decentralized
8
M. Sayed-Mouchaweh
resource management entail a significant increase in the likelihood of failures in these systems. Finally, it is worth to mention that computing the reachable set of states of HDS is an undecidable matter due to the infinite state space of continuous systems.
1.3 Contents of the Book This edited Springer book presents recent and advanced approaches and techniques that address the complex problem of analyzing the diagnosability property of cyberphysical systems and ensuring their security and safety against faults and attacks. The CPS are modeled as hybrid dynamic systems using different model-based and data-driven approaches in different application domains (electric transmission networks, wireless communication networks, intrusions in industrial control systems, intrusions in production systems, wind farms, etc.). These approaches handle the problem of ensuring the security of CPS in presence of attacks and verifying their diagnosability in presence of different kinds of uncertainty (uncertainty related to the event occurrences, to their order of occurrence, to their value, etc.).
1.3.1 Chapter 2 This chapter treats the problem of fault diagnosis of complex dynamic systems and its application to the aid of conditional maintenance of wind turbines. The proposed approach is an abductive model-based diagnosis where the system behavior is described logically by a set of propositions as premises (hypotheses) that entail conclusions (diagnoses). The set of propositions, describing how failures affect the system variables (i.e., components), represents the knowledge base that the diagnosis engine uses to provide an explanation for an observation. The proposed approach is based on three main phases: offline/online model development, online fault detection, and fault identification. The offline portion of the model is automatically built by exploiting failure assessments (e.g., Failure Mode Effect Analysis (FMEA) allowing failure modes characterizations and their manifestations or symptoms). The online portion of the model is built based on the fault detection phase. When the latter detects a new, unrecorded, abnormal behavior, the knowledge base of the model is updated in order to integrate this new abnormal behavior. The fault identification engine is triggered when an incorrect behavior is detected. It uses both the observed symptoms and the offline constructed model to compute the abductive diagnoses or explanations. The latter are continuously refined over time thanks to the discovered new symptoms and to the interactions with the maintenance technicians. The interactions with the latter are achieved through an interface allowing supporting the service technicians in preparing all spare parts and
1 Prologue
9
tools necessary before traveling to a wind turbine to ensure minimal downtime. In addition, on site, this interface provides contextual information as well as interactive questions in order to facilitate the diagnosis refinement or/and the discovering of new failures/abnormal behaviors. The advantage of the proposed approach is mainly related to its capacity to explain the occurrence of a failure using taxonomy adapted to the mode reasoning of technicians, operators, and supervisors. It is also related to its capacity to interact with technicians in order to facilitate the refinement of diagnosis candidates and the discovering of new failures/abnormal behaviors. The latter lead to enrich continuously the model/knowledge base over time. However, the computational complexity of the model/fault identification computation remains an issue to be solved in particular for large scale systems.
1.3.2 Chapter 3 This paper presents a solution in order to perform abductive fault diagnosis using the popular modeling language Modelica. The latter is an object-oriented, open, and multi-domain language used for representing hybrid cyber-physical systems. It allows generating intuitively models that are used to simulate the system normal behavior. The presented solution enriches Modelica models by creating a knowledge base (i.e., rules) used for abductive diagnosis. The knowledge base is built by extracting cause-effect rules from Modelica models. These rules are intuitive to designers familiar with failure mode and effect analysis (FMEA). When the difference is significant between a simulation of the normal (fault-free) model and the one where individual fault’s effects are integrated, a rule is extracted automatically. This rule is used for the identification of this individual fault. Therefore, the set of these extracted rules represents the behavior deviations of a system in response to a set of pre-defined individual faults. To compute the difference (deviation) between the signals representing normal (fault-free) and faulty behaviors, three different approaches are used: the average values (reference signal) and pre-defined tolerances, temporal band sequences, and the Pearson correlation coefficient. The time when a significant deviation is detected represents the fault detection time. The chapter uses two examples (voltage divider circuit and a switch circuit with a bulb and capacitor) in order to illustrate the proposed solution. The switch circuit is an example of a hybrid dynamic system where the continuous dynamics is represented by the current and voltage of the capacitor and resistor and the discrete dynamics is represented by the switch state (on/off). The represented solution is flexible in the sense that the augmented Modelica models can be reused for other components. The adaptation of these Modelica models to perform the abductive diagnosis allows them to be used in decision support tools to aid human operators to make decisions as the conditional/predictive maintenance. However, the robustness of the proposed solution against outliers, noises, environment, and load variations as well as other types of uncertainties need to be improved.
10
M. Sayed-Mouchaweh
1.3.3 Chapter 4 This chapter presents a data-driven modeling approach in order to perform the fault diagnosis of a production system represented as a multi-sensor network. The sensors are spatially distributed and measuring one or more channels. The different data streams, generated by the sensors, are gathered in a central data sink together with the discrete control events coming from the process control system. The fault diagnosis is performed by investigating the causal relations and dependencies between the system variables (channels). These causal relations and dependencies characterize the system states (normal/faulty) and are represented as causal relation network where nodes and vertices denote the channels and input/output variables, respectively. The use of such causal relation network may help operators and experts to gain insights into the interpretation and understanding of a failure mode development and causes. The reference models, representing the causal relations and dependencies, in fault-free conditions are used to compare the measurements issued from current conditions with the normal ones. This comparison generates residuals that are used to form a feature space. In the latter, the data measurements corresponding to normal operation conditions occupy restricted zones. When a fault occurs, the measurements occupy spaces far from these “normal” zones. The reference models are regularly updated in order to include the novelties of the system (new normal operation conditions, fault operation conditions) and thus to omit high false alarms and missed fault detection due to wrong predictions and quantifications on new samples. This update leads to more reliable fault detection and more stable model updates. A drift detection mechanism is used in order to detect a fault in early stage before becoming more severe or downtrends in the quality of components or products. The main advantages of the proposed approach are its capacity to update the generated models in response to the novelties and changes in the system environment and internal state and the use of drift detection mechanism to detect a fault in its early stage. However, updating the causal relation network (vertices and nodes) is a time-consuming task in particular for a hybrid dynamic system with multiple discrete modes.
1.3.4 Chapter 5 This chapter presents an approach for detecting intrusions in industrial control systems (ICS). These intrusions exploit the vulnerabilities of ICS in order to alter intentionally the integrity, confidentiality, and availability of the production system. These cyberattacks affect the control commands of the controller (Programmable Logic Controller PLC), the reports (sensors readings) coming from the plant as well as the communication between them. The proposed approach is based on the use of two filters: control and report filters. The control filter aims at verifying the consistency of control commands according to the current state and the sensor outputs while the report filter checks the integrity of the measurements (reports)
1 Prologue
11
according to the current state and the predicted (sent) control commands. Both filters are integrated out the production system and communicate through an independent and secured network in order to limit the cyberattack surface. They are based on three steps. The first and second steps aim at identifying offline the critical and prohibited states as well as the sequences (actions/sensor outputs) required to reach the prohibited states from any other normal (initial) state. The third state aims at detecting the cyberattacks online by observing a deviation from the normal behavior. The latter is described by the sequences of the pairs (state/action). This deviation is computed as a distance between the current state and the prohibited one. This distance is characterized by the minimal number of actions that can be applied on the process in order to reach a prohibited state. The main advantages of this approach are: (1) it is a non-intrusive approach since it does not need to install new probes in the production system and (2) it limits the cyberattack surface since the control and report filters are installed out the PLC and the plant and they communicate through an independent and secure network. However, the proposed approach cannot adapt to new cyberattacks (new prohibited and dangerous states) and its computation complexity grows exponentially for large scale systems with multiple discrete states.
1.3.5 Chapter 6 This chapter presents an approach to fault diagnosis of switched systems modeled by a Mealy machine (an automaton with inputs and outputs). Some transitions of this automaton, including those corresponding to faults, may occur in the absence of a control input and therefore are unobservable. Consequently, several states of the diagnoser (the model performing the fault diagnosis) are uncertain in the sense that they contain several fault labels (indicating several faulty components) and therefore they cannot isolate the responsible component of the fault occurrence. The diagnosis, performed by the proposed approach, is active in the sense that it ensures simultaneously the control and the diagnosability of the system. In order to remove the diagnosis uncertainty, the proposed approach computes the fault isolating sequence that leads to reach a certain diagnosis state with one fault label indicating the fault component. The proposed approach considers the initial conditions of the system are known and the initial mode is without fault. It considers also that only control input that drives the evolution of the system is represented by the switching function. This function specifies the active mode. Furthermore, discrete outputs are also available, as a result of each mode transition, in order to detect and isolate the fault. When a fault is detected, the nominal control is suspended for safety reason. The proposed approach requires building the system model with all its nominal and fault discrete states. This model, called diagnose, is then associated with the minimal fault isolating sequences for all uncertain states in the diagnoser. The proposed approach is illustrated and tested using a multicellular power converter in order to diagnose the discrete faults (switch blocked-on/blockedoff). The discrete dynamics of the power converter is represented by the on/off
12
M. Sayed-Mouchaweh
switches, while the continuous dynamics are related to the charging/discharging of the capacitors. The advantage of the proposed approach is related to its improvement of the system diagnosability by applying the fault isolating sequences. However, the proposed approach does not scale well according to the system size in particular with multiple discrete modes.
1.3.6 Chapter 7 This chapter studies and investigates the security issues for hybrid dynamic systems in presence of attacks. The latter are represented by compromised sensor measurements exchanged by the means of a wireless communication network. The challenge to ensure the system security and safety is to correctly estimate or reconstruct instantaneously or within a finite time interval the system internal state despite the presence of corrupted sensor measurements by external hackers. This chapter formalizes the required conditions in order to perform a secure state estimation despite the malicious attacks. To this end, the continuous dynamics is abstracted in order to generate discrete events that are used to enrich the system model. Only a subset of sensors of fixed size is considered to be corrupted. However, the sensors of this subset are unknown. Then, the link between diagnosability notion, developed for discrete event systems, and the secure state estimation is explored. The goal is to determine the conditions that allow distinguishing normal states from corrupted ones. The main advantage of the proposed approach is its capacity to estimate the secure state in the context of hybrid dynamic states where problems of decidability and computational complexity arise because of the cohabitation of both discrete and continuous dynamics. However, the proposed approach is restricted to linear nonZeno systems with fixed size of corrupted sensors.
1.3.7 Chapter 8 This chapter proposes a hierarchical component-based approach for fault diagnosis and prognosis as well as failure mitigation in critical cyber-physical systems (CPS). The latter is composed by the physical system (plant), the actuators, the protection devices, and the discrete controllers. The latter try to arrest the failure effect if detected. The proposed approach uses temporal causal diagrams in order to describe the consequences of fault occurrence and propagation in physical and cyber components. The built model accounts for normal and fault behaviors. Each behavior is represented by a sequence of states and events. A state is characterized as safe or harmful. When a primary fault occurs, its propagation enforces the protection devices to be activated. This leads the system to a blackout state. The fault prognosis is based on the determination whether the system at its current state satisfies the constraints to reach a blackout state. If the answer is yes, then the trajectory (set of
1 Prologue
13
states and transitions) reachable from the current state is computed. The proposed approach is used for fault diagnosis and prognosis as well as failure mitigation of power systems, in particular electric transmission networks. The physical system consists of generators, buses, transmission lines and loads, while actuators are the circuit breakers and the protection devices are the relays. The latter cause system reconfiguration by instructing actuators to change their state. Two types of faults are considered: phase to phase faults and phase to ground fault. The faults of the breakers (breaker stuck-closed and breaker stuck-opened) and distance relays (missed detection faults and false alarms) are considered. The observable events in the case of power transmission system are commands sent by relays to breakers, messages sent by relays to each other, state change of breakers, physical fault detection alarms, etc. The faults are the unobservable events. The proposed approach is modular in the sense that each component has its proper model and its proper local diagnoser and a reasoner is used to compute the global diagnosis decision. However, a global model is needed to build the diagnosers and the reasoner.
1.3.8 Chapter 9 This chapter presents a model-based and data-driven approach to perform the fault diagnosis of cyber-physical systems that are safety critical, yet prone to system failures. Fault refers to any fault, attack, or anomaly. The nominal behaviors and the fault modes are represented by hidden-mode switched affine models with timevarying parametric uncertainty subject to process and measurement noise. The proposed approach is based on three steps: model invalidation, fault detection, and fault isolation. Model invalidation aims at determining whether an input–output sequence over a horizon T is compatible with a switched affine model. If the data is not compatible with the nominal model, then the model is invalidated and a fault is detected. When a fault is detected, the fault isolation step aims at uniquely determining which specific fault model is validated or rather, not invalidated, based on the measured input–output data. The proposed approach is evaluated for the simple and multiple fault diagnosis of the Heating, Ventilating, and Air Conditioning (HVAC) system. The considered faults are: faulty fan (the fan rotates at half of its nominal speed), faulty chiller water pump (the pump is stuck and spins at half of its nominal speed), and faulty humidity sensor (the humidity measurements are biased by an amount of C0.005). The advantage of this approach is its ability to diagnose simple and multiple faults either discrete or continuous (parametric) by taking into account the noises and parameter uncertainties. However, it requires an important effort and depth knowledge to build offline the nominal and fault models. Moreover, it needs to determine the horizon T (time delay) required to distinguish between normal and fault behaviors (T-detectability) and between two different fault models (I-isolability).
14
M. Sayed-Mouchaweh
1.3.9 Chapter 10 This chapter proposes an extension of the diagnosability notion proposed for discrete event systems (DES) for hybrid dynamic systems (HDS) with uncertain observation. This chapter provides the answer to the question: can diagnosability be achieved even if the observation is uncertain? that is, when the order of the observed events and/or their (discrete) values are partially unknown. Indeed, in many applications, the temporal order of the observable events that have occurred within the DES is not always known, in particular when they occur in a short time span. In addition, the occurrence of some events is not certain in the sense that they may have occurred or not. Therefore, the developed diagnosability notion in this chapter considers the combination of these two types of uncertainties. However, the time delay required to verify the diagnosability for HDS is not proved to be bounded.
1.3.10 Chapter 11 This chapter proposes a diagnosability notion adapted to hybrid dynamic systems (HDS). It proposed an algorithm able to verify at design state if a fault that would occur at runtime could be unambiguously detected within a given finite time using only the allowed observations. The proposed algorithm is based on the abstractions that discretize the infinite state space of the continuous variables into finite sets. It starts by generating the most abstract discrete event system (DES) model of the HDS and checking diagnosability of this DES model. A counterexample that negates diagnosability is provided based on the twin plant. This counterexample is obtained when there is a path in the twin model with at least one ambiguous state (state with two different diagnosis labels) cycle. The model is then refined in order to try to invalidate the counterexample and the procedure repeats as far as diagnosability is not proved. If the counterexample is validated, then the system is not diagnosable. If there is no validated counterexample, then the system is diagnosable. The refinement is based on the understanding of the causes that entailed the refusal of the counterexample. Then, this spurious counterexample or any close spurious counterexamples will be eliminated next time. This will make the best out of computation. The proposed algorithm was illustrated using two examples. The first example is a server with 4 buffers. Each buffer is assigned a workflow and the switching between buffers is controlled by a user input. The second example is a classical thermostat with two different faults. The first fault is discrete fault due to a bad calibration of the temperature sensor, while the second fault is parametric due to a problem in the heater. The proposed algorithm has the advantage to account explicitly for the hybrid dynamic nature of the system and to verify the diagnosability in cost-effectiveness analysis. However, the time delay required to verify the diagnosability is not proved to be bounded.
1 Prologue
15
References 1. R. Rajkumar, I. Lee, L. Sha, J. Stankovic, Cyber-physical systems: the next computing revolution, in Proceedings of the 47th Design Automation Conference (ACM, 2010), pp. 731– 736 2. A.J. Van Der Schaft, J.M. Schumacher, An Introduction to Hybrid Dynamical Systems, vol 251 (Springer, London, 2000) 3. A. Benveniste, T. Bourke, B. Caillaud, M. Pouzet, in Hybrid systems modeling challenges caused by cyber-physical systems, cyber-physical systems (CPS) foundations and challenges, ed. by J. Baras, V. Srinivasan. Lecture Notes in Control and Information Sciences (2013) 4. Y. Yalei, X. Zhou, Cyber-physical systems modeling based on extended hybrid automata, in 5th IEEE International Conference on Computational and Information Sciences (ICCIS) (2013) 5. M.S. Branicky, V.S. Borkar, S.K. Mitter, A unified framework for hybrid control: model and optimal control theory. IEEE Trans. Autom. Control 43(1), 31–45 (1998) 6. L. Rodrigues, S. Boyd, Piecewise-affine state feedback for piecewise-affine slab systems using convex optimization. Syst. Control Lett. 54(9), 835–853 (2005) 7. H. Louajri, M. Sayed-Mouchaweh, Decentralized diagnosis and diagnosability of a class of hybrid dynamic systems, in Informatics in Control, Automation and Robotics (ICINCO), 2014 11th International Conference on, vol. 2 (2014) 8. M. Shahbazi, E. Jamshidpour, P. Poure, S. Saadate, M.R. Zolghadri, Open-and short-circuit switch fault diagnosis for nonisolated dc dc converters using eld programmable gate array. IEEE Trans. Ind. Electron. 60(9), 4136–4146 (2013) 9. R. David, H. Alla, Discrete, Continuous, and Hybrid Petri Nets (Springer, Berlin, 2010) 10. D. Wang, S. Arogeti, J.B. Zhang, C.B. Low, Monitoring ability analysis and qualitative fault diagnosis using hybrid bond graph. IFAC Proc. 41(2), 10516–10521 (2008) 11. T.A. Henzinger, The theory of hybrid automata, in Verification of Digital and Hybrid Systems, (Springer, Berlin, 2000), pp. 265–292 12. M. Sayed-Mouchaweh, E. Lughofer, Decentralized fault diagnosis approach without a global model for fault diagnosis of discrete event systems. Int. J. Control. 88(11), 2228–2241 (2015) 13. M. Sayed-Mouchaweh, Discrete Event Systems: Diagnosis and Diagnosability (Springer, New York, 2014) 14. T. Kamel, C. Diduch, Y. Bilestkiy, L. Chang, Fault diagnoses for the Dc filters of power electronic converters, in Energy Conversion Congress and Exposition (ECCE) (IEEE, 2012), pp. 2135-2141 15. M. Tabatabaeipour, P.F. Odgaard, T. Bak, J. Stoustrup, Fault detection of wind turbines with uncertain parameters: a set-membership approach. Energies 5(7), 2424–2448 (2012) 16. L. Hartert, M. Sayed-Mouchaweh, Dynamic supervised classification method for online monitoring in non-stationary environments. Neurocomputing 126, 118–131 (2014) 17. M. Sayed-Mouchaweh, N. Messai, A clustering-based approach for the identification of a class of temporally switched linear systems. Pattern Recogn. Lett. 33(2), 144–151 (2012) 18. H. Toubakh, M. Sayed-Mouchaweh, Hybrid dynamic data-driven approach for drift-like fault detection in wind turbines. Evol. Syst. 6(2), 115–129 (2015) 19. M. Sayed-Mouchaweh, Diagnosis in real time for evolutionary processes in using pattern recognition and possibility theory. Int. J. Comput. Cognit. 2(1), 79–112 (2004) 20. L.K. Carvalho, Y.C. Wu, R. Kwong, S. Lafortune, Detection and prevention of actuator enablement attacks in supervisory control systems, in 13th International Workshop on Discrete Event Systems (2016), pp. 298–305 21. D. Thorsley, D. Teneketzis, Intrusion detection in controlled discrete event systems, in 45th IEEE Conference on Decision and Control (2006) pp. 6047–6054
Chapter 2
Wind Turbine Fault Localization: A Practical Application of Model-Based Diagnosis Roxane Koitz, Franz Wotawa, Johannes Lüftenegger, Christopher S. Gray, and Franz Langmayr
2.1 Introduction The increasing complexity and magnitude of technical systems is leading to a demand for effective and efficient automatic diagnosis procedures to identify failure-inducing components in practice. This is especially true in application areas experiencing excessive service costs and idle time revenue loss. In the industrial wind turbine domain operation and maintenance constitute significant factors in terms of turbine life expenditure. Given the remote locations onshore and offshore of wind turbine installations, accurate fault identification is essential for reducing costs and risks of component failures as well as turbine downtime [14]. Wind turbine diagnosis is complicated, however, since their overall reliability is affected by a multitude of failure modes concerning all major sub-systems and environments and furthermore the load conditions change regularly [36]. While electrical and control systems account for most wind turbine failures, other sub-systems, such as gearboxes, cause extensive downtimes due to the complexity of maintenance and thus pose a higher cost risk [35]. Unfortunately, currently implemented standard alarm systems deliver a large number of false alarms and thus are not suitable for standalone fault detection and identification [14]. Wind turbine operators often rely on time-based maintenance, where turbines are inspected periodically to assess their condition. This practice may lead to unnecessary turbine downtime for healthy systems, while failure-inducing conditions between services remain unnoticed. Due to these disadvantages predictive R. Koitz ( ) · F. Wotawa · J. Lüftenegger Graz University of Technology, Graz, Austria e-mail: rkoitz@ist.tugraz.at; wotawa@ist.tugraz.at; jlueften@ist.tugraz.at C. S. Gray · F. Langmayr Uptime Engineering GmbH, Graz, Austria e-mail: c.gray@uptime-engineering.com; f.langmayr@uptime-engineering.com © Springer International Publishing AG 2018 M. Sayed-Mouchaweh (ed.), Diagnosability, Security and Safety of Hybrid Dynamic and Cyber-Physical Systems, https://doi.org/10.1007/978-3-319-74962-4_2
17
18
R. Koitz et al.
and condition-based maintenance have become increasingly popular. Both rely on condition monitoring software and diagnosis methods [23]. Condition monitoring software utilizes the signals transmitted from the sensors integrated within the turbine and further processes the information to derive health information of critical components, e.g., gearboxes and main bearings. Some specification indicators of the subsequent failure are observable prior to around 99% of equipment fault occurrences [35]. Unnecessary maintenance activities can be avoided by scheduling repair or replacement of components based on their present or impending failure risks [1]. Numerous approaches exist for wind turbine diagnosis. Signal processing techniques analyze the multidimensional turbine data without considering an a priori developed mathematical model to extract faults based on spectral analysis or trend checking techniques, while machine learning methods, such as neural networks, can rely on historic data for failure identifications [16]. Zaher and McArthur [41] introduce a fault and degradation detection system for entire wind turbine installations based on supervised learning of the nominal behavior. By collecting data from various downtimes their system computes an overall turbine operational behavior model. Schlechtingen et al. [29] also adopt machine learning by creating neural networks based on the normal wind turbine behavior in combination with fuzzy rules representing expert knowledge on faults. Their approach requires the availability of Supervisory Control And Data Acquisition (SCADA) operation logs, which provide 10-min values of various measurements, such as power output, rotor speed, or gearbox oil temperature [40]. Anomalies can be detected by comparing the normal-behavior model with the actual performance. The fuzzy inference system then automatically identifies the faulty components. Gray et al. [14, 15] describe a combination of diagnostic and prognostic techniques exploiting the relations between operational and environmental loads as also damage accumulation rates. Based on an analysis of potential component-based failure modes and their damage driving physics, a mathematical model is computed that can be used to calculate the rate at which damage accumulates in response to the operating environment. This method offers a means for determining current and projected failure probabilities based on the derived damage model and statistical failure model. Statements about absolute remaining useful life cannot be made, since the load capacity would have to be known in advance to a high degree of accuracy. This is very rarely the case, and considerable variation occurs due to, e.g., variations in material quality, tolerances in component manufacture, influence of transport, installation and configuration. Therefore the prognostic method focuses on quantification of the applied loads instead, and uses probabilistic methods to relate the said loads to damage accumulation. Model-based diagnosis (MBD) techniques have been developed in the Fault Detection and Isolation (FDI) and the Artificial Intelligence (AI) community. Approaches stemming from the FDI field usually depend on quantitative models, while the AI methods utilize qualitative/symbolic representations of the underlying system to draw conclusions regarding the state of the system and its components
2 Wind Turbine Fault Localization: A Practical Application of Model-Based. . .
19
[1, 6, 9, 28]. For instance, Echavarria et al. [10] utilize qualitative physics to formalize the behavior of wind turbines. Combined with a solver, their model-based system is capable of detecting and identifying faults. In this chapter, we focus on the AI variation of MBD. MBD relies on a description of the system, which together with abnormal observations of the current system state is exploited to derive root causes [9, 28]. Multiple applications have been developed for diverse fields, e.g., the automotive industry [32] or environmental decision support systems [39]. Even though MBD can look back on decades of research, a widespread dissemination in practice is still lacking. The reasons for this include the effort associated with the development of diagnosis models [4] and the difficulties of adequately integrating MBD into current industrial work processes [12, 24, 31]. In cooperation between Graz University of Technology and Uptime Engineering1 a project was initiated with the aim of providing a methodology and framework for MBD in the industrial domain. Due to Uptime Engineering’s many years of experience and expertise in the field of wind power plant maintenance, industrial wind turbines constitute an ideal test bed and application area for MBD. In this chapter, we present some of the results of the collaboration as well as the ongoing realization of MBD in practice. With this work we aim to bridge the gap between the theory of MBD and its practical application. We first discuss the foundations of MBD and present a general process for integration of this approach in real-world fault identification. Subsequently, we introduce an application designed to facilitate diagnosis in the industrial wind turbine domain. In particular, the application’s graphical user interface (GUI) is presented, which has been created taking the needs, work processes, and environments of the maintenance personnel into consideration. Subsequently, we discuss the current status of the integration of an MBD engine in the industrial wind turbine domain and provide some concluding remarks.
2.2 Model-Based Diagnosis Model-based reasoning fosters the idea of reusing knowledge by relying on a formalization of the system under consideration. The model together with a set of observed symptoms can be exploited to obtain diagnostic hypotheses for the observations. Two variations have emerged in the literature: consistency-based and abductive MBD. Consistency-based diagnosis utilizes a description of the correct system behavior and identifies root causes through inconsistencies arising from the model in combination with the given symptoms. A diagnosis is then a set of abnormality assumptions about the system such that the observations and assumptions are consistent [9, 28]. In contrast, abductive diagnosis is based on the notion of logical entailment [26]. A set of premises logically entails a 1 Uptime Engineering GmbH provides consulting services as well as software tools in the field of technical reliability.
20
R. Koitz et al.
conclusion if and only if for any interpretation in which holds is also true. We write this relation as ˆ and call a logical consequence of . A set of abnormality assumptions entailing the observations constitutes an abductive diagnosis or explanation. To derive causes for observed anomalies by utilizing this type of inference, the abductive MBD approach depends on a model representing the links between faults and their manifestations. Even though both variations are based on different reasoning techniques, Console et al. [5] showed the close relation between consistency-based and abductive diagnosis. In the upcoming portion of the chapter, we focus on abductive MBD. First, we describe an abductive diagnosis problem and its solution based on a subset of propositional logic, namely Horn clauses. Subsequently, we discuss a process facilitating the incorporation of abductive MBD in real-world applications.
2.2.1 Propositional Horn Clause Abduction A Horn clause is defined as a disjunction of literals featuring at most one positive literal and can be described by a rule, e.g., f:a1 ; : : : ; :an ; anC1 g can be written as a1 ^ : : : ^ an ! anC1 . Similar to Friedrich et al. [13], we define a knowledge base (KB) representing the abductive diagnosis model in the context of propositional Horn clause abduction. Definition 1 A knowledge base (KB) is a tuple (A,Hyp,Th) where A denotes the set of propositional variables, Hyp A the set of hypotheses, and Th the set of Horn clause sentences over A. A hypothesis, also referred to as an assumption, is a propositional variable for which we can presume a certain truth value. Hypotheses are the propositions which can be part of a diagnosis, while the Horn theory depicts the relationships between the variables.
Example 1 Gearbox lubrication is an essential aspect of industrial wind turbine reliability as it protects the contact surfaces of gears and bearings from excessive wear and prevents overheating. Considering a simplified scenario, insufficient lubrication can be caused by a damaged oil pump, which leads to loss of oil pressure and therefore a reduction in the flow rate of oil through the system. Furthermore, a blockage of the filter in the oil cooling system may cause overheating of the oil, which also negatively affects the lubrication due to a reduction in the film thickness at the bearing and gear contacts. (continued)
2 Wind Turbine Fault Localization: A Practical Application of Model-Based. . .
21
Example 1 (continued) Starting from this description, we can identify two root causes of insufficient lubrication: a blocked filter or a damaged oil pump. These causes constitute the faults we want to identify during diagnosis and thus their corresponding variables form the set of hypotheses: ˚ Hyp D damaged_pump; blocked_filter The set of propositional variables A comprises all hypotheses as also propositions representing effects: 8 9 < damaged_pump; blocked_filter; reduced_pressure; overheating; = AD reduced_film_thickness_bearing_contacts; : ; reduced_film_thickness_gear_contacts; poor_lubrication Given the set of propositional variables, the circumstances leading to an insufficient greasing of the gearbox can be represented by a Horn theory: 8 9 damaged_pump ! reduced_pressure; ˆ > ˆ > ˆ > ˆ > ˆ > reduced_pressure ! poor_lubrication; ˆ > ˆ > ˆ > ˆ > blocked_filter ! overheating; < = Th D overheating ! reduced_film_thickness_bearing_contacts; ˆ > ˆ ˆ overheating ! reduced_film_thickness_gear_contacts; > > ˆ > ˆ > ˆ > ˆ > reduced_film_thickness_bearing_contacts^ ˆ > > :̂ ; reduced_film_thickness_gear_contacts ! poor_lubrication
Definition 2 Given a knowledge base (A,Hyp,Th) and a set of observations Obs A then the tuple (A,Hyp,Th,Obs) forms a propositional Horn clause abduction problem (PHCAP). A diagnosis problem involves a KB plus a set of observations for which the explanations are to be computed. In our context, these observables may only be a conjunction of propositions and not an arbitrary logical sentence. The solution to a PHCAP or diagnosis is a set of hypotheses explaining the propositions in Obs, i.e., entailing them together with the theory Th. In other words, the observations are a logical consequence of the failure relations described in the theory and the determined explanation. An additional requirement is that only consistent diagnoses are permitted, thus solutions leading to a contradiction are disregarded. Imposing a parsimonious criterion on the solutions is a principle commonly used in diagnosis. From our practical point of view only subset minimal explanations are of interest.
22
R. Koitz et al.
Definition 3 Given a PHCAP (A,Hyp,Th,Obs). A set Hyp is a solution if and only if [ Th ˆ Obs and [ Th 6ˆ ?. A solution is parsimonious or minimal if and only if no set 0 is a solution. -Set contains all solutions obtained from a PHCAP.
Example 1 (continued) Considering the PHCAP and assuming we detect an insufficient lubrication of the gearbox, i.e., Obs = fpoor_lubricationg, we can derive two minimal explanations; either the oil pump is faulty inducing a reduction of the oil flow ( 1 D fdamaged_pumpg) or a blocked filter causes poor cooling which leads to insufficient greasing of the bearing and gear contacts ( 2 D fblocked_filterg), i.e., Set D ffdamaged_pumpg; fblocked_filtergg.
While abductive reasoning provides an intuitive approach for fault localization, its computational complexity for general propositional theories is located within the second level of the polynomial hierarchy [11]. Focusing on a less expressive modeling languages, such as in our case Horn clause models, reduces the complexity. Yet, Friedrich et al. [13] showed that computing the solution to a PHCAP is still NPcomplete. Thus, for practical applications efficient solvers are essential to compute diagnoses in a reasonable time frame.
2.2.2 Incorporating Diagnosis into Practice While MBD offers several attractive features, such as allowing the reuse of already created system models and a clean separation between the problem description and its solving mechanism, the dissemination of implementations in practice is limited [31]. The integration of MBD in industrial applications is impeded by two main drawbacks associated with this type of reasoning. First, the computational complexity as mentioned in the previous section discourages the practical use in cases where diagnoses are to be computed within short periods of time. Second, model-based reasoning techniques always demand the existence of a system description, whether it is of the correct behavior in consistency-based diagnosis or how failures affect system variables in the context of the abductive variation. Developing a model is associated with an initial effort and acquiring a technical description of a system suitable for diagnostic purposes can be challenging and is often hindered by organizational issues. Further, a lack of tools facilitating the model generation and integrating it into existing work processes complicates the modeling phase [2].
2 Wind Turbine Fault Localization: A Practical Application of Model-Based. . . OFFLINE
23
ONLINE
Model Development Failure Assessment
Fault Detection
Mapping
Diagnoses
KB
Diagnosis Engine
Data Analysis
Observations
Data Acquisition
Probing Results Additional Measurements
Ranking
Ranked Diagnoses
Observation Discrimination
Fault Identification
Ranked Probing Points Repair or Replacement
Fig. 2.1 Abductive MBD process (adapted from [18])
To counteract these factors Koitz and Wotawa [18] have defined a process based on discussion with Uptime Engineering for applying abductive MBD to real-world applications. The method relies on an automated model creation on the one hand and on the other hand ensures an efficient diagnosis computation by limiting the models to Horn logic. A graphical representation of the activities is depicted in Fig. 2.1. The process is divided into three phases: model development, fault detection, and fault identification. In order to lower the entrance barrier for implementing the MBD approach, models are automatically built by exploiting failure assessments frequently used in practice. Such analyses must characterize failures and their manifestations to be mapped to a KB required for abductive reasoning. A wide-spread tool which can be used for this purpose is Failure Mode Effect Analysis (FMEA) incorporating expert knowledge on component failures [38]. Constructing the model from a fault assessment only needs to be performed once—given that there are no updates to the analysis—and can be accomplished offline. The online portion of the MBD process is prompted once the presence of a fault has been detected by a mechanism discovering the existence of an incorrect system behavior such as a condition monitoring system. Based on the observed symptoms and the offline constructed model, the fault can be identified by deriving the abductive diagnoses. Various approaches are capable of computing abductive explanations [17] and further additional refinements to the initial diagnoses can be made via additional observations and prioritization of the results in regard to certain objectives, e.g., diagnosis likelihood or maintenance cost.
2.2.2.1
Model Development
Automatic generation of a suitable knowledge base using information available a priori is an essential feature of the proposed diagnosis process as it reduces additional modeling efforts. While different assessments can be utilized, we show
24
R. Koitz et al.
Table 2.1 Example 2: FMEA excerpt (adapted from [27]) Component Yaw drive
Failure mode Fails to rotate
Yaw drive
Drive shaft blocked
Failure effect No yaw, safety system failure, decrease of efficiency No yaw, decrease of efficiency
Likelihood 2.2E 5
Severity V
1.3E 5
IV
here an example of a conversion based on FMEA [38]. FMEA is an established standardized reliability approach, in which an expert group analyzes a system and determines potential component-based single faults. Each failure mode is examined in regard to its causes, consequences, and various other characteristics [3]. Table 2.1 depicts an excerpt of an FMEA for the yaw drive of a wind turbine presenting two failure modes. While in general an FMEA can feature more columns than the ones shown in Table 2.1 depending on the standard followed, the modeling methodology as proposed by Wotawa [38] focuses on three aspects which have to be considered within the analysis: the components (COMP), the failure modes (MODES), and moreover, the existence of propositions PROPS corresponding to observable failure effects. Definition 4 An FMEA is a set of tuples (C, M, E) where C 2 COMP is a component, M 2 MODES is a failure mode, and E PROPS is a set of effects. As the FMEA typically contains a description of how each fault affects a set of system variables, we can convert this information in a straightforward manner to a logical KB, where the hypotheses comprise the component-based failures and the theory consists of propositional Horn clause sentences describing the causeeffect relation depicted in the FMEA. A variable mode(C,M) is constructed for each component-failure mode pair in the analysis, where C is the component and M is the failure mode. These propositions compose the set of hypotheses: [
Hyp Ddef
fmode.C; M/g
(2.1)
.C; M; E/2FMEA
To form the set of all variables A, the union over all hypotheses as well as propositions representing effects is constructed: A Ddef
[
E [ fmode.C; M/g
(2.2)
.C; M; E/2FMEA
Each record in the FMEA describes the effects of a single fault, thus, the relations between defects and their manifestations can be transformed into a Horn model in a straightforward way. Let HC be the set of Horn clause sentences, then the mapping function M W 2FMEA 7! HC generates a set of Horn clauses which are a subset of HC for each record in the FMEA.
2 Wind Turbine Fault Localization: A Practical Application of Model-Based. . .
25
Definition 5 Given an FMEA, the function M is defined as follows: [ M.FMEA/ Ddef M.t/
(2.3)
t2FMEA
where M.C; M; E/ Ddef fmode.C; M/ ! e je 2 E g
(2.4)
Example 2 (continued) The FMEA in Table 2.1 features the two componentfault mode pairs (Yaw Drive, Fails to rotate) and (Yaw Drive, Drive shaft blocked). Their corresponding propositional variables are added to Hyp. Hyp D
mode( Yaw_Drive; Fails_to_rotate); mode( Yaw_Drive; Drive_shaft_blocked)
The set of all propositions then contains the hypotheses as well as all variables corresponding to effects. 8 <
9 mode( Yaw_Drive; Fails_to_rotate); = AD mode( Yaw_Drive; Drive_shaft_blocked); : ; no_yaw; safety_system_failure; decrease_of _efficiency For each manifestation contained in a record, a rule is built such that the single hypothesis representing the component-fault mode pair implies the effect. The theory is then simply a union over all these Horn clauses. 8 ˆ ˆ ˆ ˆ ˆ <
9 > mode( Yaw_Drive; Fails_to_rotate) ! no_yaw; > > > mode( Yaw_Drive; Fails_to_rotate) ! safety_system_failure; > = Th D mode( Yaw_Drive; Fails_to_rotate) ! decrease_of _efficiency; ˆ > ˆ > ˆ > mode( Yaw_Drive; Driveshaft_blocked) ! no_yaw; ˆ > > :̂ ; mode( Yaw_Drive; Drive_shaft_blocked) ! decrease_of _efficiency
Due to the structure of an FMEA, the resulting logical system description is acyclic2 and consists of bijunctive Horn clauses, i.e., implications always lead from one hypothesis to a single effect variable. This results in an efficient diagnosis computation and a system considering the single fault assumption [19].
2 A logical theory is acyclic in cases when it can be represented by a directed acyclic graph where each proposition is represented as a node and an edge is drawn from a proposition to another in case the former directly implies the latter.
26
R. Koitz et al.
A shortcoming of FMEA is that it does not take into account the potential interdependencies between various manifestations, which might be essential from a practical point of view to describe how a fault affects the system. Thus, other failure assessments, such as fault trees, can be used as the basis of the automatic modeling [20]. Depending on the underlying failure analysis type, the resulting diagnosis system description may feature different characteristics, such as being a non-bijunctive Horn model. Independent of the assessment type, the accuracy and composition of the failure analysis largely impact the quality of the automatically generated diagnosis model. It is apparent that failures and manifestations disregarded in the failure review are missing from the system description and thus cannot be considered during fault identification. Hence, to achieve precise diagnoses, model completeness is an essential premise [24]. Furthermore, manifestations must be detectable in order to be useful in a diagnostic context and effects as well as failures have to be coherently reported throughout the assessment to allow automatic processing.
2.2.2.2
Fault Identification
Abductive diagnosis, even in the case where the system description is restricted in expressiveness to Horn sentences, is at least NP-complete [13]. Hence, efficient methods for deriving diagnoses are required in practice. There are various techniques capable of computing abductive diagnoses such as SATbased approaches [17], consequence finding procedures [30], or the well-known Assumption-based Truth Maintenance System (ATMS) [8]. Internally the ATMS operates on a directed graph representing the logical relations contained in the theory, where propositions and the contradiction are nodes and implications determine the edges. Each node is equipped with a label recording for the corresponding variable the sets of assumptions, i.e., hypotheses, it can be derived from. Thus, the ATMS documents the entailment relations characterized within the theory [22] and ensures that labels are consistent and minimal with respect to subsumption. To compute the abductive diagnoses an additional implication is added to the ATMS such that o1 ^ : : : ^ on ! ex, where fo1; : : : ; on g D Obs and ex represents a new propositional variable not yet contained within A. The label of ex then comprises all solutions to the PHCAP. MBD may yield an exponential number of explanations in the worst case. Thus, techniques assisting in distinguishing diagnoses are required to allow for an effective decision making in regard to repair and replacement activities. Subsequently, we present two methods aiming at supporting fault identification: observation discrimination and diagnosis ranking. Observation Discrimination Probing has been proposed as a means to decrease the solution space by supplying additional facts to the diagnostic reasoner. While Friedrich et al. [13] propose an interleaved process between diagnosis, probing and repair, Wotawa [37] suggests computing all explanations and subsequently adding new symptoms, which allows either the removing or confirming of diagnoses.
2 Wind Turbine Fault Localization: A Practical Application of Model-Based. . .
27
Definition 6 Given a PHCAP (A,Hyp,Th,Obs) and two diagnoses 1 and 2 . A new observation o 2 A n Obs discriminates two diagnoses if and only if is a diagnosis for (A,Hyp,Th,Obs [ fog) but 2 is not. Once discriminating observations have been selected, probes are taken and the fault identification process is restarted with the additional measurement information. Determining an ideal probing point is essential in order to converge to a plausible solution efficiently. The best new observation o is the probing point with the highest entropy H(o). Entropy represents the information gain; thus, a higher entropy value indicates a measurement with a greater discrimination capability [9]. H.o/ D p.o/ log2 p.o/ .1 p.o// log2 .1 p.o//
(2.5)
Equation (2.5) defines the entropy value for an observation o, where p(o) is the probability of o defined as the ratio between the diagnoses entailing the symptom together with the theory and the total number of explanations: p.o/ D
jf j 2 Set; [ Th ˆ foggj j Setj
(2.6)
Diagnosis Ranking Depending on the underlying logical theory and set of observations, there might not be a single solution available. Thus, in these cases a prioritization of the diagnosis results can be useful to initiate appropriate maintenance activities. A common strategy is to exploit probabilities. Considering Bayes rule for conditional probability, we can define the probability of an explanation given an observation o as p. j o/ D
p.o j /p. / : p.o/
(2.7)
Under the presumption that there is no uncertainty in the measurement, i.e., the data has not been subjected to errors or noise, we can state for any o 2 Obs that p.o/ D 1. As we known from the entailment relation required by abductive diagnosis that the explanation logically implies the observation, we can assign p.o j / D 1. Consequently from these two assignments and Eq. (2.7) it follows that p. j o/ D p. /. Assuming independence amongst faults, the probability of each diagnosis can be computed based on the a priori probabilities p(h) of the hypotheses: p. / D
Y h2
p.h/
Y
.1 p.h//
(2.8)
h…
Given a PHCAP’s solutions we compute p( ) for all diagnoses in -Set and subsequently assign ranks accordingly. FMEA, for instance, holds additional information such as failure likelihoods, which can be utilized for prioritization. Other criteria, such as repair and replacement costs or fault and diagnosis severity, i.e., seriousness of consequences in regard to safety or monetary considerations, could also be considered instead [34].
28
R. Koitz et al.
2.3 Industrial Wind Turbine Diagnosis Wind turbine reliability presents a very interesting use-case for the application of diagnostic methods. The cost of electrical energy produced depends strongly on the operational efficiency of the machines as also on the availability. Component faults leading to unplanned downtime have been shown to impact the overall energy production significantly, the financial motivation for optimization in this respect is thus high. The use of remote detection and diagnostic technology is an area which is receiving an increasing level of focus, in particular in the offshore wind energy industry, where turbine failures are even more critical due to the difficulties related to accessing and repairing the machines in potentially harsh environmental conditions. All modern wind turbines use sensors, data acquisition, and on-board processing as part of the closed-loop control system. Furthermore, a range of diagnostic functions is typically included within the system controller, so that at least basic status information can be provided in case of faulty operation. However, such onboard diagnostics are limited by the computing resources of the turbine controller and the absence of instant access to a long-term historical database. The turbine continuously stores operational data logs (SCADA logs), which can be retrieved and transferred to a central data store. The use of such data stores for detailed performance analysis and diagnostic work is becoming standard in the wind industry, since the data is readily available and provides information about a number of systems and components within the turbine. Today most medium to large scale operators of wind turbine fleets have installed centralized data management systems to collect and store such SCADA logs. Uptime Engineering has developed a software application that is capable of performing automated and continuous analysis of such data, typically with the aim of detecting anomalies in the behavior of individual turbines. Continuous advances have been made in the capabilities of the analytic models, and it is now possible to detect outlying behavior with a high degree of sensitivity. The results of such analysis are combined with the above-mentioned on-board diagnostic results together with general information concerning the turbine age, type and build status, in order to support the turbine operator in efficiently reacting to detected anomalies. However, such analysis activities often produce a high volume of information (multiple turbines monitored, multiple alarms originating from many systems and accompanied by a range of heterogeneous supporting information). The main challenge facing the user of such a system is efficient interpretation of the results and the derivation of an effective response strategy. Therefore a strong need has been identified to provide the software user with “decision support”; i.e., an additional layer of intelligence built in to the software, which combines all generated observations and produces clear recommendations for action. MBD is a highly relevant solution, due to the strong capability of the approach in combining state information from a multitude of sources and identifying the root cause with the highest likelihood.
2 Wind Turbine Fault Localization: A Practical Application of Model-Based. . .
29
The project between Uptime Engineering and Graz University of Technology aims at integrating an abductive MBD engine, created by the university, taking into account the process shown in Sect. 2.2.2 into Uptime Engineering’s wind turbine condition monitoring software. Uptime Engineering continuously extends and updates a comprehensive failure assessment of industrial wind turbines, providing a structured evaluation of faults and their manifestations. This analysis can be exploited by the MBD engine as the basis of the model development phase to construct a suitable diagnostic system description. Once an anomaly has been detected by Uptime Engineering’s condition monitoring software, an MBD computation is triggered taking into account of the abductive KB created and the symptom discovered. Given the results of the diagnosis, the MBD engine then provides additional information on the next best measurement based on entropy values. To ensure a suitable integration into the actual wind turbine maintenance workflow, we collaborate with an energy provider employing Uptime Engineering’s condition monitoring. In this section, we first describe the interface and interaction design of the diagnosis engine and how it will be incorporated into the work processes of the maintenance personnel of the energy provider. We then describe the phases of the integration and its current status.
2.3.1 Abductive Model-Based Diagnosis Prototype To enable MBD in industrial practice as proposed in Sect. 2.2.2, the necessary failure information must be available to automatically extract a suitable diagnostic model and an anomaly detection method is needed to initiate the fault identification phase. Furthermore, in order to yield benefits from deploying such a system, solutions need to be computed efficiently3 and effectively reflecting defects present in the system. We argue, however, that these technical features are not the only deciding factors determining the success of a newly integrated diagnosis software. While current research frequently focuses on developing and improving reasoning techniques, the suitable integration of MBD in operational processes is rarely addressed [24, 33]. In addition, it is well known that the acceptance of new technology is tightly linked to the perceived usefulness of the product as well as its perceived ease of use [7]. The former refers to the benefits for the users and other stakeholders in regard to the performance of work tasks, whereas the latter is on par with the usability of a product. Hence, in developing an MBD application for use in the field within our project, we focus not only on the technical aspects of feasibility but further account for the human factor. An interface and interaction design was incrementally developed
3 Here, efficiency is subjective to the application domain, e.g., in the context of wind turbines deriving explanations in minutes is sufficient, while for automotive on-board diagnosis this computation time is unacceptable.
30
R. Koitz et al.
for an abductive MBD engine, which should function as a template for the actual implementation of the tools which will be integrated into Uptime Engineering’s software. Various prototypes were created iteratively, starting from a low-fidelity paper mock-up to a clickable prototype depicting a usual fault identification scenario. These prototypes reflect the above-described general process of abductive MBD in the context of wind power plants. Particular attention was paid to respecting current work processes and accounting for a usable design. The design process started with eliciting the requirements of the diagnosis application in consideration of the stakeholders involved in the project, who were: • the service technicians, who are the users, will operate the diagnosis software for troubleshooting from the service center as also in the field, and are responsible for performing the turbines’ planned maintenance, repair as well as replacement activities • the management of a wind energy provider, planning on extending their selfmaintenance activities for their wind turbine plants in the future • Uptime Engineering, who currently develops condition monitoring software for wind turbines and will extend their portfolio with usable and extendable diagnosis software
2.3.1.1
Requirements
A list of requirements in regard to the final diagnosis application was established during the course of the various design iterations. The three distinct stakeholder groups have differing requests, which were analyzed in order to resolve conflicts and prioritize the resulting requirements. Since the success of the application depends to a great extent on being used by the service technicians, special attention was given to their suggestions and needs. An important observation is that current fault detection activities performed by the service personnel typically rely on visual inspection. Hence, in order to support diagnosis, images should be used for easier recognition. Once a fault has been identified, the repair or replacement task is executed according to the wind turbine manufacturer’s instruction manuals. Therefore, such documents need to be easily accessible via the software. After the maintenance activities have been completed, the service technicians are required to create a report of the task and the actions performed. The software should thus support automation of the reporting step to reduce the overall effort. The working environment inside a wind turbine is often uncomfortable and limited in space, and work is performed under time pressure in potentially difficult weather conditions. The user interface therefore needs to be intuitive in use and must guide the user through a strictly defined sequence with minimal user interactions. Considering the overall work process, the software should feature a desktop software part operated in the service center as well as a mobile application, which should be used within the turbine itself.
2 Wind Turbine Fault Localization: A Practical Application of Model-Based. . .
31
The management of the energy provider is interested in promoting digitalization as well as increasing the productivity and safety of their wind operations in use. On the one hand, the software should support the service technicians in preparing all spare parts and tools necessary before traveling to a wind turbine to ensure minimal downtime, while on the other hand given the hazardous environment in the field it should support the safety processes already in place, e.g., the service technicians personal safety equipment. In addition to the user and management requirements, the specifications of Uptime Engineering needed to be satisfied. To extend and update the knowledge base, i.e., abductive model, the users should be able to report new fault modes, which have not been previously contemplated. Further, the user interface should be extendable and adaptable to satisfy other customers as well as other domains for future projects.
2.3.1.2
Design Process
To ensure a user friendly end product, the diagnosis engine GUI was developed using an iterative process [25]. Each iteration starts with a definition or adaptation of the requirements, a design is then created and subsequently a prototype is implemented. This prototype is evaluated by users from the target group to determine usability issues, which must be fixed in the design of the proceeding iteration. According to Nielsen [25] due to the various repetitions of this cycle, this type of design process allows gaining sufficient insight into usability issues even given a limited number of test users. In our case, the first iteration was kicked-off with a meeting between Graz University of Technology and Uptime Engineering to elicit the first set of requirements. One of the main goals identified was that the software should be designed in a way that supports the service technicians’ current work processes without causing additional effort. Facilitating the service personnel’s work tasks is essential as this assures usefulness, which is a key aspect in technology acceptance [7]. Furthermore, we defined the overall workflow for the application, the general structure for an initial paper mock-up, and a small set of features, which should be realized. The initial paper prototype and all following designs were evaluated at meetings with the management of the energy provider and service technicians. At these meetings the current prototype was presented and a more detailed knowledge about the users and their work process was gained, usability issues could be uncovered and useful features, which would aid the maintenance personnel throughout the fault correction process, were identified. During the first iterations predominant usability problems were detected and some more drastic changes to the design were introduced, while in the later iterations only minor issues were found and as a result only slight GUI alternations were necessary. The product of the design phase is a clickable prototype which has undergone a small-scale qualitative usability test involving five service technicians. In the test scenario the users performed a mock-up fault identification process from start to finish. The design of the final prototype is presented below together with the overall application workflow.
32
R. Koitz et al.
Service Center Uptime Employee Engineering Diagnosis Condition Engine 1.send alarms 2.triggers 3.provides Monitoring diagnosis
diagnoses
5. prepares tools, spare parts and safety equipment at service center 6a.performs part inspection at turbine, adds additional measurements
7.repairs/replaces faulty component(s)
4.selects work assignments based on diagnosis results
6b.manually restarts diagnosis
Service Technician
8. creates report of maintenance activities
Fig. 2.2 Workflow of the diagnosis application (adapted from [21])
2.3.1.3
Workflow and GUI Design
The workflow of the diagnosis application was created in consideration of the current functionality of Uptime Engineering’s condition monitoring software, the maintenance process of the energy provider, and the general abductive diagnosis procedure. Figure 2.2 depicts the identified activity sequence of the diagnostic process. The interface and interaction design decisions of the application were taken based on the workflow and requirement analysis. As mentioned in the previous section, a diagnosis computation is invoked once an anomaly has been encountered. Each wind turbine includes a set of sensors and a basic on-board system that triggers alarms whenever measurements fall outside certain limits (Step 1 in Fig. 2.2). Uptime Engineering’s condition monitoring software extends and refines the fault detection by further processing the available sensor information. Once a symptom of a faulty turbine has been identified, the Uptime Engineering’s software triggers the root cause identification by supplying the previously created system description as well as the observations to an MBD engine (Step 2 in Fig. 2.2). After the computation, the results are accessible to the employees at the service center (Step 3 in Fig. 2.2). The diagnosis results are displayed as part of Uptime Engineering’s web interface, i.e., at the Operations Center, which is depicted in Fig. 2.3 and designed for desktop or laptop computers. The results are available per turbine instance and displayed as collapsible panels. For each triggering symptom,
2 Wind Turbine Fault Localization: A Practical Application of Model-Based. . .
33
Fig. 2.3 Operations Center [21]
e.g., Error Converter Bus, the panel contains the possible root causes4 as well as diagnosis likelihood expressed as percentages. Based on the outcome, the service center employee can create and assign repair tasks for the service technicians preselecting some of the possible faults for consideration during the field work (Step 4 in Fig. 2.2). Each repair task is either preformed in conjunction with the next planned maintenance, scheduled, or immediately executed. Figure 2.4 depicts an example for a scheduled repair tasks, consisting of the anomaly and the corresponding error codes. The service center employee can then define a trouble shooting task, schedule the activity under consideration of the time table depicting the availability of service technicians, assigning both a supervisor and a team for the task, add the corresponding parts and tools to the work assignment depending on the failures proposed by the engine, and provide additional auxiliary information to the work assignment such as previous issues with the targeted wind turbine. Several repair tasks can be scheduled for the same day and the same maintenance team. In the context of diagnosis within the field, we concluded that the software would be most usable on a mobile device since the technicians prefer not to carry a laptop. Thus, once work assignments have been created, the rest of the diagnosis process is conducted by the service technician teams over a mobile application. An essential aspect of the software design, was to follow guidelines and best practices for mobile user interfaces to ensure an easy-to-use application. A flat navigation was thus chosen for the prototype featuring little nesting of sub-levels, and thus warranting minimal user interaction and proving a defined role in the work process of the 4 In the case of wind turbines, there is generally a strong single fault assumption. Thus, each depicted root cause in this example only consists of a single failure, e.g., IGBT module: Diode/IGBT wire bonding—TMF. Yet, the diagnosis engine is of course capable of determining multiple fault diagnoses.
34
R. Koitz et al.
Fig. 2.4 Repair task screen
technicians. A simple navigation drawer is used to allow the user to switch quickly between the top-level sites. In Fig. 2.5a the Home screen of the mobile application is shown, where the technician can see all work assignments for the day as collapsible panels with additional information. Based on their tasks the technicians can obtain a list containing all necessary spare parts, tools, and safety equipment required for all maintenance activities planned on that day from the Preparation view depicted in Fig. 2.5b. The preparation would usually be performed at the service center, where the stockroom is also located (Step 5 in Fig. 2.2). Once at the turbine, an overview of the maintenance task for this particular instance and assignment is shown in the overview screen (see Fig. 2.5c), where an enforcement permit for the activity must be acquired.5 5
A notification for the person responsible for the entire installation is automatically generated containing the request. Only after the permission has been granted, may the technicians perform the maintenance work.