Jump to content
Main menu
Main menu
move to sidebar
hide
Navigation
Main page
Recent changes
Random page
Help about MediaWiki
Special pages
Niidae Wiki
Search
Search
Appearance
Create account
Log in
Personal tools
Create account
Log in
Pages for logged out editors
learn more
Contributions
Talk
Editing
Fault management
Page
Discussion
English
Read
Edit
View history
Tools
Tools
move to sidebar
hide
Actions
Read
Edit
View history
General
What links here
Related changes
Page information
Appearance
move to sidebar
hide
Warning:
You are not logged in. Your IP address will be publicly visible if you make any edits. If you
log in
or
create an account
, your edits will be attributed to your username, along with other benefits.
Anti-spam check. Do
not
fill this in!
{{Multiple issues| {{more citations needed|date=October 2017}} {{technical|date=October 2017}} }} In [[network management]], '''fault management''' is the set<!--non-technical use, don't link--> of functions that detect, isolate, and correct malfunctions in a telecommunications network, compensate for environmental changes, and include maintaining and examining [[Computer glitch|error]] [[Data logging|logs]], accepting and acting on error detection notifications, tracing and identifying faults, carrying out sequences of diagnostics tests, correcting faults, reporting error conditions, and localizing and tracing faults by examining and manipulating [[database]] [[information]].<ref>{{Cite web|title = What is fault management? - Definition from WhatIs.com|url = http://searchnetworking.techtarget.com/definition/fault-management|access-date = 2015-10-06}}</ref> When a fault or event occurs, a network component will often send a notification to the network operator using a protocol such as [[Simple Network Management Protocol|SNMP]]. An alarm is a persistent indication of a fault that clears only when the triggering condition has been resolved. A current list of problems occurring on the network component is often kept in the form of an active alarm list such as is defined in RFC 3877, the Alarm [[Management information base|MIB]]. A list of cleared faults is also maintained by most [[network management]] systems.<ref>{{Cite web|date=2020-04-07|title=What Is Fault Management? A Definition & Introductory Guide|url=https://www.xplg.com/what-is-fault-management-2/|access-date=2020-11-15|website=XpoLog Log Analysis, Management & Viewer|language=en-US}}</ref> Fault management systems may use complex filtering systems to assign alarms to severity levels. These can range in severity from debug to emergency, as in the [[syslog]] protocol.<ref>RFC 3164</ref> Alternatively, they could use the ITU X.733 Alarm Reporting Function's perceived severity field. This takes on values of cleared, indeterminate, critical, major, minor or warning. Note that the latest version of the syslog protocol draft under development within the [[IETF]] includes a mapping between these two different sets of severities. It is considered good practice to send a notification not only when a problem has occurred, but also when it has been resolved. The latter notification would have a severity of clear. A fault management console allows a [[network administrator]] or [[system operator]] to monitor events from multiple systems and perform actions based on this information. Ideally, a fault management system should be able to correctly identify events and automatically take action, either launching a program or script to take corrective action, or activating notification software that allows a human to take proper intervention (i.e. send [[e-mail]] or [[Text messaging|SMS text]] to a [[mobile phone]]). Some notification systems also have escalation rules that will notify a chain of individuals based on availability and severity of alarm. ==Types== There are two primary ways to perform fault management - these are active and passive. Passive fault management is done by collecting alarms from devices (normally via [[SNMP]] traps) when something happens in the devices. In this mode, the fault management system only knows if a device it is monitoring is intelligent enough to generate an error and report it to the management tool. However, if the device being monitored fails completely or locks up, it won't throw an alarm and the problem will not be detected. Active fault management addresses this issue by actively monitoring devices via tools such as [[Ping (networking utility)|ping]] to determine if the device is active and responding. If the device stops responding, active monitoring will throw an alarm showing the device as unavailable and allows for the proactive correction of the problem. Fault management includes any tools or procedure for testing, diagnosing or repairing the network when a failure occurs. ==See also== *[[Alarm management]] *[[Alarm fatigue]] ==Notes== {{Reflist}} ==References== *{{FS1037C MS188}} [[Category:Network management]]
Summary:
Please note that all contributions to Niidae Wiki may be edited, altered, or removed by other contributors. If you do not want your writing to be edited mercilessly, then do not submit it here.
You are also promising us that you wrote this yourself, or copied it from a public domain or similar free resource (see
Encyclopedia:Copyrights
for details).
Do not submit copyrighted work without permission!
Cancel
Editing help
(opens in new window)
Templates used on this page:
Template:Cite web
(
edit
)
Template:FS1037C MS188
(
edit
)
Template:Multiple issues
(
edit
)
Template:Reflist
(
edit
)
Search
Search
Editing
Fault management
Add topic