Outages, evidence and planned maintenance
Understand what each outage detector observes, what can confirm recovery, and how planned work and management-network uncertainty affect alerts and history.
Read an outage as an evidence-based finding
Cenovel records conditions it can observe through network management and server management controllers. An outage may identify lost contact, a missing stack member, a failed part or a capacity risk. These are useful operational signals, but they do not all mean users have lost service. A responding switch does not prove that traffic is forwarding, and a responding server controller does not prove that the operating system or an application is healthy.
Check the cause, affected device or part, observation time and supporting detail before deciding the response. Critical and Warning are triage priorities. They are not measurements of how many people are affected. Whether a site is a priority site comes from its site type, not from its name or notes. In Settings › Locations › Site types, an Administrator EX sets each type's outage urgency to Normal or Priority. A network outage that monitoring opens at a site whose type is Priority is raised to Critical, its alert is marked URGENT, and the alert names the site type as a priority facility.
Prepare the record and the management read
- Record equipment at the correct site and closet, with a usable management address. Keep stack membership and access-point connections accurate: several detectors compare current evidence with those records.
- Configure the applicable credential, trusted SSH key or server certificate, and permitted management networks. A successful login does not guarantee that every hardware table or command is supported or complete.
- Network lost-contact detection requires a successful monitoring baseline. Cenovel reports a switch as not answering only after it has read that switch successfully at least once; a retained discovery identity alone is insufficient. Server-controller lost-contact detection does not have that same prior-answer requirement.
- Viewing current outages, viewing history and downloading evidence are separately permissioned. Scheduling, editing or canceling maintenance requires current-outage edit permission and permission to write.
Network detectors
The following are the implemented causes and their default priorities. A single eligible assessment can open a condition; there is no common requirement for three failed checks or another fixed failure count. A detector runs only when the read supplies the evidence it needs.
Scroll the table sideways to read all columns.
| Detector and priority | Evidence and threshold | Recovery and interpretation |
|---|---|---|
| Switch not answering Cenovel · Critical | A previously observed switch stops answering with a failure classified as transport silence, such as timeout or no route. | A successful system read can clear it. Credential or trust problems do not establish network recovery. This identifies lost management contact, not its physical cause. |
| Stack member missing · Critical | A complete membership reading disagrees with the recorded roster. Where only counts are usable, the reported positive count is lower than the recorded count. Serial-only ambiguity is reported without naming an arbitrary missing member. | A complete, sufficiently identifiable roster must account for the member. Keep replacements and renumbering accurate in the record; an incomplete roster cannot clear the incident. |
| Power supply failed · Warning | A read identifies a failed supply. A supply the device reports as disabled counts as failed; testing, unknown or missing states count as not reported. Supported environmental output can also identify a failed supply. | Recovery needs a complete reading that shows the original identified supply healthy again. Omitted, unrelated, unknown or duplicate supplies cannot establish recovery. Management contact alone does not prove forwarding or redundancy. |
| Access point not answering Cenovel · Warning | A previously observed access point stops answering. A known upstream switch or stack/uplink incident may explain it and group it under that incident instead of opening another standalone outage. | A successful access-point system read can clear it. An answering upstream switch narrows the investigation but does not prove the access point itself is the root cause. |
| Uplink down · Critical | Complete interface and neighbour readings show a previously recorded neighbour missing from its local port, with that port administratively up and operationally down. | Every originally lost port remains a recovery obligation, even if later neighbour snapshots omit it or another port fails. Complete interface and neighbour readings must uniquely identify each port with known states and show it up, administratively disabled, or with a neighbour. Clearing does not prove restored forwarding. Malformed original evidence blocks automatic recovery. |
| Mist reports this device disconnected · Warning | A Juniper Mist read reports a linked device on a mapped Mist site as disconnected from the Mist cloud. A device with no status from Mist, an unlinked device or an unmapped site concludes nothing, and a failed Mist read opens nothing. | A later Mist read reporting the device connected closes it. This is Mist's report, not Cenovel's own reachability check; see the Juniper Mist guide. |
| Switch restarted · Warning | A supported uptime read reports a positive value lower than the previous positive value. | A subsequent complete uptime assessment without another decrease can clear it. Treat this as restart evidence, not a proven reboot cause; the detector does not independently exclude counter wrap. |
| Cooling failed · Warning | An environmental reading reports a fan in a recognized failure state, including failed, fault, critical, stopped, off, down or degraded. Unknown and absent are not failure states by themselves. | A complete environmental reading must positively identify the previously affected fan as healthy. Missing or renamed fan evidence can leave recovery unconfirmed. |
| PoE budget exhausted · Warning | Allocated power reaches or exceeds reported capacity, or a port reports power-deny, denied, faulty or faulted. Member-specific budgets are used when available. The separate low-headroom indicator begins at 85%; that alone does not open this outage. | Recovery needs complete usable budget evidence below exhaustion and known non-denied states for previously affected ports. Spare capacity on another stack member is not assumed transferable. |
| PoE will not survive a supply failure · Warning | At least two supplies are listed, some but not all are healthy, and allocated power exceeds estimated surviving capacity. The estimate divides total capacity equally between listed supplies. | Complete PoE and supply evidence must no longer meet the rule. Unequal supplies, power-sharing modes and vendor reserve policies can invalidate the estimate; it is a resilience warning, not proof of current or inevitable loss of service. |
Server-controller detectors
These five causes use supported Redfish hardware readings. Controller health describes hardware and reported conditions. It does not test a hosted application, user login or operating-system service.
Scroll the table sideways to read all columns.
| Detector and priority | Evidence and threshold | Recovery and interpretation |
|---|---|---|
| Server not answering Cenovel · Critical | The controller read returns a not-answering outcome. This can include transport failures and other errors that prevent a usable controller read. | A successful controller answer or explicit credential refusal clears lost-contact status. No credential or an unaccepted changed certificate does not clear it. Read the failure detail before concluding the server is powered off. |
| Server power supply problem · Warning | A supply reports failed or degraded health, including supported vendor configuration problems. Lost redundancy also qualifies when at least two supplies are installed and no individual failed supply already explains it. | The same identified supply must be read in a recognized non-failing state, or the redundancy group must report redundant. An explicitly absent supply can close its fault incident; that does not mean a replacement was installed. |
| Server fan failed · Warning | An installed fan reports Redfish Warning or Critical health. A speed reading alone is not the failure threshold. | The same identified fan must report OK or explicitly absent. A fan omitted from a partial collection is not confirmed recovered. |
| Server too hot · Critical | A sensor reports Critical health, or its value reaches or exceeds its reported upper critical or fatal threshold. There is no universal temperature in degrees used for every server. | The same sensor must provide a recognized non-critical reading. Warning can clear the critical incident while still contributing a hardware warning. |
| Server hardware warning · Warning | Read storage health is non-OK, a read temperature reaches a caution condition, or overall health is non-OK and non-unknown without a more specific part failure explaining it. | Clearing requires complete healthy system, storage and temperature evidence. A healthy system summary alone cannot clear an earlier warning when supporting collections are unread or capped. |
Review and confirm recovery
For example, a switch may answer while its fan command fails. Its lost-contact incident can resolve while its earlier fan incident remains open. Similarly, a controller may return three healthy fans and fail to read the fourth: the missing fourth fan is not automatically declared repaired.
Incident times are detection and confirmation times, not exact physical failure times. Outage duration excludes time held for management-network uncertainty. Site downtime merges overlapping incident periods instead of adding every device's minutes as if they were separate periods of site-wide disruption; it still does not measure actual user impact.
When a fresh successful read clears a switch or access-point lost-contact incident, its recovery entry can include the complete current uptime value and its source. The entry keeps a physical restart verdict unknown. A short uptime alone does not prove a reboot; a long uptime does not prove that only the network path failed. Stored fallback readings and unrelated hardware recoveries do not receive this contact-recovery annotation.
- Open the incident detail and compare the named device, member or part with the site record. Inspect the underlying failure and latest evidence rather than relying on the title alone.
- Check whether the required portion of the read completed. Positive fault evidence can still be useful in a partial read; missing data cannot be treated as a healthy replacement for an earlier fault.
- Check reported user impact separately. For an uplink incident, confirm the traffic path and redundancy. For a server incident, check the application through its normal service checks.
- After repair, obtain the relevant supported reading and verify the incident's recovery entry. A general successful connection is insufficient for an unread failed fan, supply or storage condition.
- Use history and available evidence downloads to preserve the timeline. If equipment was removed from Cenovel, distinguish the subject-removed closure from restored service. When a unit is swapped out with Replace, the old unit's open outages close as Replaced by the new serial on that date; for a stack member, only that member's outages close. That records the swap, not a repair of the old unit.
When Cenovel may have lost its own management path
Cenovel raises a management-network condition when at least two distinct addresses have transport-silence evidence across the current check and the recent 15-minute window, at least one was silent in the current check, and no genuine answer appears in either set. The failures need not all have occurred at exactly the same instant.
A successful read or explicit sign-in refusal counts as an answer for this condition. A changed key, certificate response, closed port or negotiation failure is treated as ambiguous for deciding whether the wider management path recovered. One isolated silent address cannot establish this shared condition.
While the condition is open, new switch and access-point lost-contact outages are withheld. Existing incidents of those two kinds are neither confirmed nor resolved, and their duration clocks pause. Hardware and server-controller incidents are outside this particular hold. A genuine answer releases the hold; a return within 15 minutes can remain the same management-network episode.
If Cenovel could not use its own stored credential and therefore never contacted a device, it reports a separate console-credential problem. Correct access through Monitoring rather than treating this as evidence that the device failed.
Schedule maintenance and keep the history
Maintenance suppresses alerts while monitoring continues and incidents remain recorded. Covered outages are excluded from ordinary current-outage lists, Today, actionable incident summaries, morning-report incident sections and ordinary outage totals. Moving an incident into maintenance is not reported as a recovery. The site's outage list retains covered records with maintenance context, while its outage statistics exclude them.
An incident still open when coverage ends becomes eligible for ordinary visibility and alerts again. An incident resolved while covered retains its maintenance classification in history. Ordinary evidence exports exclude maintenance incidents, so a site's visible history and an ordinary export can contain different rows. Maintenance is not a promise that a fault was repaired or that all of its duration will be subtracted from a later unsuppressed incident.
Scope matters. A site window covers everything monitored at that site, including devices added during the window. A closet window covers the switches and servers in that closet and the access points behind those switches; other closets at the same site still alert. A switch window covers that switch or stack, not access points behind it that fail on their own. An access-point window covers only that access point. A server or address window covers the device at that address. An access point that Juniper Mist reports is offered as Juniper Mist's disconnected report, and its window quiets that report. Ended and canceled windows cannot be edited.
A future window quiets nothing until it starts. When a window ends or is canceled while an outage it covered is still open, the outage returns to Outages and its alert is sent once, even if several parts of Cenovel notice the end. That includes an alert that was waiting to be sent when the window began. An alert already delivered before the window is not sent again, and an outage that cleared during the window sends nothing.
- Choose what the work affects. Only things Cenovel monitors are offered, each once: a site, a closet, a switch or stack, an access point, a server, or a polled address. Inventory records without a management address raise no alerts, so they are not listed. Read the line under the choice: it says what the window will quiet and what will still alert.
- Enter a readable reason, start and end times, and the intended time zone. Add a work reference when useful. The end must follow the start, and a single window cannot exceed 180 days.
- Confirm the scheduled window before starting work. It becomes active at its start time and stops applying at its end time. Edit or cancel a current window when plans change; cancellation requires a reason.
- Use the site's outage history to review incidents observed during the work. After maintenance ends, inspect any still-open conditions and obtain recovery evidence.
Interpret protocol evidence conservatively
The official Interfaces MIB describes an interface's operational state, not an end-to-end application's availability. Missing neighbour advertisements can also reflect the neighbour protocol or its configuration. Cenovel therefore combines a recorded neighbour, complete reads and a down operational port for its uplink rule.
SNMP system uptime measures time since the management portion was initialized. Its TimeTicks representation wraps after approximately 497 days. A decrease can suggest a management-agent restart or a wrap rather than a whole-device reboot. Corroborate the event with device logs when the distinction matters.
The ENTITY state standard allows enabled to mean partially or fully operable; it does not prove redundant power. Redfish likewise distinguishes a component's health from the health of the containing system. A failed redundant supply can need attention while the system continues operating. Use the vendor's reported detail and actual service checks to establish impact.
Further reading
- RFC 2863: The Interfaces Group MIB
- RFC 3418: SNMP system uptime
- RFC 2578 section 7.1.8: TimeTicks
- RFC 4268: Entity State MIB
- DMTF Redfish Data Model Specification 2025.2: Status and HealthRollup
Limits of this guide
- This guide is based on the application source, automated checks and the linked protocol references. It does not claim validation against live production devices.
- Which detectors can run depends on supported commands, tables, controller resources and a trustworthy inventory baseline. A successful management connection does not mean every detector ran.
- There is no universal consecutive-failure threshold, recovery delay or user-editable threshold. Protocol retries are not an incident debounce policy.
- Some incident summaries state forwarding loss, downstream loss, restarts or future PoE impact more firmly than the evidence supports. This guide separates those inferences from what was observed.
- The PoE redundancy calculation assumes the listed supplies have equal capacity; it does not model vendor-specific power sharing.
- Server-controller reads use the first system and chassis and limit how much they collect. Failed, capped or paged collections cannot be assumed complete.
- Server not answering Cenovel can mean a usable controller read failed, not only that the network was silent. Server-controller incidents are not held by the switch and access-point management-network condition.
- Maintenance coverage follows the incident's recorded scope and location. An access-point window covers only that access point's own incidents; use the server's own entry, its address, its closet or its site to cover a server-controller incident.
- Maintenance flags mark covered incidents; they do not subtract every maintenance overlap from every incident duration. Site history keeps planned records, while ordinary outage exports leave them out.
- The Switch restarted detector relies on a drop in reported uptime and needs further validation.
- If the original uplink-recovery evidence is malformed, Cenovel will not close the incident automatically even after healthy readings. That state needs evidence repair or administrative removal; it does not confirm a continuing physical fault.
- The correction that leaves an address unmonitored until a first successful manual SSH read may not be in your installed version.
Reviewed against the current source code. Confirm the behavior and available actions on your installed release.