Appearance
Devices that go quiet, and the daemon's own health
Almost every alarm in this system is a device complaining about itself. That leaves one blind spot: a device that has lost power, lost its wiring, or died outright reports nothing at all, and silence looks exactly like health.
Two alarms exist to close it. Both are raised by muxen-alarms itself rather than by any device.
| Code | Text | Severity | Means |
|---|---|---|---|
65000 | Device not detected on CAN-Bus | warning | a piece of equipment the boat is supposed to have did not answer |
65001 | Alarm data source lost | error | the alarm system itself cannot see the boat |
For a crew member
"Device not detected" means a box on the boat has stopped answering. The equipment is off, unpowered, or disconnected. It is not a fault inside the equipment — the equipment is not saying anything at all. Something switched off deliberately at the panel will produce this alarm, and that is normal.
"Alarm data source lost" is the serious one. It means the alarm page is not being fed. While it is showing, an otherwise empty alarm page proves nothing — the boat could be in perfect order or on fire, and this system cannot tell. Treat it as "the alarm system is out of service" and get it looked at.
How offline detection works
Once a minute the daemon:
- runs
muxen-uds uid --scan-to-json <tmpfile>as a child process — an active scan that asks every device on the MUXEN bus to identify itself; - reads the boat's configuration from
/etc/muxen/deploy.json; - raises alarm
65000against every device in the configuration that the scan did not find.
A 65000 entry lives for 90 seconds rather than the 15 of a device alarm, which is long enough to survive one missed scan without the alarm flickering off and back on between minutes.
The scan is spawned asynchronously, and a scan still running when the next minute comes round is not queued behind itself — one round is skipped instead. If muxen-uds cannot be started, or exits non-zero, or produces something that is not JSON, the round is abandoned with a line in the journal and no alarm is raised:
offline: scan command failed (status = 256), skipping
offline: failed to start muxen-uds: Failed to execute child process
offline: failed to parse the scan JSONNote the direction of that failure: a broken scan produces no 65000 alarms, not spurious ones. Offline detection fails quiet.
The configuration decides what "missing" means
65000 is raised strictly from /etc/muxen/deploy.json. A device that is on the boat but not in that file is never reported missing, and a device in that file that was never fitted is reported missing forever. Keeping the deployment file honest is therefore part of commissioning, not an optional tidy-up.
Two per-device parameters suppress the alarm:
| Parameter | Value | Effect |
|---|---|---|
VirtualDevice | "1" | the device is not real hardware; never scanned for |
DisableOfflineAlarm | "1" | real hardware that is allowed to be absent |
Both are compared as the string "1". A JSON number 1 does not match and the exemption silently does not apply — which is the usual cause of "I disabled the offline alarm and it is still there".
DisableOfflineAlarm is the setting for equipment that is legitimately switched off most of the time: a generator that is isolated at sea, a shore-power charger, a piece of seasonal kit. Without it, that equipment produces a permanent warning that trains the crew to ignore the page.
How the source-lost alarm works
65001 reports on the daemon's own input, and it is checked once a second.
In the shipped configuration (--mqtt) it is raised whenever the mosquitto client is not connected to 127.0.0.1:1883. On a Brain started with --interface canX it is raised whenever the CAN link monitor reports the interface as unhealthy — down, deleted, or a controller that has wedged.
Each check refreshes the alarm for 5 seconds, so it clears within a few seconds of the source coming back, and one late timer tick does not make it flicker.
The alarm is attributed to the daemon itself, at function code 9 (Display), instance 0 — device id 576. Those are compile-time values with no command-line override, so 65001 always arrives as device 576.
The daemon does not give up
An absent data source is not a startup failure. If the broker or the interface is not there when the daemon starts, it does not exit: it retries with a backoff that doubles from 1 second up to 30 seconds, and raises 65001 in the meantime.
That is deliberate. The REST API and the WebSocket work perfectly well without a bus, and a daemon that exited would take them down and leave the screens with a dead endpoint and no explanation. This way the screens keep their connection and are told, in the alarm list itself, exactly what is wrong.
On the CAN route there is one extra behaviour: a socket bound to an interface that is then deleted and recreated never recovers, so when the link monitor says the interface has gone bad the daemon drops the socket and rebuilds it on the same backoff rather than waiting for a recovery that cannot come.
If the link monitor itself cannot be started — no netlink socket — the daemon logs it and then stays quiet rather than raising an alarm it cannot substantiate:
source: can monitor unavailable on can0, bus health unsupervisedBoth are ordinary alarms once raised
65000 and 65001 differ from device alarms only in where they come from. Once in the list they behave identically: they appear in dtcs[], they are counted in active, they are translated through the same catalogue, and they can be muted with the same filters.
sh
# stop counting offline alarms for a day, during a refit
curl -X PUT http://127.0.0.1:12001/api/filter \
-H 'Content-Type: application/json' \
-d '{"code": 65000, "duration": 86400}'Muting 65001 is possible and is almost always the wrong thing to do: it silences the one alarm that says the rest of the page is untrustworthy.
Both codes are catalogue entries with no functionCode, which makes them wildcards — they resolve to the same text whichever device id they arrive against. See The alarm catalogue.
