Skip to content

The life of an alarm ​

An alarm on the boat's screen has a short and entirely mechanical life: a device reports a fault, the report is held for 15 seconds, and it is either renewed or forgotten. Understanding those 15 seconds explains almost everything a crew member sees, including why an alarm sometimes seems to clear itself and why acknowledging it does not make it go away.

What a crew member needs to know ​

An alarm on the page is happening now. The device reporting it said so within the last 15 seconds. Nothing on this page is historical.

Nothing needs clearing. When the condition ends, the device stops reporting it and the alarm disappears from the page within 15 seconds, with no action from anybody.

Acknowledging silences the count, not the fault. It is a timer: by default the alarm stops being counted as active for one hour, then comes back if the device is still complaining. It never touches the equipment.

Nothing is remembered. There is no alarm history and no logbook here. If an alarm needs recording, it has to be recorded elsewhere while it is live.

The rest of this chapter is the mechanism behind those four statements.

How an alarm is raised ​

A MUXEN device decides on its own that something is wrong and starts broadcasting a fault code. muxen-alarms never decides that a device is faulty — with the two exceptions in Devices that go quiet, and the daemon's own health, it only relays what devices say about themselves.

The report reaches the daemon by one of two routes, chosen at startup:

RouteFlagWhat arrives
MQTT (shipped default)--mqtta message on device/<functionCode>/<instance>/error/<code>
CAN--interface canXthe DTC broadcast frame, identifier 170 (0xAA), 3 bytes

Both carry the same three things: which device, which code, and when.

  • On CAN, the device is the low 12 bits of the frame's CAN identifier, the code is the 16-bit errorCode field of the frame, and the time is the kernel's receive timestamp.
  • On MQTT, the device is rebuilt from the topic as functionCode * 64 + instance, the code is the last topic level, and the time is the payload's metadata.rxdate.

A message whose payload is not valid JSON, or which carries no metadata.rxdate, is dropped silently. That is the single most common reason for "the device is publishing but the alarm never appears"; see Troubleshooting.

Only the topic levels are read on the MQTT route. Beyond rxdate, nothing in the payload body affects the alarm.

Identity: one entry per device and code ​

The daemon holds a list of live alarms, keyed on the pair (deviceId, alarmCode). That has three consequences worth knowing:

  • A device reporting the same code ten times a second produces one entry, refreshed ten times a second.
  • The same code from two different devices is two entries — battery 0 and battery 1 raising "Low voltage" are two separate alarms.
  • Two different codes from one device are two entries.

The list lives in memory only. It is never written to disk, and restarting muxen-alarms.service empties it. It refills from the next round of reports, so on a boat with a live fault the page is repopulated within seconds.

The 15-second window ​

Every report sets the entry's expireTime to 15 seconds after the report's own timestamp. Once a second the daemon sweeps the list and removes every entry whose expireTime has passed.

creationTime behaves differently, and deliberately: an arriving report for an entry that is still live keeps the original creationTime. So creationTime answers "since when has this been wrong?", not "when did the last message arrive?". If an entry had already expired but had not yet been swept, the next report resets creationTime — a fault that genuinely stopped and restarted reads as new.

report ──► entry created,  creationTime = t0, expireTime = t0 + 15
report ──►                 creationTime = t0, expireTime = t1 + 15
report ──►                 creationTime = t0, expireTime = t2 + 15
 (silence)
        ──► 15 s after the last report, the entry is removed

The practical reading: expireTime minus 15 seconds is when the fault was last confirmed. A screen that wants to show freshness compares that against the clock.

The window is 15 seconds for device alarms on both routes. The two alarms the daemon raises itself use different windows — Devices that go quiet, and the daemon's own health.

Translation: code to sentence ​

The daemon does not know what any code means. It looks the pair (functionCode, alarmCode) up in the catalogue loaded at startup from /usr/share/muxen-alarms/dtc.json, and copies the matching entry — description, French description, severity, and any extra field it carries — into the published alarm.

When there is no matching entry the alarm is still published, with:

json
"alarmDescription": "unknown",
"severity": "unknown"

That is a deliberate choice: an untranslatable alarm is still an alarm, and hiding it would be worse than showing it without words. The alarm catalogue covers the catalogue and what to do about unknown.

Acknowledging: what a filter is ​

A filter is a request to stop counting matching alarms as active for a period. It is what an HMI's "acknowledge" or "mute" button creates.

A filter carries up to three constraints, and an omitted constraint matches anything:

FieldMeaning when presentMeaning when absent
deviceIdonly this boxany box
functionCodeonly this kind of equipmentany kind
codeonly this alarm codeany code
durationseconds the filter lasts3600 (one hour)

At least one of the three constraints must be present. A filter with none of them — which would silence the entire boat — is refused with 400. duration must be between 1 second and 365 days; anything outside that is also refused with 400, rather than being clamped.

Useful shapes:

jsonc
{"deviceId": 320, "code": 1}          // one alarm on one box
{"deviceId": 320}                     // everything that box reports
{"functionCode": 5}                   // every battery on the boat
{"code": 65000, "duration": 86400}    // offline alarms, for a day

What a filter does and does not do ​

A matched alarm is not removed. It stays in dtcs[] with:

  • "muted": true,
  • a filterBy object describing the filter that matched it,

and it is counted under muted instead of active. Everything else about it — creationTime, expireTime, the 15-second refresh, the disappearance when the device stops reporting — is unchanged.

A filter therefore changes presentation only. It sends nothing to the device, alters no threshold, and disables no protection. Whatever the equipment does about its own fault, it goes on doing.

When several filters match the same alarm, the first one created wins, and it is the one reported in filterBy.

Filters expire, and they do not survive a restart ​

Filters are swept once a second, exactly like alarms: when expireTime passes, the filter is deleted and every alarm it was hiding goes back to active on the next broadcast.

Filters also live in memory only. Restarting muxen-alarms.service clears every mute on the boat. So does a MUXEN deployment, since the unit is PartOf=muxen.target and that target is restarted wholesale. This is the safe direction to fail in — a restart can only make the boat noisier, never quieter — but it does mean a long mute placed before a software update will not be there afterwards.

The identifier is derived, not random ​

The field is called uuid, and it is not one. It is built from the three constraints as deviceId-functionCode-code, zero-padded to %03u-%03u-%06u, with an absent constraint written as 0:

FilterIdentifier
{"deviceId": 320, "code": 1}320-000-000001
{"functionCode": 5}000-005-000000
{"code": 65000}000-000-065000

Two things follow.

Re-acknowledging refreshes rather than duplicates. A PUT with the same three constraints finds the existing filter and resets its duration and creation time. Pressing acknowledge twice does not leave two filters behind, and there is no way to accumulate them.

A constraint whose value is 0 is indistinguishable from its absence.{"deviceId": 0}, {"code": 0} and {"functionCode": 0, "code": 0} all derive 000-000-000000, so the second PUT refreshes the first rather than creating a second filter — and the constraints that stay in force are the first filter's, not the ones just sent. Any filter whose present constraints are all zero shares that one slot.

This is reachable in practice: device id 0 is button panel instance 0, function code 0 is the button panel, and code 0 is a real alarm code on several equipment families. It only bites when every constraint in the filter is zero — {"deviceId": 320, "code": 0} derives 320-000-000000 and is unambiguous. Where a mute must target a zero-valued field, add a non-zero constraint alongside it.

Spoken alarms ​

On a Brain with an audio output the daemon says each alarm out loud, in English, with the voice files shipped in muxen-alarms-database.

When. An alarm is spoken as soon as it appears. After that, every five minutes (--audio-interval), the whole list of what is still active is spoken again: most severe first, oldest first, with two seconds (--audio-gap) of silence between two sentences. With --audio-interval 0 each alarm is spoken once and never repeated.

Which. Only alarms at or above the minimum level, error by default. The ladder is error above warning above notice, so out of the box the catalogue's warning entries stay silent and only its error entries speak. An alarm with no catalogue entry ranks as notice.

Acknowledging silences the voice too. A muted alarm is not spoken. When its mute expires and the device is still reporting, it is spoken again, as if it had just appeared.

Over other sounds. An error sentence plays through the muxen_alarm output, anything lower through muxen_notification. The platform's audio configuration defines those two names and decides what each does to the other sounds on board. The daemon never changes a volume. Where a name is not defined, the sentence plays through the default output.

Switching it off does not last. The voice can be turned off, or its minimum level changed, with a message on app/alarm/settings (Reference). Nothing about that is stored: any restart — a reboot, a deployment — speaks again with the defaults. That is deliberate. An alarm system that stays mute because somebody once silenced it is worse than one that speaks again.

Testing the speaker. muxen-alarms-test-audio, run on the Brain, makes the daemon play a short welcome clip through the alarm output — even when the voice is switched off, since the point is to check the speaker. A screen can ask for the same test over the WebSocket. The test never cuts off an alarm: it is ignored while anything is being said, and an alarm that falls due while it plays stops it (Reference).

It is never a dependency. A Brain without an audio output reports and displays alarms exactly as one with it; its audio state simply says available: false. The daemon finds out by playing a second of silence, at startup and then every 30 s for as long as that fails, so audio that comes up later is picked up; it speaks nothing until one plays (Reference). The same goes for a single sentence: one whose voice file cannot be decoded in full is not spoken at all, rather than cut off halfway, and the alarm is still displayed.

If nobody acknowledges ​

Nothing degrades. The alarm stays in active for as long as the device reports the fault, the WebSocket keeps broadcasting it every 5 seconds, and whatever a screen or an off-Brain client does about a non-zero active count, it keeps doing. There is no escalation, no timeout that turns an unacknowledged alarm into something worse, and no state that gets stuck.

The system is designed so that the only way to make the page quiet is for the underlying condition to end — or for somebody to make a deliberate, time-limited, self-expiring decision to stop counting it.

Integration of multiplexed solutions
MUXEN and the MUXEN logo are trademarks of MUXEN SAS.