Appearance
Troubleshooting
The daemon has no control socket and no status command. Everything it knows it either prints to the journal or publishes on its own API, so those two are the whole toolbox:
sh
systemctl status muxen-alarms
journalctl -u muxen-alarms -n 100 --no-pager # the startup lines, and any errors
journalctl -u muxen-alarms -f # follow it
curl -s http://127.0.0.1:12001/api/dtcs # the list, right now
curl -s http://127.0.0.1:12001/api/filter # the mutes in force
mosquitto_sub -h 127.0.0.1 -t 'device/+/+/error/#' -v # what is actually being reported
muxen-dtc -f # the same, decoded
mosquitto_sub -h 127.0.0.1 -C 1 -v -t app/alarm/audio # the spoken alarms
mosquitto_sub -h 127.0.0.1 -C 1 -v -t app/alarm/info # version, features, onlineBefore anything else: nothing here is made worse by retrying. The daemon writes nothing to the bus and nothing to the broker but its own audio state and identity, mutes expire by themselves, and the alarm list rebuilds from live reports within seconds of any restart.
The alarm page is empty and I do not believe it
Check for alarm 65001 first, which is the same as checking that active is 0:
sh
curl -s http://127.0.0.1:12001/api/dtcsWhile 65001 is present the daemon can see nothing, and an empty list means nothing. If it is absent and the list is genuinely empty, the daemon is being fed and is reporting no faults.
If neither is true — no 65001, no alarms, and you are certain a device is complaining — work down An alarm is being reported but never appears, below.
65001 — "Alarm data source lost"
| Route | Cause | Check |
|---|---|---|
--mqtt (shipped) | the broker is not reachable on 127.0.0.1:1883 | systemctl status mosquitto; the journal has no main: mqtt connected |
--mqtt | the broker dropped the client | repeated main: mqtt connected lines after main: mqtt connection failed (rc = …) |
--interface canX | the interface is down, absent, or the controller has wedged | ip -details link show can0 |
The daemon retries by itself, doubling the delay from 1 second to a ceiling of 30, and says so:
source: broker not available yet, retrying
source: broker opened
source: can0 unusable, dropping the socket to rebuild itIt never exits over this. A daemon that is up and serving an API while raising 65001 is behaving correctly.
One case produces no alarm at all: if the CAN link monitor could not be started, the daemon logs source: can monitor unavailable on can0, bus health unsupervised and then deliberately stays silent rather than raise an alarm it cannot substantiate. On that Brain, 65001 is not available as an indicator on the CAN route.
An alarm is being reported but never appears
Work down this list; each step rules out the one above it.
1. Is the report on the broker at all?
sh
mosquitto_sub -h 127.0.0.1 -t 'device/+/+/error/#' -vNothing there means the problem is upstream — the device, or muxen-boat, which is what turns CAN frames into these topics. Confirm on the bus with candump can0.
2. Does the topic have exactly the right shape? The daemon matches device/<digits>/<digits>/error/<digits> and nothing else. The subscription is wider than the match, so a topic with an extra level after the code, or with a non-numeric segment, arrives and is discarded in silence. So is one whose numbers are out of range: the function code and the instance must be 0 … 63 and the code 0 … 65535, since a device id is functionCode * 64 + instance in 12 bits. Out of range they are dropped, not masked — device/192/57/error/1 would otherwise pass for an alarm of device 57, and error/65537 for code 1.
3. Does the payload carry metadata.rxdate? This is the common one. A payload that is not valid JSON, whose metadata is missing or is not an object, or which has no readable rxdate inside it, is dropped without a message. muxen-dtc -f applies the same checks — plus metadata.expireAfterSec and a freshness test — so a report visible in mosquitto_sub and absent from muxen-dtc -e is a malformed payload.
4. Is the Brain's clock sane? expireTime is derived from the report's own rxdate. A report timestamped 15 seconds or more in the past is added to the list and swept on the very next second, so it appears and vanishes faster than any screen refresh. muxen-dtc -e prints reports the freshness check would otherwise drop, which is how to confirm it.
5. Is it muted? A muted alarm is in dtcs[] with "muted": true and is counted under muted, not active. A screen showing only the active count will not display it.
sh
curl -s http://127.0.0.1:12001/api/filterAn alarm appears and disappears every few seconds
The device is reporting the fault intermittently, at intervals longer than 15 seconds. The entry expires between reports and is recreated by the next one — with a new creationTime, because the previous entry was already gone.
This is the system working as designed, and it is diagnostic: a fault that flickers on the page is flickering on the wire. Watch the raw topic to confirm the reporting interval.
Every alarm says unknown
The catalogue did not load. The daemon checks only that the file is readable at startup, not that it parses, so a corrupt or truncated dtc.json produces a daemon that starts cleanly and translates nothing.
sh
journalctl -u muxen-alarms | grep 'config: dtc' # which file it opened
jq empty /usr/share/muxen-alarms/dtc.json # does it parse
jq 'type' /usr/share/muxen-alarms/dtc.json # must be "array"A file that is not a top-level JSON array is rejected outright, with the same result. Reinstalling muxen-alarms-database restores the shipped file.
If only some alarms say unknown, the catalogue is fine and those codes simply have no entry — see The alarm catalogue.
65000 — devices reported offline that are not
65000 is raised for every device in /etc/muxen/deploy.json that did not answer the once-a-minute scan.
| Cause | Check | Fix |
|---|---|---|
| The device really is off or unplugged | muxen-uds -i can0 uid --scan | power and wiring |
| It is legitimately off most of the time | — | set DisableOfflineAlarm to "1" on it |
| It was never fitted | it is in deploy.json but not on the boat | correct the deployment file |
| It is a virtual device | it should carry VirtualDevice "1" | set it |
| The exemption was set as a number | the parameter's value is 1, not "1" | both parameters are compared as strings; quote them |
| Two devices share an address | muxen-uds uid --scan marks the id with X | a colliding pair answers as neither; fix the addressing |
The last one is worth ruling out early: an address collision makes both units invisible to the scan, so both are reported offline while both are perfectly healthy.
65000 never appears, even for a device that is off
The scan round failed and was abandoned. The journal names which way:
| Message | Cause |
|---|---|
offline: failed to start muxen-uds: … | muxen-uds is not installed, or not on the daemon's PATH |
offline: scan command failed (status = …), skipping | the scan ran and exited non-zero — usually no usable CAN interface |
offline: failed to parse the scan JSON | the scan produced something unexpected |
offline: failed to read '…' | the temporary scan file could not be read back |
Offline detection fails quiet: a broken scan raises no alarms rather than raising false ones. Silence from this feature is therefore not evidence that everything answered.
A scan that takes longer than a minute does not stack up — the next round is skipped instead, so the effective interval stretches rather than the scans overlapping.
The service restarts every 30 seconds
Restart=always with RestartSec=30, so anything that makes the daemon exit loops on that period. Two causes:
It printed its usage block and exited. A configuration error is reported by printing the usage and calling exit(0) — the exit status is 0, which makes systemctl status unhelpfully cheerful. The journal shows the usage text. Causes:
- neither
--mqttnor--interfacewas given; - the file named by
--dtcis not readable, which printsconfig: dtc file … not readable !just above the usage block.
The shipped unit passes --mqtt and no --dtc, so this needs a modified unit or a missing muxen-alarms-database.
main: lws init failed, exit status 1. The daemon could not bind its port. Something else is on 12001:
sh
sudo ss -lntp | grep 12001The API answers on 127.0.0.1 but not through nginx
The daemon binds the loopback interface only, on purpose, and is never reachable from the network directly. The two snippets in /etc/nginx/snippets/ have to be included by a site configuration; they do nothing on their own.
sh
curl -s http://127.0.0.1:12001/api/dtcs # the daemon
curl -s http://127.0.0.1/api/alarms/dtcs # through nginx
sudo nginx -tNote the two different paths: nginx rewrites /api/alarms/<x> to /api/<x>. A request to /api/alarms/dtcs on the daemon's own port is a 404, and so is /api/dtcs through nginx.
A path that is new in a snippet — /api/alarms/sound/… after the upgrade that added it — answering 404 through nginx means nginx is still running the previous snippet. The package reloads nginx itself at the end of every install and upgrade — dpkg takes the trigger from the version being installed, so that includes the upgrade that introduces it. What it cannot cover is a snippet copied by hand, or a Brain still on a snippet installed by an older version with nothing reloading since: sudo nginx -t && sudo systemctl reload nginx.
The WebSocket connects and then goes quiet
Three possibilities, in decreasing order of likelihood:
- Nothing is wrong. The daemon sends the full list at least every 5 seconds, whether or not it is empty, and the
audiostate every second, so a genuinely quiet feed is adtcspayload with"dtcs": []every 5 seconds and anaudiopayload every second. No payload at all is the symptom; the same payload repeatedly is not. - The client was culled. A client that cannot keep up fills the 16-slot send ring, and the daemon disconnects the slowest one rather than let it stall the broadcast to everybody else. The daemon logs
Killing lagging client. A client that reconnects in a loop under load is this. - A proxy timed out. The shipped snippet sets
proxy_read_timeout 15d; a different reverse proxy in front will have its own idea.
The socket is one-way. Anything a client sends on it is discarded; mutes are placed over REST, not over the WebSocket.
A PUT /api/filter is refused with 400
| Body | Why |
|---|---|
| not valid JSON | it cannot be parsed |
{} | at least one of deviceId, functionCode, code is required — a filter with no constraint would silence the whole boat |
{"duration": 0} | below the 1-second minimum |
{"duration": 40000000} | above the 365-day maximum |
only duration given | same as {}: duration is not a constraint |
duration is validated, not clamped: an out-of-range value is refused outright rather than quietly adjusted.
A body larger than 64 KB is refused with 413. The shipped nginx snippet caps the same request at 128 KB so an oversized body is stopped before it crosses the loopback.
A mute was placed but the alarm is still active
| Cause | Check |
|---|---|
| It expired | GET /api/filter — the default duration is one hour |
| The service restarted | filters live in memory only; a restart, or a MUXEN deployment, clears every one |
| The constraints do not match | compare the filter against the alarm's deviceId, functionCode and code |
deviceId was confused with instance | deviceId = functionCode * 64 + instance |
| It was absorbed by an existing filter | see below |
The last one is specific to this daemon. A filter's identifier is derived from its three constraints, not random, so a PUT whose constraints match an existing filter refreshes that filter's duration and keeps its constraints instead of creating a new one. Two filters whose present constraints are all zero collide on the identifier 000-000-000000. GET /api/filter shows what is actually in force.
A DELETE of a filter returns 404
The identifier is the derived one, not something the client chose. Read it back from GET /api/filter, or from the filterBy object on the muted alarm. It looks like 320-000-000001, never like a real UUID.
A request returns 405 rather than 404
The path exists but not for that method, and the response carries an Allow header listing what it does accept. /api/dtcs is GET only; /api/filter/{uuid} is GET and DELETE, with no PUT — a single filter is updated by PUTting the same constraints to /api/filter.
An alarm shows functionName: "Unknown"
The function code has no readable name in the shared MUXEN function table (muxen_function_name_from_id in libstdmuxen). The alarm itself is fine, and its description and severity are whatever the catalogue gave. Since libstdmuxen 9.4.1 every defined function code has a name, including thermal engine alarms (function code 28, "Thermal engine"), which read "Unknown" before; so "Unknown" now means a function code newer than the libstdmuxen this daemon was built with. Such an alarm carries functionName.<LANG> only if functions.json already has an entry for its code. Read functionCode rather than functionName when the distinction matters.
Alarms are not spoken
First tell the speaker from the alarms: muxen-alarms-test-audio plays a short clip through the alarm output, whatever the settings. If it is heard, the output works and the cause is in which alarms are spoken; if not, the journal says why the test did not play.
Then work down the audio state and the journal; the first line that fits is the cause.
sh
mosquitto_sub -v -C 1 -t app/alarm/audio
journalctl -u muxen-alarms | grep -E 'config: audio|audio:|app/alarm/(settings|test)'| What you see | Cause | Fix |
|---|---|---|
config: audio = off | started with --no-audio | remove it from the command line |
audio: cannot read …/manifest.json | the voice files are not installed | reinstall muxen-alarms-database |
"available": false since the start | no output check has played yet; it is tried every 30 s, with no limit | wait: it turns true within 30 s of the output working |
audio: no audio output (PCM …: …); alarms are not spoken, checking again every 30 s | the output check failed. Logged once; the daemon keeps checking | on a Brain without an audio output, nothing: that is expected. Otherwise check the output with aplay -L and that the audio drop-in of Reference is in place. No restart needed: audio: output checked, PCM …, audio is on follows within 30 s of the output working |
audio: off, the output check through PCM … did not answer | the audio output accepted the stream but never played it | restart the audio service of the platform, then systemctl restart muxen-alarms |
"available": false later on, and audio: cannot open PCM … | the output worked at startup and stopped | check the output as above; the next alarm or speaker test tries it again |
audio: the player is stuck | the audio output stopped answering the daemon | restart the audio service of the platform; alarms are reported and displayed meanwhile |
"enabled": false | somebody turned it off over MQTT | mosquitto_pub -t app/alarm/settings -m '{"enabled": true}', or restart the daemon |
ignoring a retained app/alarm/settings | a settings message was published retained | clear it: mosquitto_pub -r -n -t app/alarm/settings |
ignoring a retained app/alarm/test | a speaker test was published retained | clear it: mosquitto_pub -r -n -t app/alarm/test |
audio: speaker test from … ignored, audio is off | --no-audio, no readable manifest.json, or an output check stuck in alsa-lib | as the lines above |
audio: speaker test from … ignored, no audio output yet | no output check has played yet | see the no audio output line above; try again once available is true |
audio: speaker test from … ignored, an alarm is being spoken | the test never interrupts an alarm | wait for the round to end, and run it again |
audio: speaker test failed, no audio output | the alarm output would not open | as cannot open PCM above |
the alarm is a warning and "minLevel": "error" | below the minimum level, the default | mosquitto_pub -t app/alarm/settings -m '{"minLevel": "warning"}' for this run, or --audio-level warning |
the alarm has "muted": true | acknowledged: a muted alarm is not spoken | nothing, or delete the filter |
audio: no voice file for "…", not spoken | that sentence has no voice file | regenerate the voice files; see doc/internal/build-and-release.md |
audio: …: MPEG audio is not played (format 0x…), refused | an MP3 voice file | MP3 is never played; convert it to Ogg Opus, FLAC or WAV |
audio: …: decoding failed after N of M frames (…), not played | a damaged voice file | replace the file. A sentence is played whole or not at all, never cut off |
audio: …: truncated, the file is shorter than its header says, not played | a WAV cut short, typically an interrupted copy or upload | replace the file |
audio: PCM muxen_alarm is not defined, announcing through default | the platform does not define the alarm output | not a fault: the sentence plays, through the default output |
An alarm is spoken when it appears and then once per round, every --audio-interval minutes. Silence between two rounds is normal.
FAQ
An alarm is showing. Is it safe to acknowledge it? Acknowledging changes what the screen counts and nothing else. It sends nothing to the equipment, changes no threshold, and disables no protection. Whatever the equipment is doing about its own fault, it goes on doing. What acknowledging costs you is the reminder: for the next hour, that alarm no longer contributes to the active count.
What happens if nobody acknowledges it? Nothing changes. The alarm stays active for as long as the equipment keeps reporting the fault, and clears by itself when the fault ends. There is no escalation and nothing gets stuck. The risk of ignoring an alarm is whatever the equipment is complaining about, never the alarm system itself.
The alarm disappeared on its own. Did somebody clear it? Almost certainly not — there is no way to clear one. An alarm vanishes 15 seconds after the equipment stops reporting it. Either the condition ended, or the equipment was switched off, or it stopped talking. If it stopped talking and it is in the boat's configuration, a "device not detected" alarm appears within the minute.
Why is the same fault showing twice? Two devices, or two codes. Alarms are held one per device and code, so two entries mean two distinct reports. Compare the deviceId field: two batteries reporting low voltage are genuinely two alarms.
How long does it take for an alarm to appear? As long as the equipment takes to report it. The daemon adds it on arrival and the screens are updated at the next one-second broadcast.
Can I see what was wrong last night? Not from this system. It holds only what is true now, in memory, with no history and no logbook. Recording alarms is a separate job, done by something subscribing to the feed while the alarms are live.
We are doing work on the generator all week and it alarms constantly. What do we do? Mute it for a bounded period: a filter on that device, or on function code 4, with a duration in seconds. It expires by itself, so there is nothing to remember to undo. For equipment that is legitimately off long-term, DisableOfflineAlarm in the boat's configuration is the right tool for the "device not detected" half.
Why does everything say "unknown" after an update? The alarm code catalogue did not load. See Every alarm says unknown above; reinstalling muxen-alarms-database restores it.
Does muting one battery mute them all? Only if the mute was placed on the function code rather than the device. {"functionCode": 5} silences every battery; {"deviceId": 320} silences one. Check with GET /api/filter.
Do the mutes come back after a reboot? No. They are in memory and are cleared by any restart of the service, including a MUXEN deployment. That is the safe direction: a restart can only make the boat noisier.
The screen went quiet — is the alarm system still running?systemctl status muxen-alarms, and check that 65001 is absent. A silent page with no 65001 and a live WebSocket is a genuinely quiet boat.
Tips
At commissioning, work through 65000 first. Let the offline scan run for a couple of minutes and read the list. Every device reported missing is either genuinely absent, badly addressed, or wrongly listed in /etc/muxen/deploy.json. Clearing that list is what stops the crew learning to ignore the alarm page.
Set DisableOfflineAlarm on equipment that is meant to be off. A shore charger, a seasonal watermaker, an isolated generator. A permanent warning that everybody knows to ignore is worse than no warning.
Quote the exemption values. VirtualDevice and DisableOfflineAlarm are compared as the string "1". A JSON number is silently ignored.
Note which alarms read unknown at handover and get catalogue entries written for them. An unknown alarm is a code with no sentence, and the crew cannot act on a number.
Prefer narrow mutes. {"deviceId": …, "code": …} silences one alarm. {"functionCode": …} silences a whole class of equipment, and it is easy to forget it is in force for an hour.
Give a mute an explicit duration. The default is 3600 seconds. For maintenance work, say so — a duration matching the job means nothing has to be undone.
Do not mute 65001. It is the alarm that says the rest of the page is untrustworthy; hiding it hides everything.
Check the Brain's clock. Freshness is computed from the report's own timestamp against the Brain's clock, so a clock that has jumped forward makes live alarms look expired and one that has jumped back leaves them on the page.
Read active, not the length of dtcs[]. The array carries muted alarms too. A client that counts the array will still be shouting about alarms the crew deliberately silenced.
Each dtcs frame is the whole list, not a diff. It comes on connect, on every change and every 5 seconds; a client replaces its list with it. The WebSocket is where a client belongs; polling GET /api/dtcs gives the same object with more overhead.
-vv on a live boat is a firehose. It dumps every frame or message the daemon sees. Use it against a quiet bus, or for a few seconds.
Keep muxen-uds installed. Removing it does not break the daemon, but it silently removes offline detection — and the failure mode is no alarms rather than obviously wrong ones.
