Skip to content

Troubleshooting ​

This chapter is organised by what you see. Start with what the daemon can tell you, which is very little: there is no status command, no control socket and no MQTT topic. The journal and a hand-run bus scan are the whole toolbox.

sh
systemctl status muxen-blocswap
journalctl -u muxen-blocswap -n 100
muxen-uds uid --scan-to-json /tmp/scan.json && cat /tmp/scan.json

That last command is the exact one the daemon runs every 30 seconds, and the file it writes is the daemon's entire view of the bus:

json
{
  "devices": [
    { "deviceId": 65, "uid": "27005E000B504256" },
    { "deviceId": 66, "uid": "2D005E000B504256" }
  ],
  "duplicate": false
}

Run it yourself and you see what the daemon sees. Nothing else it does depends on anything outside that file.

First, be sure the journal is telling the truth ​

The daemon writes to standard output and never flushes. Under systemd that output is fully buffered, so its lines reach the journal in blocks of a few kilobytes rather than as they occur. A startup block for a boat with a couple of dozen devices fits entirely inside one buffer.

One thing does arrive promptly: anything muxen-uds itself writes to standard error. That is the child's output, not the daemon's, and it is not buffered here. A journal holding muxen-uds complaints and no muxen-blocswap lines at all is the normal shape of a bad day.

An empty journalctl -u muxen-blocswap is not evidence of anything. Before concluding that the daemon is not working, stop the service and run it in a terminal, where output is line-buffered and immediate:

sh
sudo systemctl stop muxen-blocswap
sudo -u muxen /usr/bin/muxen-blocswap

Give it a minute — the first scan is 30 seconds in.

Nothing was deployed after a swap, and nothing was logged ​

The commonest report, and it has four distinct causes. Work through them in this order.

Likely causeWhat to checkFix
The feature is offFeature BlocSwap enabled missing from the startup blockset FeatureBlocSwap to "1" — see Getting started
The address is contended"duplicate": true in the scan fileresolve the collision; the daemon is locked until you do
The replacement is at a different addressthe deviceId is absent from devices[] in the scan filegive the spare the right instance, then deploy by hand
The device is not monitoredno Monitoring device <id> line at startupit is missing from devices[], or flagged VirtualDevice

The second one deserves its own section, because it is silent.

Two devices at the same address — the silent lock ​

While the bus scan reports that two different units are answering at the same deviceId, the daemon does nothing at all: no deploy, no reset, no state change, and no log line, for any device — not just the contended one.

That is deliberate. "The device at deviceId 65" is not a well-defined thing while two of them answer there, and writing a configuration to it would be a guess. The interlock is conservative to the point of locking when the scan result does not carry the flag at all.

It is also the hardest state to recognise, because the daemon looks exactly like a daemon that has nothing to do.

Detection: run the scan yourself and read the flag.

sh
muxen-uds uid --scan-to-json /tmp/scan.json && cat /tmp/scan.json

"duplicate": true is the lock. So is the absence of the key altogether: the daemon defaults to locked when the scan result does not say.

How it happens during a swap: a replacement that was configured for a different position on the same boat comes up alongside an existing unit of the same product and lands on its address.

Fix: give one of them the correct instance, using muxen-uds, so the scan resolves to one unit per address. Then the daemon unlocks by itself on the next tick.

The daemon logged has been changed, but the bloc is not configured ​

The deploy ran and failed, or did not run at all, and the daemon did not notice. It checks only that the command could be started; the command's own result is not examined. Having then recorded the new UID as the one it expects, it will not try again.

The usual mechanism: muxen-uds deploy performs its own bus scan before writing anything, and skips a device that scan does not see. The replacement was present when muxen-blocswap scanned and absent a few seconds later when muxen-uds did — still booting, badly seated, a connector that is marginal under vibration.

Check: the device answers now.

sh
muxen-uds uid --scan-to-json /tmp/scan.json

Fix: run what the daemon would have run.

sh
sudo systemctl stop muxen-blocswap
muxen-uds --device-id 65 deploy
muxen-uds --device-id 65 reset
sudo systemctl start muxen-blocswap

If it fails again, muxen-uds will say why — its output is the diagnostic, not this daemon's.

Every device suddenly goes become OFFLINE ​

All of them, within the same few minutes, on a bus that is plainly alive.

The daemon marks every monitored device offline at the start of each scan and un-marks the ones the scan reports. If the scan itself fails — muxen-uds cannot be run, cannot reach the bus, or its output cannot be read or parsed — the un-marking never happens, and after three such scans every device is logged offline.

What to checkMeaning
muxen-uds uid --scan-to-json by hand as user muxenif this fails too, the problem is the bus or muxen-uds, not this daemon
cmd: failed to read '<path>' in the journalthe scan file could not be read back — the scan wrote nothing
cmd: failed to parse the JSON in the journalthe scan file was not valid JSON
muxen-uds errors in the journalthe daemon discards the scan's normal output but not its errors, so they land here

Nothing is deployed while this lasts: a device that is never seen online is never compared against anything.

One device is permanently offline, the rest are fine ​

It genuinely is not answering at that address. The daemon logs become OFFLINE once, after three scans, and then waits — indefinitely, silently, with no repetition.

Confirm with a scan of your own. If the unit is on the bus but at a different deviceId, its instance is wrong: the boat's configuration expects it at one address and it is answering at another. That is a muxen-uds job, not something this daemon can repair — it has no instance assignment of any kind.

A device is re-deployed at every startup ​

The startup block shows an empty right-hand side for it:

Loading deployed.json
65 =>

/var/lib/muxen/deployed.json holds no UID for that deviceId, so the expected value is empty, matches nothing, and the unit currently at that address is treated as a replacement on the second tick.

This is self-correcting: muxen-uds deploy writes the UID into that file as part of the deploy, so the next startup has a value to compare and the device settles. If it recurs on every boot, the deploy is not completing — see the daemon logged has been changed above — or the file is not writable.

On a boat where devices carry hand-set parameters that are not in the deployment file, this is destructive rather than merely noisy: a deploy overwrites the device's configuration. Check for empty right-hand sides before turning the feature on.

The same device is deployed over and over ​

device <id> has been changed repeats every 30 seconds.

The daemon records the UID it just saw and compares against it next time, so a repeat means it is seeing a different UID at that address on each scan. Two devices alternating at the same deviceId is the shape that produces this — and the duplicate interlock does not catch it, because each individual scan sees only one of them.

Take several scans in a row and compare their devices[] entries. Meanwhile, stop the service: every cycle resets the device.

The service restarts every 30 seconds ​

Restart=always with RestartSec=30. Since the daemon never exits on its own — a failed startup still enters the main loop — a restart loop means it is being killed, not exiting. systemctl status gives the signal.

Note that a normal startup produces the same visible pattern in one respect: nothing appears in the journal. Check systemctl status for the uptime rather than inferring from the log.

systemctl stop takes a long time ​

The daemon shells out synchronously and does not handle signals while a child is running. A stop issued during a bus scan, a deploy or a reset waits for that command to finish. A deploy on a device with many parameters is the long case.

This is not a hang. Let it finish rather than escalating to SIGKILL mid-deploy.

The bus is busy every 30 seconds even though the feature is off ​

It is. The scan timer is installed unconditionally, so a Brain carrying this package runs a bus-wide UID scan every 30 seconds whether or not FeatureBlocSwap is set — the feature flag decides whether anything is done with the result, not whether the scan happens.

If that scan is in the way — during a firmware upload, a long deploy, or any diagnostic session driven from muxen-uds — stop the service for the duration:

sh
sudo systemctl stop muxen-blocswap

FAQ ​

I replaced a module. Do I need to do anything? No, provided the replacement is the same product and is set to the same position number as the one it replaced. Fit it, power up, and wait about a minute. The module restarts once by itself when the settings are applied.

How long does it take? Up to about a minute if the boat is live, and about a minute after the Brain finishes booting if the work was done with the power off. There is no way to make it happen sooner other than restarting the service.

How do I know it worked? By the module doing its job. The daemon publishes no status anywhere and there is no indicator on any screen. The journal line device <id> has been changed says the daemon started the work, not that it succeeded.

Can I replace two modules at once? Yes, if they are different products or different positions. Not if the second one arrives set to the same position as an existing module — the Brain then stops touching anything at all, deliberately, until the clash is resolved.

Will it wipe the settings on my new module? Yes. A deploy resets the module's configuration and writes the boat's version of it. That is the point. Do not hand-configure a spare expecting those settings to survive.

Does it install firmware? No. It writes configuration only. A replacement that needs a firmware update needs muxen-uds, run deliberately.

Nothing appears in the journal. Is it broken? Probably not. The daemon buffers its output, so a healthy one can look completely silent for a long time. Run it by hand in a terminal to see what it is doing.

Why did nothing happen when I fitted a brand-new spare? Most likely it is not answering at the address the boat expects. A spare that has never been commissioned sits at the parking address, not at the position of the module it is replacing. Give it the right instance first.

Can I turn it off? Yes — clear the FeatureBlocSwap setting in the boat's configuration. The daemon keeps running and keeps scanning the bus, but does nothing with what it finds. To stop the scanning as well, stop the service.

It deployed a module I did not touch. Why? Because it had no record of which unit was at that address. That happens on a device that was never deployed through muxen-uds. It is a one-off: once the deploy has recorded the UID, it settles.

Tips ​

Check the startup block before you enable the feature. Run the daemon by hand once, read every Monitoring device line and every => UID line. Any empty UID is a device that will be deployed — and therefore have its configuration overwritten — within a minute of the feature going live. That is the single check worth doing at commissioning.

Record the deviceIds of the blocs on the boat. They are the only handle you have in the field, they are in the startup block, and they are not obtainable from the daemon at any other time.

Bring the spare to the right instance before it joins the bus. The daemon matches on address and assigns nothing. A spare on the wrong instance either does nothing or, worse, collides with a working unit and locks the daemon silently.

Stop the service before any deliberate muxen-uds session. Firmware uploads, deploys, instance assignment, long scans: all of them share the bus with a scan this daemon runs every 30 seconds, and nothing coordinates the two.

Fit one bloc at a time. Detection is per-address and the interlock is global; a single bad address stops the whole mechanism for every device.

Do not rely on the journal as an audit trail. Buffered output means timestamps are not when things happened, and there is no record of a deploy succeeding or failing at all. If you need evidence that a swap was handled, check /var/lib/muxen/deployed.json — its UID for that deviceId is what the last successful deploy wrote.

Watch for repeats. One has been changed per swap is correct. A second one on the next tick is a symptom, not a retry.

Leave the feature on. Its cost when nothing is swapped is one bus scan every 30 seconds, which happens anyway.

Integration of multiplexed solutions
MUXEN and the MUXEN logo are trademarks of MUXEN SAS.