Skip to content

Swapping a bloc ​

This chapter is the field procedure. It says what to do when a module has to be replaced, what the Brain does on its own afterwards, and how long each part takes.

For the crew ​

If a MUXEN module has failed and a spare is aboard, fitting the spare is the whole job. Switch the circuit off, unplug the failed module, fit the replacement, switch back on, and wait. Within about a minute the boat puts its own settings back onto the new module and restarts it. Nothing has to be typed anywhere.

Two things are worth knowing while you wait:

  • The new module restarts once, by itself. Whatever it drives goes off and comes back. That is the configuration being applied, not a fault.
  • Fit one module at a time. If two identical modules are on the boat and the new one arrives with the wrong address set, the Brain stops rather than guess which is which — and then nothing happens at all.

The rest of this chapter is for whoever fits the module.

Before you go ​

The one thing that has to be right is the address — the deviceId, which is the module's function code and its instance number. The daemon recognises a swap as the same address answered by a different serial number. A replacement that comes up on a different instance is not a swap; it is an unknown device, and the boat's configuration has no entry for it.

So, before the visit:

  • Note the deviceId of the unit being replaced. The journal lists them at startup, one per line — see Getting started.
  • Confirm the spare is the same function — the same product family. Nothing checks this. A different product answering at the same address would be deployed the boat's configuration for the old one.
  • Confirm the spare's instance matches, or plan to set it. This daemon does not assign instances; that is a muxen-uds job, done before the spare joins the bus.

The hot swap — the bus stays live ​

The Brain is running and watching. The sequence, with the daemon's own 30-second tick as the clock:

What happensJournal
tThe old bloc is unplugged—
t+0…30 sA scan finds nothing at that deviceId—
Up to two more scans may pass with it absent—
The new bloc answers at the same deviceIddevice <id> has been changed
muxen-uds --device-id <id> deploy—
muxen-uds --device-id <id> reset—
The bloc reboots and comes back configured—

Detection is on the first scan that sees the new unit, not on a timeout. If the replacement is in place and answering by the next tick, the deploy starts within 30 seconds of the swap.

An absence of three consecutive scans — 90 seconds or more — moves the device into a "waiting" state and logs device <id> become OFFLINE. That is not a failure. The comparison when it returns is the same one; only the log line differs, and you get a matching device <id> become ONLINE if the same unit comes back.

Both paths converge on the same two commands.

The cold swap — the boat is dead ​

The blackout case: the Brain is off, or the whole boat is, while the module is changed. Nothing observes the transition.

What saves it is that the expected UID is not held only in memory. At startup the daemon reads /var/lib/muxen/deployed.json and takes the UID recorded there for each monitored deviceId. That file is the record of what was last deployed where, and it survives the power cut.

So on the way back up:

What happens
Brain boots, daemon startsReads the device list, then the last-known UID for each
+30 sFirst scan. Devices that answer are marked present — no comparison yet
+60 sSecond scan. Now the UIDs are compared, and the swap is found

A cold swap is therefore detected about a minute after the Brain finishes booting, not immediately. The first tick is deliberately a settling pass: it establishes that the device is on the bus before anything is judged about which device it is.

What the deploy actually does ​

Two commands, run in order as ordinary shell commands, with the option before the subcommand:

sh
muxen-uds --device-id <id> deploy
muxen-uds --device-id <id> reset

The daemon does not inspect what deploy writes. Everything about the content of the configuration — which parameters, from where, in what order — belongs to muxen-uds and to the boat's deployment file. What this daemon contributes is when and to which address.

deploy also refreshes /var/lib/muxen/deployed.json, which is why a swap the daemon has handled is not detected all over again after the next reboot: the file now names the new unit.

After the reset the daemon records the new UID as the one it now expects and returns the device to its settling state, exactly as at startup: one tick to see it come back, then normal watching resumes.

The daemon is blocked while all this runs. It shells out synchronously: no scan, no timer and no signal is handled until each command returns. A systemctl stop issued mid-deploy therefore takes effect only when the deploy finishes.

Timings ​

StepDuration
Scan interval30 s, fixed
One bus scana few seconds — muxen-uds listens for a fixed window
Offline threshold3 consecutive scans, so 90 s or more
Hot swap, best casedetected on the first scan that sees the new unit
Cold swaptwo ticks after the daemon starts, so about 60 s
Deploy and resetseconds to minutes, depending on how many parameters the device carries

None of these are configurable. There is no command line, no environment variable and no configuration key that changes any of them.

What stops it ​

Three conditions stop a deploy that would otherwise happen. Two are deliberate; the third is the interesting one.

1. Two devices at the same address. The scan reports this, and while it holds the daemon runs nothing for any device — not the swapped one, not the others. Nothing is logged. This is the state to suspect when a swap is plainly visible on the bus and the daemon is not reacting.

2. The feature is off. Without FeatureBlocSwap set to "1" in the boat's configuration, no device is monitored at all.

3. The replacement never answers at that deviceId. It is then just an address that has gone quiet. After three scans the device is logged become OFFLINE and the daemon waits, indefinitely and without further logging, for something to answer there.

The daemon does not check that the deploy worked ​

This is the one thing to know before trusting the automatic path on a boat you are about to leave.

The daemon checks only that the muxen-uds command could be launched — that the process started and was not killed by a signal. What the command reported is not examined. A deploy that ran and failed is followed by the reset anyway, and the new UID is recorded as the one now expected. From the daemon's point of view the swap is done.

The realistic way this happens: muxen-uds deploy performs its own bus scan before writing anything, and skips a device that scan does not see. The replacement was present when muxen-blocswap scanned, and absent — still booting, badly seated, a marginal connector — a few seconds later when muxen-uds scanned. Nothing is written, nothing is logged by this daemon, and it will not try again, because as far as it is concerned the unit at that address is now the expected one.

The symptom is a bloc that is on the bus and answering but behaving as if it were factory-fresh. The fix is to run the two commands by hand.

There is one asymmetry worth knowing. A swap detected while the daemon still considered the device present — the fast path, where the replacement appeared before three scans had missed it — records the new UID before running the commands, so a launch failure there is never retried. A swap detected after the device had gone offline records it only afterwards, so a launch failure there is retried on the next tick. Neither path retries on a command that ran and failed.

Doing it by hand ​

If the automatic path is not available — the feature is off, the address is contended, or the replacement had to be given a different instance — the same result is obtained by running the two commands yourself on the Brain:

sh
sudo systemctl stop muxen-blocswap
muxen-uds --device-id 65 deploy
muxen-uds --device-id 65 reset
sudo systemctl start muxen-blocswap

That is precisely what the daemon would have run. Doing it by hand is also the way to recover a swap the daemon has already decided it has handled; see Troubleshooting.

Stop the service first. It scans the bus every 30 seconds through muxen-uds, and nothing coordinates two muxen-uds invocations sharing a bus. The overlap window is small but it is real, and a deploy is the worst moment to lose a response.

Note the argument order: --device-id comes before the subcommand. muxen-uds deploy --device-id 65 is not the same command.

After the swap ​

The daemon publishes nothing, so there is no "swap succeeded" indicator anywhere. Three checks, in the order that costs least:

1. The module is doing its job. The outputs it drives, the values it reports, on whatever screen normally shows it. This is the only check that proves the configuration actually landed.

2. /var/lib/muxen/deployed.json names the new unit. Its uid for that deviceId is what the last successful deploy wrote, so it is the closest thing to a receipt.

3. The journal shows the change once.

sh
journalctl -u muxen-blocswap -n 50

Expect device <id> has been changed once and nothing after it. A second one 30 seconds later means the device is being re-deployed in a loop — see Troubleshooting.

The journal is the weakest of the three, because the daemon's output is buffered and can be a long way behind. An empty journal proves nothing either way.

Integration of multiplexed solutions
MUXEN and the MUXEN logo are trademarks of MUXEN SAS.