Skip to content

Troubleshooting ​

A portal session can fail in a small number of ways, and several of them produce the same symptom. This chapter is organised by what you see.

Before anything else: a failed session never damages a gateway. The image is unconfirmed, the upload targets the secondary slot, and a power-cycle always brings the product firmware back. Diagnose at leisure.

open refuses immediately ​

Identification runs synchronously, so these come back before anything is written to the device.

MessageCauseAction
muxen-uds is not reachablemuxen-uds is down, its WebSocket has not come up yet, or it is older than 10.0.0 and sends no helloThe daemon waits briefly for the link on a cold start, so this means muxen-uds really is unavailable. Check its service; journalctl -u muxen-portal says muxen-uds 10.0.0 or later required when the version is the cause.
device not found on the busthe deviceId or UID does not answer the scanCheck muxen-portal list-targets, and prefer the UID — it is the stable identity.
readconfig failed — only products exposing HardwareId are supportedthe unit did not answer readconfigSee readconfig answers nothing below — a deviceId collision is the usual cause.
firmware does not expose HardwareIdthe product firmware predates the identification nodeBring the unit onto a current product firmware first. This is a deliberate scope limit.
unknown HardwareId <id> — no portal image for this boardthe board is not in the portal fleetSee Supported gateways.
serial gateway without a ProductIda combined serial board whose ProductId could not be readThe VE.Direct/Modbus tie-break needs it; fix the identification first.
no portal image installed for this boardthe .deb shipped without that imageReinstall muxen-portal; a source-only build produces a valid package with no images in it.
gateway still in portal firmware, reverting — retry in ~30 san orphaned portal from a lost session, and the reset did not landWait for the keep-alive to expire and retry.

The gateway is not in list-targets ​

list-targets lists only devices whose HardwareId the daemon resolves to a board. A gateway missing from it is either not a supported board, or did not answer readconfig. Three causes, in decreasing order of likelihood:

  1. A deviceId collision. Two devices presenting the same function and instance answer at the same deviceId, and neither can be read.
  2. A parked unit. Instance 63 answers a passive scan but will not serve UDS until activated, so it cannot be classified passively and is listed unclassified, with that as its reason. The daemon activates it itself at identify, so it opens normally by UID. Its parking deviceId is shared with every other parked gateway on the bus and addresses none of them — open a parked unit by UID, never by deviceId.
  3. An unsupported or too-old firmware.

You can always open by UID, which bypasses the deviceId entirely.

readconfig answers nothing — rule out a collision first ​

This is worth checking before you diagnose anything else, because nothing in the portal can see the difference between a contended address and a dead segment.

Two gateways that both present, say, function 17 Generic IO at instance 0 answer at the same deviceId. While that lasts, readconfig on that id returns nothing at all, muxen-uds uid --instance fails, and neither unit reaches list-targets. The daemon's identify step cannot get a HardwareId, so open fails exactly the way a dead bus would.

Detection: muxen-uds uid --scan marks a contended address with a trailing X.

Confirming it is easy if one unit can be taken off the bus: the scan then shows a single device at that id, readconfig starts answering, and open proceeds normally. That is also how to run a session on one of a colliding pair without fixing the addressing first.

Clearing it properly is a fleet addressing job, and muxen-uds uid --instance has two preconditions worth knowing before blaming the tool:

  • It parks the target at instance 63 of its function and reads it back, so the parking address must be free. If devices with an unassigned instance are sitting there, every attempt times out until they are given real instances, one at a time.
  • It needs the target uniquely addressable, so it cannot repair the very collision you want gone.

A unit running a factory application does not accept an instance write at all, so plan on assigning instances while units are on their product firmware.

no READY from the portal firmware (upload rejected by MCUboot?) ​

The upload succeeded, the reset was sent, and the gateway never came back as a portal. The message points at MCUboot, and MCUboot is often innocent. In decreasing order of likelihood:

1. The MUXEN bus is genuinely dead or the gateway is off. The message is indistinguishable from this case. Check that other devices still answer on can0.

2. A deviceId collision (above) that made identification resolve the wrong device.

3. The unit's bootloader predates the portal image's swap mode. A board's swap algorithm is a property of the MCUboot binary programmed at manufacture. It lives below the executable slot and no firmware upload replaces it — uploads go to the secondary slot. So a portal image built for swap-using-offset is staged one sector into the secondary slot, with its trailer at the end of that slot; a bootloader still using the scratch algorithm reads neither location, finds both erased, resolves "no swap", boots the old image unchanged, and logs nothing. The upload reports success, the gateway comes back as its product firmware, and the session times out at "no READY" with no other symptom anywhere.

This has no in-band fix and no in-band diagnosis: reflashing MCUboot needs SWD and physical access, by design, and the bootloader generation is not readable over UDS. The nearest proxy is the product's SoftwareVersion from readconfig — a unit well behind the fleet is a candidate. If you have SWD access, dumping the bootloader region and reading its banner string settles it outright, and reflashing the board's factory -firmware-usine image (which carries a bootloader) fixes it.

Take a full-flash dump before erasing a configured unit: its application and its configuration do not survive a mass erase, and the unit comes back parked with factory defaults.

The upload dies at 99 % ​

firmware upload failed, with the transfer refused at the very last block. On a board using MCUboot's swap-using-offset mode this is the signature of a missing --slot-size in the firmware build: imgtool pads to the larger secondary-slot size and the --pad trailer lands past the end of the executable partition, which the device refuses to write.

This is a firmware packaging bug, not an operational one — it cannot happen with the shipped images. It matters if you build your own; see Gateway firmware.

The gateway is never at risk: the transfer targets the secondary slot, so the unit stays on its product firmware and answers readconfig immediately afterwards.

The session is active but nothing moves ​

Every counter reads zero, no error is reported, and the mirror is silent. This combination — healthy state, clean counters, silence — points nowhere by itself. Work through it in this order:

1. Is your observer filtering? A bare candump portal0 does not deliver CAN_ERR_FLAG frames. On a LIN mirror that hides every NO_RESPONSE and CHECKSUM record, which is most of what you are looking for. Use:

sh
candump portal0ch1,0:0,#FFFFFFFF     # data frames AND error frames

The far-end counter tells the two apart: if noResponse is incrementing in muxen-portal stats while your candump shows nothing, the frames are being dropped by the reader, not missing from the wire.

2. Is the channel a monitor? A monitor channel is counted by the gateway but its stream is paused and it has no PTY. stats shows it as monitor and flowPaused is true. Promote it:

sh
muxen-portal open-channel 2

3. Is anything reading the PTY? After ~3 s without a reader the daemon pauses the channel's stream deliberately. Attaching a reader resumes it.

4. Is the remote simply silent? On an RS-485 Modbus segment this is the expected state: the polling master was the gateway firmware the session just replaced. muxen-portal timing <ch> distinguishes "quiet because the master is gone" from "quiet because nothing is connected" — see Monitoring a session.

A port receives, but the data is garbage ​

stats shows a nonzero RX rate together with a high error rate, and the note "the configured baud is probably wrong". That is baudSuspect, and it is a real answer rather than a fault: something is transmitting on that line, at a rate other than the configured one.

sh
muxen-portal set-rate 9600 --channel 3

Re-reading the wire-timing histogram at the wrong rate is pointless — shape means nothing until the baud is right.

The session keeps dying and recovering ​

Repeated active → recovering → active cycles mean STATUS replies are being missed three times in a row. Causes, in rough order:

  • Something restarted the daemon. muxen-portal.service is PartOf=muxen.target, and a MUXEN deployment restarts that target wholesale — writing /etc/muxen/deploy.json is enough. A deploy during a session will end it.
  • The gateway CPU is saturated. Monitor-opening four UARTs at a high rate on a live or noisy line costs one interrupt per received character. The default monitor rate is 19200, which is not a problem; an explicit --rate 115200 on the slowest board is the one configuration that can get there. rxDropped climbing on the opened channel beforehand is the tell. The escape hatch is monitor-channels = false.
  • The bus is saturated, or the tunnel is competing with a firmware upload elsewhere.
  • A debugger is attached. Halting the CPU over SWD during a live session starves the 1 Hz PING past the three-missed detector within about 20 s. Dump before an open, or after it has failed.

Recovery is designed to be survivable: the mirrors linger for 180 s and applications see a data gap, not a dead file descriptor.

set-term is refused ​

Termination control is a property of the backend, not of the board. A board can carry termination-enable lines and still refuse the command, because the VE.Direct (RS-232) backend does not drive them — only the Modbus (RS-485) backend does. The refusal is local and immediate, with a clear message, and is not a session fault.

100 % data loss with a perfectly healthy session ​

If allow-fd is on and the data plane negotiated CAN-FD, this is the signature of a classic CAN node on the bus: it destroys every FD frame with an error flag while the classic control plane keeps working.

The bring-up probe bounds this to a few milliseconds, but a classic node hot-plugged mid-session degrades the bus for the ~3 s detection window before the automatic fallback to ISO-TP. The session, mirror and applications continue at reduced throughput; data in flight during the window is lost and counted.

If it recurs, leave allow-fd off. FD is a 100 %-U5-segment feature.

After a close, the product has not come back ​

close --wait waits for the verified revert. If it warns instead:

  • The gateway may still be swapping. The revert swap takes a few seconds and the daemon then waits for the next scan tick.
  • If the CLOSE was lost, the gateway's own keep-alive ends the session within ~30 s regardless.
  • A power-cycle always works, and is the guaranteed path.

If the gateway is still answering as a portal device (function 0x3E) with no session owning it, the next open on that target resets it automatically and proceeds.

muxen-detect says busy or exits 5 ​

Another process holds the PTY. Two readers split a byte stream, so the tool refuses rather than corrupt both captures. Close the other reader.

Exit 5 is also returned when the session is recovering: during it the mirror is alive but silent, which would be recorded as "the device went quiet". Wait for active.

Getting more detail ​

sh
journalctl -u muxen-portal -f          # daemon log, one line per state change
muxen-portal status --json             # full session object
muxen-portal stats --watch             # live per-port counters
muxen-detect --attach <mirror> -v      # per-stage audit detail on stderr
mosquitto_sub -t portal/session -v     # the retained session state

The daemon logs every state transition, every NAK with its translated reason, and every dropped or malformed frame with a counter.

FAQ ​

Is this going to break my gateway? No. The firmware a session uploads is deliberately never confirmed, so any reset at all — the close, a crash, a power-cycle, or simply the Brain going quiet — makes the bootloader put the original firmware back. There is no state in which the portal firmware survives a reset. Worst case, switch the gateway off for ten seconds.

Something on the boat stopped working while an engineer was connected. That is expected while a session runs, and it ends with the session. The gateway is running the portal firmware instead of the firmware that drives its equipment, so whatever hangs off it goes quiet for the duration. On a gateway that bridges a whole segment, everything behind it goes quiet the same way. muxen-portal open prints which product identities will go offline before it starts, and publishes them on portal/session so screens and alarms can say "under maintenance" instead of raising data-loss alarms.

How long does it take? Roughly 30 to 60 seconds to open — the firmware upload alone is about 16 seconds — and a few seconds to close and confirm the product firmware is back.

Do I have to be next to the gateway? No, that is the point of it. You need access to the Brain; the gateway can be anywhere on the boat. It becomes a physical job again only when a board's bootloader has to be reflashed, which needs SWD.

What happens if the Brain reboots, or my connection drops, mid-session? The gateway reverts on its own. It expects a PING every second and reboots itself — which reverts it — after roughly 30 seconds of silence from the Brain. There is nothing to clean up and no state to recover; the next open starts fresh.

Can I leave a portal open to keep an eye on a bus? No. A session is a deliberate maintenance act, not a monitoring mode: for as long as it runs, the gateway is not doing its job. Close it when the work is done.

Two of us want to work on different gateways at once. One session per Brain. A second open is refused with a session is already running. Close the first, then open the next. The protocol itself would support several at a time; the single-session limit is daemon policy.

Does a session change anything on the gateway — settings, calibration? No. Portal firmware deliberately refuses readconfig and writeconfig. Config storage is shared with the product firmware and changes there would only apply after a restart, which is exactly what reverts the portal. Session parameters — rate, framing, termination — live on the tunnel control plane and disappear with the session.

Do I have to flash my gateways before I can use this? No. Installing the package drops the signed images into the muxen-uds firmware store on the Brain and touches no gateway. A gateway receives firmware only during a muxen-portal open, and never keeps it.

I upgraded muxen-portal. Did that update the firmware in my gateways? No — it updated the images on the Brain. The next session uses the new one. The fleet is never mass-flashed, and no gateway carries portal firmware between sessions.

Why does it ask for a UID rather than the device number? Either works. The deviceId is what a scan shows; the UID is the identity that does not move. A UID still opens a gateway parked at instance 63, or one sharing a deviceId with another device — both cases in which the deviceId is useless.

How do I tell a quiet line from a dead one?muxen-portal stats reports what the gateway itself sees on each of its ports, opened locally or not, so a port whose RX count is moving is alive even when your terminal shows nothing. On a serial gateway muxen-portal timing goes further and describes the shape of the traffic. And on a Muxen-native RS-485 segment, quiet is the expected answer: the polling master was the gateway firmware the session just replaced.

Could somebody use this to get into the boat's systems? A session carries no authentication, deliberately, but it grants nothing new. Starting one requires access to the Brain, and anyone with physical access to the MUXEN bus can already reset devices and upload signed firmware over UDS. The barrier to running code on a gateway is the signature chain: MCUboot only boots images signed with that board's key, portal images included.

Tips ​

Prefer the UID to the deviceId. list-targets prints both. The UID is the stable identity and keeps working when the deviceId does not — a parked unit, or two devices contending for the same address.

Say what is going offline before you open. The warning on open names the product identities, and portal/session carries them for the interfaces. On a boat in use, that list matters more than the session does.

Do not deploy during a session. muxen-portal.service is PartOf=muxen.target, and writing /etc/muxen/deploy.json restarts that target wholesale. The session ends with the deploy.

Arm the error mask on every candump.

sh
candump portal0,0:0,#FFFFFFFF

A bare candump does not deliver CAN_ERR_FLAG frames, so remote bus errors — and on LIN every NO_RESPONSE and CHECKSUM record — are invisible. A failing bus then reads as an idle one.

Read stats before you open a terminal. Every channel the gateway advertises is monitor-opened and counted from the start of the session, so is anything talking on port 3? is answered without creating a single PTY.

On RS-485, measure before you transmit. muxen-portal timing <ch> with bursts > 0 on a Modbus channel means a foreign master owns that segment. bursts == 0 is the expected reading on a Muxen-native one, and is not a broken port.

Fix the baud before reading anything else. baudSuspect in stats means something is transmitting at a rate other than the configured one. Histogram shape means nothing until the rate is right, and set-rate is a millisecond re-OPEN, so hunting for it is cheap.

Close explicitly, and with --wait. The keep-alive is the safety net, not the procedure. close --wait verifies by scan that the product identity is back, which is the only positive confirmation there is.

Use --wait in scripts. open, open-channel and close all submit and return; without it a script races the bring-up.

Sort the bus addressing out first, on product firmware. A unit running a factory application does not accept an instance write at all, and muxen-uds uid --instance cannot repair a collision it has no way to address uniquely. Instances assigned once, while the units are on their product firmware, remove the single most confusing failure in this chapter.

Do not attach a debugger to a gateway during a live session. Halting the CPU over SWD starves the 1 Hz PING past the three-missed detector in about 20 seconds. Dump before an open, or after it has failed.

Leave allow-fd off unless the whole segment is U5. One classic CAN node destroys every FD frame on a shared bus. FD is a 100 %-U5-segment feature, and the fallback costs a detection window of lost data.

Budget a walk. muxen-detect --walk opens and closes a session per gateway — roughly 40 s of fixed overhead each, on top of the measuring window — and takes each gateway's products offline in turn. P pauses after the current gateway and q closes and reverts, so neither leaves a session open.

Power-cycle means fully off for about ten seconds. A brief dip is not guaranteed to clear all MCU state; the revert guarantee is stated on a true power-on reset.

A session on a factory-state unit proves the tunnel, not the product. Such a board answers with a real HardwareId and opens like any other, but it is not running its product function, so nothing you see through the mirror says anything about the equipment behind it.

Integration of multiplexed solutions
MUXEN and the MUXEN logo are trademarks of MUXEN SAS.