Appearance
Sessions
A portal session is the interval during which a gateway runs portal firmware instead of its product firmware. It is opened deliberately, ended deliberately, and ends by itself if the Brain stops talking.
Nominal flow
muxen-portal open <deviceId|uid> --rate 250000
│
├─ 1. Identify via muxen-uds (scan / uid / readconfig): record the UID,
│ function/instance and the identification strings; select the portal
│ image by HardwareId — plus ProductId on the two combined serial
│ boards. An unknown board is refused here, before anything is written.
├─ 2. muxen-uds `firmware` → upload into the MCUboot secondary slot
├─ 3. muxen-uds `reset` → MCUboot swaps in test mode → portal firmware boots
├─ 4. Wait for READY on the parking control pair (timeout 20 s → abort)
│ READY carries the UID and must match the recorded target
├─ 5. SET_INSTANCE(uid, instance) → the portal re-binds its tunnel
│ endpoints; wait for HELLO (timeout 5 s → abort)
├─ 6. OPEN(type, rate, mode/cfg) → OPEN_ACK (timeout or NAK → abort)
├─ 7. Create the mirrors, start bridging, start the 1 Hz PING
│
│ ── active: the mirror is usable, STATUS is polled every second ──
│
muxen-portal close
│
├─ 8. Destroy the mirrors (stop accepting TX)
├─ 9. CLOSE → the gateway reboots → MCUboot reverts to the product image
│ (fallback: a UDS reset on the portal device, sent by the daemon)
├─ 10. Verify the revert: a UID scan shows the same UID back under the
│ recorded product function/instance, and the portal 0x3E entry gone
└─ 11. Report the session statisticsIdentification is synchronous, so a refusal comes back immediately. Everything from the upload onward runs in the daemon; open returns at once and the session is observed through muxen-portal status or the portal/session MQTT topic. --wait on open, open-channel and close turns the submit into a blocking call, which is what a script wants.
Bring-up is typically 30–60 s, dominated by the ~16 s firmware upload. A verified close takes a few seconds: the MCUboot revert swap itself is 2–4 s, and the daemon then waits for the next scan tick to confirm the product identity is back.
What the gateway looks like during a session
While a session runs the gateway is visible to muxen-uds as a portal device — function code 0x3E, instance assigned by the daemon at bootstrap — not under its product identity.
- Available: identification, routines, and remote reset (which ends the session and reverts).
- Not available: everything product-specific, and
readconfig/writeconfig. Config storage is shared with the product firmware and changes only apply after a restart — which would revert the portal — so the portal firmware deliberately does not touch stored settings.
Session and bridge parameters live on the tunnel control plane instead, driven by set-rate, open-channel, set-term, and termios on a PTY.
Ending a session
close is honoured in any phase:
- During identification, the chain simply stops.
- During the upload, the daemon aborts it actively: it sends a UDS reset to the product device, the gateway reboots, and the in-flight transfer fails and is discarded. The partial image left in the secondary slot is inert — its trailer is invalid, so MCUboot never swaps it.
- After the swap reset,
closefollows the normal teardown.
One benign race is accepted: if the abort reset lands just after the upload completed, the gateway swaps once and the portal self-reverts on keep-alive loss.
Three other things end a session, all of them safely:
| Event | What happens |
|---|---|
| Keep-alive timeout (default ~30 s without a PING) | the gateway reboots and reverts on its own — the central safety net |
| Gateway power-cycle | revert by construction, since the image is unconfirmed |
| Daemon crash or Brain reboot | the gateway's own keep-alive reverts it; there is no state to recover, and the next open starts fresh |
The daemon also fires a best-effort CLOSE from systemd's ExecStopPost, so a clean stop does not make the gateway wait out its timeout.
Keep-alive and loss recovery
The daemon sends PING every second and expects a STATUS reply. Three missed replies declare the session lost. STATUS rides the dedicated, higher-priority control ID pair, so a data backlog cannot starve it into a false loss.
On loss:
- The daemon fires one best-effort
CLOSE. If it lands, the gateway reverts immediately instead of waiting out its keep-alive, which shortens the recovery window. - The session enters
recovering, and the mirrors survive formirror-lingerseconds (default 180). Acandumporpicocom— or an external vendor diagnostic tool — holding the interface or the port open sees a data gap, not a dead file descriptor or a vanished interface. - The daemon polls for the product identity to reappear, then re-runs the full open: upload, handshake, and one
OPENper previously open channel with the recorded parameters, including any termios changes applied mid-session. Bridging resumes on the same mirrors. - Recovery →
active. Grace expiry, or a target that cannot be recovered → the mirrors are destroyed and the session islost.
Frames and bytes written during the gap are discarded and counted — the natural failure model of a dead bus or a silent serial device.
An FD session lost outright recovers on ISO-TP; re-open with --fd to probe FD again.
Failure modes
| Failure | Detection | Outcome |
|---|---|---|
| Upload fails, or MCUboot rejects the image | muxen-uds error, or no READY after the reset | Abort. The device still runs (or reverts to) the original firmware. Nothing to clean up. |
| Portal boots but the bootstrap stalls | READY seen, no HELLO after SET_INSTANCE | Abort. The gateway self-reverts on its keep-alive. |
| Portal firmware crashes | Software supervision timer, or Zephyr's fatal-error handler | The MCU resets → MCUboot reverts. The daemon sees ping loss and closes the session. |
Brain stops pinging (daemon crash, reboot, can0 down) | Gateway keep-alive timeout | The gateway reboots and reverts. |
| Gateway power-cycled mid-session | PING/STATUS loss on the Brain | Revert by construction. The daemon closes the session. |
| Remote CAN bus-off | STATUS report | The session stays up; the error surfaces on portal0 and in status. Automatic recovery is attempted. |
| Remote serial line errors (framing, parity, overrun) | STATUS counters | The session stays up; visible in status and stats. |
| CLOSE lost | No revert confirmation | Warn. The gateway's own keep-alive ends the session anyway. |
Unknown or unsupported product at open | Identity check | Refused before anything is uploaded. |
| Target found as an orphaned portal | Identity check | The daemon resets it, waits for the product identity, and proceeds — the user just sees a slower open. If the reset does not land: "gateway still in portal firmware, reverting — retry in ~30 s". |
Orphaned portal means a device answering as 0x3E with the target's UID and owned by no active session: the leftover of a lost session still waiting out its keep-alive.
Protocol errors are handled asymmetrically on purpose. A NAK during bring-up fails the command, with the reason spelled out in the CLI error message; an unexpected NAK mid-session is logged and counted, and the session continues.
What the crash net does and does not cover
The firmware reboots itself — and therefore reverts — on CLOSE, on keep-alive expiry, and on an unrecoverable internal error. That last one is a software watchdog: a kernel timer fed by the main loop, plus Zephyr's fatal-error handler turning a HardFault, a panic, a failed assert or a stack-sentinel trip into an immediate cold reboot.
The STM32 hardware watchdog (IWDG) is deliberately not used, because on STM32 it survives a system reset: it would keep running through MCUboot's revert swap and then through the reverted product firmware, reset-looping the gateway until someone power-cycled it. The full reasoning is in Gateway firmware.
The accepted consequence: an MCU hung hard enough that no ISR runs at all is not auto-reverted. It stops bridging, the Brain sees ping loss and closes the session, and the gateway waits in portal firmware until someone power-cycles it. Safe by the invariant, just not self-healing.
Concurrency rules
- One active session per gateway, enforced by the daemon.
- One active session per Brain. The tunnel derives dedicated ID pairs per session instance, so the protocol supports several gateways at once; the single-session limit is daemon policy.
- Bring-ups are serialised. At most one gateway sits between its reset and its
SET_INSTANCEat any time, because the parking control pair is shared by every booting portal. Loss-recovery re-inits obey the same rule.
The invariant
At every step, a power-cycle of the gateway returns the product to its original firmware. There is no state in which the portal firmware persists across a reset.
"Power-cycle" means fully off for ~10 s, not a brief dip: only a true power-on reset is guaranteed to clear all MCU state. That is cheap to arrange during a debug session, which is what makes the software watchdog's coverage gap above an acceptable trade rather than a real operational cost.
