Appearance
Tunnel protocol
The wire reference for the portal tunnel: the ID layout, the opcodes and every payload. It is the contract between the Brain daemon and the gateway firmware, and what a third client would have to implement.
The definitions are frozen in rtos/portal/.../libportal/include/portal/tunnel.h and vendored byte-identically into linux/src/tunnel.h, so both trees share one definition.
The baseline transport is ISO-TP (ISO 15765-2) over classic CAN on the MUXEN bus — the MUXEN Zephyr stack (libuds / libisotp) already provides it, and the current gateway fleet is classic bxCAN (STM32L4) anyway. Raw CAN-FD framing is negotiated per session where both ends support it: new products are STM32U5-based (FDCAN) and the new Brain is an i.MX8MP with an FD-capable controller — see Transport negotiation below. All payload layouts are transport-agnostic; one tunnel message = common header + payload, carried as one ISO-TP message or one 64-byte FD frame.
Tunnel CAN IDs
The MUXEN 29-bit ID layout (libstdmuxen/include/stdmuxen/header.h):
| Bits | Field |
|---|---|
| 26–28 | NETWORK (only 0 in use) |
| 24–25 | ACCESS: 0 = broadcast, 1 = UDS/diag, 2 = command, 3 = free → portal tunnel |
| 18–23 / 12–17 | DST function / DST instance (or 12-bit broadcast identifier) |
| 6–11 / 0–5 | SRC function / SRC instance |
The tunnel claims ACCESS = 3 with the standard addressed layout, split into two ID pairs per session — control plane and data plane — forming two independent ISO-TP connections:
| Direction | Plane | ID |
|---|---|---|
Brain → Gateway (PORTAL_ID_CTL_DOWN) | control | 0x03000000 | ctlAddr << 12 | brainAddr |
Gateway → Brain (PORTAL_ID_CTL_UP) | control | 0x03000000 | brainAddr << 12 | ctlAddr |
Brain → Gateway (PORTAL_ID_DATA_DOWN) | data | 0x03000000 | dataAddr << 12 | brainAddr |
Gateway → Brain (PORTAL_ID_DATA_UP) | data | 0x03000000 | brainAddr << 12 | dataAddr |
with ctlAddr = (0x3E << 6) \| instance, dataAddr = (0x3F << 6) \| instance and brainAddr = 0x27F (the client address muxen-uds already uses on the DIAG band). All four IDs derive from the portal instance — multi-gateway sessions need no extra allocation. 0x3F is not a device function code: it exists only inside the ACCESS = 3 band as the data-plane address of portal instance n; the portal's UDS identity stays 0x3E only.
The split exists for priority, at two layers:
- Wire arbitration — in both directions the control ID differs from the data ID only in the portal function bits, and 0x3E < 0x3F: control frames always win arbitration over data frames of the same session.
- No head-of-line blocking — separate IDs are separate ISO-TP connections, so a
PING,STATUSorCLOSEnever serializes behind a multi-frameDATAtransfer or a data backlog. Senders must also drain their control queue before their data queue locally (08-reference.md,10-firmware.md), so keep-alive and teardown latency stay bounded regardless of mirror load.
DATA and SDATA ride the data pair; every other opcode rides the control pair. A message received on the wrong pair is malformed (NAKed by the gateway; counted and dropped by the Brain).
Bootstrap — parking control pair
At boot the firmware has no instance yet, so the addressed pairs are not computable. Bootstrap runs on the parking control pair: the two control-pair formulas with instance = PARKING_INSTANCE = 63 — the standard MUXEN parking instance, confirmed against muxen-uds (its parking test is inst == 63). The portal transmits READY at 1 Hz from boot on this pair; the Brain answers SET_INSTANCE carrying the target UID and the assigned instance; the portal then re-binds on the instance-derived pairs and proceeds with HELLO. Silence on the parking address is an application-level convention, not a technical restriction — the parking pair sends and receives like any other.
The parking pair is shared by every portal in its boot phase, and two simultaneous ISO-TP transfers from one ID corrupt each other — so the daemon, which initiates every portal boot (upload + reset), keeps at most one gateway between reset and SET_INSTANCE at any time (bring-ups and loss-recovery re-inits are serialized). The UID echo in SET_INSTANCE makes an accidental overlap harmless: non-matching gateways ignore it. The keep-alive default (30 s) runs from boot; bootstrap is entirely daemon-driven, so no external parking-assignment latency can eat that budget. A protocol-version mismatch at READY aborts the open — daemon and firmware images ship in one package, so a skew means a stale install, not a negotiation case.
Being the numerically highest band in network 0, ACCESS = 3 loses arbitration to all broadcast, UDS and command traffic: a portal session can never starve normal MUXEN traffic. (NETWORK bits 26–28 are all free — a further escape hatch if ACCESS = 3 is ever claimed by something else.)
All multi-byte fields are little-endian.
Common header (4 bytes)
| Offset | Size | Field | Description |
|---|---|---|---|
| 0 | 1 | opcode | See below |
| 1 | 1 | seq | Per-direction, per-plane sequence number (wraps) — and per-channel on serial data, so a lost chunk on one channel never fakes a gap on another (10-firmware.md underrun rule); gap = message loss |
| 2 | 1 | count | DATA: number of encapsulated records. 0 in every other message (SDATA carries exactly one chunk — see len) |
| 3 | 1 | flags | Bit 0: overflow occurred since last frame; bit 1: EOB — end of burst (SDATA Brain → GW only, see SDATA payload); bits 2–3: channel index − 1 on the multi-channel families — serial (SDATA, FLOW, TERM, STATUS) and LIN (DATA, FLOW, STATUS); 0 on CAN sessions, which are single-channel; rest reserved |
Opcodes
| Value | Name | Direction | Purpose |
|---|---|---|---|
| 0x01 | HELLO | GW → Brain | Sent at 1 Hz after SET_INSTANCE, from the instance-derived control pair, until the first PING/OPEN: protocol version, fw version, capabilities: interface type, channel count, max rate, listen-only support, termination control, CAN-FD support. An unknown type is a hard reject on the Brain — never a guess — so a future family cannot be half-driven by an old daemon |
| 0x02 | OPEN | Brain → GW | Configure and start the remote interface |
| 0x03 | OPEN_ACK | GW → Brain | Success only: echoes the applied settings (= the requested ones — a confirmation trace, not a negotiation). Any failure is a NAK |
| 0x04 | DATA | both | Batched encapsulated classic CAN frames (CAN gateways), or batched LIN frames of one channel (LIN gateways — see LIN frames on the DATA record) |
| 0x05 | PING | Brain → GW | Keep-alive; gateway reboots (→ revert) if none within timeout |
| 0x06 | STATUS | GW → Brain | Reply to PING: line state, error counters, drop counters |
| 0x07 | CLOSE | Brain → GW | End session; gateway stops bridging and reboots |
| 0x08 | SDATA | both | Serial byte chunk (serial gateways) |
| 0x09 | READY | GW → Brain | Sent at 1 Hz from boot on the parking control pair until SET_INSTANCE: protocol version, UID |
| 0x0A | SET_INSTANCE | Brain → GW | Assigns the session instance (UID echo — non-matching gateways ignore); the portal re-binds on the instance-derived pairs and proceeds with HELLO |
| 0x0B | FD_PROBE | Brain → GW | Emit count (payload u8) one-shot FD test frames on the data pair; allowed between HELLO and OPEN |
| 0x0C | FD_PROBE_ACK | GW → Brain | Probe report: tx_ok, tx_err (u8 each) as seen by the gateway's controller |
| 0x0D | FLOW | Brain → GW | Payload u8: 0 = pause / 1 = resume the GW → Brain stream of the channel in flags; any Brain → GW SDATA on that channel implicitly resumes (see SDATA payload) |
| 0x0E | TERM | Brain → GW | Line-termination control: payload u8 (0 = off, 1 = on), channel in flags bits 2–3. Hot — no re-OPEN, state survives re-OPENs. NAK 0x0C if the board/channel cannot drive termination |
| 0x0F | TIMING | Brain → GW | Arm (or disarm) the per-channel wire-timing collector for a bounded window; channel in flags bits 2–3. Serial gateways only — NAK 0x0D elsewhere |
| 0x10 | TIMING_REPORT | GW → Brain | One inter-byte-gap / burst-length histogram per armed window |
| 0x7F | NAK | GW → Brain | Command or frame rejected: offending opcode + reason + detail (see NAK payload) |
The remote interface type is a property of the gateway product (CAN-to-CAN, CAN-to-SERIAL or CAN-to-LIN board), announced in HELLO; OPEN must match it or is NAKed.
DATA record — CAN and LIN gateways (20 bytes each, batched)
Mirrors struct can_frame semantics so the Linux side is a straight copy:
| Offset | Size | Field | Description |
|---|---|---|---|
| 0 | 4 | can_id | SocketCAN encoding: 29-bit ID + EFF/RTR/ERR flags in top bits |
| 4 | 4 | timestamp_us | Gateway µs clock at RX (wraps); 0 on Brain→GW records |
| 8 | 1 | dlc | 0–8 |
| 9 | 8 | data | Payload, unused bytes zero |
| 17 | 3 | — | Padding (alignment / future use) |
RTR and 11-bit IDs are carried because the SocketCAN encoding gives them for free; the target networks (NMEA 2000 / J1939, 29-bit, no remote frames) don't use them, but the CAN backend does transmit RTR frames (bxCAN and the Zephyr CAN API support them) — a mirrored tool polling with remote frames works. CAN-FD frames cannot be represented in the record (8-byte payload): the mirror is created MTU 16, so an FD write fails synchronously with EMSGSIZE (03-mirrors.md) — the FD transport of this document never implies an FD remote bus.
The sender batches whatever is queued into one message. Batch size and flush deadline are transport-tuning parameters:
- ISO-TP (baseline): up to 16 records per message (4 + 320 bytes), flush deadline ≤ 5 ms — larger batches amortize ISO-TP flow-control overhead.
- FD (negotiated): 3 records per frame (4 + 60 = 64 bytes), flush ≤ 1 ms.
LIN frames on the DATA record
LIN gateways reuse the record above unchanged — no third data-plane format. The mapping is the one the Linux sllin line discipline (the LIN sibling of slcan, in linux-can) already uses to expose a LIN bus as a SocketCAN interface, so the Brain-side mirror is a vxcan and the frames read correctly in candump with no LIN-aware tooling:
| Record field | LIN meaning |
|---|---|
can_id | LIN frame ID, 0x00–0x3F, standard frame (no EFF). The PID parity bits are not carried: the gateway checks them on RX and computes them on TX, so the Brain always sees the bare ID |
dlc | Response length, 0–8 |
data | Response bytes — checksum stripped and verified by the gateway |
| RTR flag | Header only. GW → Brain: a header went by with no response. Brain → GW: send this header and publish whatever answers |
| ERR flag | LIN error record, dlc = 1, class in data[0] (below) |
timestamp_us | Gateway µs clock at the break, as for CAN |
Error classes in data[0] of an ERR-flagged record:
| Value | Name | Meaning |
|---|---|---|
| 0x01 | CHECKSUM | Response checksum mismatch |
| 0x02 | FRAMING | UART framing / bit error on the channel |
| 0x03 | NO_RESPONSE | Header sent or seen, nothing answered before the timeout |
| 0x04 | PID_PARITY | Bad PID parity on an observed header |
| 0x05 | OVERRUN | UART overrun on the channel |
The record carries no channel of its own: the channel rides the common-header flags (bits 2–3) exactly as on serial, so one DATA message batches one channel, and seq is per-channel like SDATA. Batching and flush deadlines are the CAN ones; at 19200 baud a LIN bus produces a few hundred frames per second at most, so the 5 ms ISO-TP flush dominates and a batch rarely fills.
Diagnostic IDs 0x3C/0x3D get no special treatment — they appear as ordinary records, which is what a diagnostician wants to see.
Master vs monitor is the OPEN.mode field (below): in master mode the Brain drives the schedule, one header per Brain → GW record; in monitor mode the gateway never transmits and only reports what a foreign master puts on the wire. A Brain → GW record on a monitor channel is dropped and counted (tx_drops), not NAKed — it is a data-plane record, and the data plane never NAKs.
SDATA payload — serial gateways
| Offset | Size | Field | Description |
|---|---|---|---|
| 4 | 4 | timestamp_us | Gateway µs clock at RX of first byte (0 on Brain→GW) |
| 8 | 1 | len | Number of payload bytes (1–255) |
| 9 | ≤255 | data | Raw byte stream, order-preserving |
Chunk cap per transport: ISO-TP (baseline) 128 bytes; FD (negotiated) 55 bytes (fits one frame). The serial stream is opaque: no framing is imposed. The gateway flushes a chunk when it reaches the cap or after an idle gap (spec: 2 character times, min 1 ms) to preserve message boundaries of common protocols without adding latency. seq gaps on SDATA imply lost bytes — surfaced in stats; the tunnel does not retransmit. len is redundant with the transport length: receivers validate it against the actual message size, dropping and counting mismatches.
Serial sessions may have several channels open at once (see OPEN): their chunks are multiplexed on the single data pair, the channel index riding the common-header flags (bits 2–3). Chunking, idle-gap flush, EOB and seq all apply per channel.
Brain → GW chunking is streaming: the daemon reads the PTY and emits a chunk as soon as bytes are available (up to the cap), no accumulation delay. The chunk that empties the daemon's read buffer carries the EOB flag (common-header flags bit 1) — in practice the natural write() boundary of the application, e.g. one Modbus ADU. EOB matters on the half-duplex RS-485 backend only (DE keying and underrun resynchronization, 10-firmware.md); the RS-232 backend ignores it, and GW → Brain chunks never set it (the idle-gap rule above already frames that direction). The tunnel preserves byte order and contiguity, not silences: inter-frame gaps (e.g. Modbus 3.5 chars) remain the application's responsibility, exactly as on a local port with a deep buffer.
Flow pause (FLOW, serial sessions, per channel): when nothing on the Brain reads a channel's PTY, the daemon pauses that channel's GW → Brain stream instead of shipping bytes it would discard — every tunnel message loads all nodes of the MUXEN bus. Paused, the gateway keeps listening and counting (STATUS still reflects remote activity) but emits no SDATA for that channel; nothing is buffered. Resume is explicit (FLOW(1)) or implicit on any Brain → GW SDATA of the channel — a request/response master reopens the mirror with its own request, before the reply exists. Default at OPEN: active — note this means an OPEN clears an existing pause, so a re-OPEN on a paused channel has to be re-paused by the Brain. "Keeps listening and counting while paused" is also what makes per-port statistics work on ports nobody mirrors locally (04-monitoring.md, Monitor channels).
OPEN payload
| Offset | Size | Field | Description |
|---|---|---|---|
| 4 | 1 | type | 0 = CAN, 1 = serial — must match the gateway's HELLO |
| 5 | 1 | iface | Remote channel, 1-based — same numbering as the CLI and the product firmware's channels (CAN gateways: 1; serial gateways: 1–4; LIN gateways: 1–3). 0 or above the HELLO channel count → NAK |
| 6 | 2 | ping_timeout | Keep-alive timeout in units of 100 ms; 0 = default 30 s. Valid range 50–3000 (5 s – 5 min): below, PING-loss bursts cause spurious reverts; above, a dead Brain leaves the product out of service too long ("transient by construction"). Out-of-range → NAK, never silently clamped |
| 8 | 4 | rate | CAN bitrate, serial baudrate, or LIN baudrate, in bit/s resp. baud. LIN valid range 1000–20000 (20 k is the LIN 2.x ceiling; default 19200) — outside → NAK 0x06 |
| 12 | 1 | mode | CAN: 0 = normal, 1 = listen-only. LIN: 0 = master, 1 = monitor (same two values, see LIN frames on the DATA record). Serial: reserved (0) |
| 13 | 1 | serial_cfg | Serial: bits 0–1 data bits (0 = 8, 1 = 7), bits 2–3 parity (0 = none, 1 = even, 2 = odd), bit 4 stop bits (0 = 1, 1 = 2), bits 5–7 reserved (0) — DE handling is a board property (10-firmware.md), not a session parameter. LIN: checksum selector — 0 = classic, 1 = enhanced, 2 = auto (classify per frame by the LIN 2.x rule, IDs 0x3C–0x3F always classic); other values → NAK 0x07. LIN framing itself is fixed 8N1 by the protocol and is not configurable |
| 14 | 1 | transport | 0 = ISO-TP (default), 1 = raw CAN-FD — NAKed unless HELLO advertised FD support |
| 15 | 1 | stmin | ISO-TP STmin in ms for both directions (0 = back-to-back, default) — session-level, from the daemon setting (08-reference.md). Valid range 0–10; out-of-range → NAK 0x08 (see ISO-TP parameters for why) |
OPEN is per-channel and additive (serial and LIN): each OPEN opens its iface, or — if that channel is already open — reconfigures it, restarting only that channel briefly. Session-level fields (ping_timeout, transport, stmin) are set by the first OPEN; later OPENs must repeat them identically (NAK 0x0B otherwise) — with one exception: transport may be lowered 1 → 0 (the FD runtime-fallback path, Transport negotiation); raising it mid-session stays forbidden. There is no per-channel close: an unread channel costs nothing on the bus (FLOW), and CLOSE always ends the whole session.
READY / SET_INSTANCE payloads
| Offset | Size | Field | Description |
|---|---|---|---|
| 4 | 1 | version | Protocol version (SET_INSTANCE echoes it) |
| 5 | 12 | uid | Device UID: first 8 bytes = the muxen UID (uint64 derived from the STM32 unique ID, same pattern as the product apps and the UDS UID answer — 64-bit collision risk accepted); last 4 bytes zero, ignored on compare |
| 17 | 1 | instance | SET_INSTANCE only: assigned instance (≠ PARKING_INSTANCE) |
NAK payload
| Offset | Size | Field | Description |
|---|---|---|---|
| 4 | 1 | opcode | The offending opcode |
| 5 | 1 | reason | See below |
| 6 | 2 | detail | Rejected value or violated bound; 0 when unused |
| Reason | Meaning |
|---|---|
| 0x01 | Malformed frame (bad size, inconsistent len/count) |
| 0x02 | Unsupported opcode |
| 0x03 | Wrong plane (data opcode on the control pair or vice versa) |
| 0x04 | type does not match the gateway |
| 0x05 | iface invalid (0 or > channel count) |
| 0x06 | rate unsupported (> HELLO max rate, or not exactly applicable) |
| 0x07 | serial_cfg / mode unsupported (e.g. listen-only not advertised) |
| 0x08 | Timing parameter out of range (ping_timeout 50–3000, stmin 0–10) |
| 0x09 | Transport unavailable (FD not advertised) |
| 0x0A | UID mismatch / bad instance (SET_INSTANCE) |
| 0x0B | Wrong state (e.g. OPEN before SET_INSTANCE) |
| 0x0C | Termination control unavailable on this board/channel |
| 0x0D | Timing collector unavailable on this board/channel (TIMING on a CAN or LIN gateway) |
One error path, no silent clamping: a command is applied exactly as requested or NAKed with the reason — OPEN_ACK never carries an error nor a corrected value (this generalizes the ping_timeout rule). Validation is atomic: a NAKed OPEN/re-OPEN leaves the current configuration and the running bridge untouched. The Brain never emits NAK: a malformed gateway frame is dropped, counted and logged (08-reference.md).
TIMING / TIMING_REPORT payloads — wire timing
The tunnel preserves byte order and contiguity, not silences (SDATA payload): the daemon injects into the PTY as bytes arrive and original inter-byte spacing is never re-simulated (03-mirrors.md, Timing and mirroring semantics). Inter-byte gaps therefore exist in exactly one place — the gateway's UART RX interrupt — and no Brain-side measurement can recover them at any price.
TIMING ships the measurement instead of the stream. The gateway accumulates an inter-byte-gap histogram for one bounded window and returns it in a single 54-byte control-plane message. That answers the first question about any unknown serial port — what shape is the traffic? — for the cost of one round trip and no data bytes at all.
Three properties are the reason it is a protocol feature rather than a Brain-side one:
- No stream is read to get it. No PTY, no reader, no
FLOWresume, nothing on the data plane. It works on aFLOW-paused monitor channel — which is the normal state of every port the operator did not explicitly open (04-monitoring.md), and exactly where "is anything talking on port 3, and what kind of thing?" needs answering. - Gaps are bucketed in character times, computed on the gateway from the applied baud and framing. The 3.5-character Modbus inter-frame boundary is then an exact bucket edge at every rate; the Brain cannot do that correctly while it is still sweeping the line settings.
- It is the honest basis for a quiet-bus interlock. "The gateway's ISR recorded no byte" is a much stronger claim than "the Brain read nothing", and it is the claim anything about to transmit on a multidrop RS-485 segment needs first.
Serial gateways only. CAN framing is inherent to the bus and LIN is not specified yet (the payload shape would suit it — header/response turnaround timing is a natural fit); both answer NAK 0x0D.
TIMING payload (Brain → GW, 4 bytes)
| Offset | Size | Field | Description |
|---|---|---|---|
| 4 | 1 | mode | 0 = disarm (report immediately with whatever was collected), 1 = arm one-shot |
| 5 | 1 | — | Reserved (0) |
| 6 | 2 | window_ms | Measurement window, 100–60000; 0 = default 5000. Out of range → NAK 0x08 |
One collector per channel. Arming a channel that is already collecting restarts it with the new window — re-arming is how a caller extends one, so it is not an error. The channel must be open: a closed channel has its UART off and physically cannot count, the same constraint STATUS lives under (04-monitoring.md) — NAK 0x05 otherwise. A re-OPEN of the channel, CLOSE, or session loss cancels an in-flight window silently: the settings the bucket edges were derived from are gone, and a report built from half a window would be worse than none.
TIMING_REPORT payload (GW → Brain, 54 bytes)
| Offset | Size | Field | Description |
|---|---|---|---|
| 4 | 4 | window_us | Actual elapsed measurement window (shorter than requested if disarmed early) |
| 8 | 4 | bytes | Bytes received in the window (tracks the STATUS delta) |
| 12 | 2 | char_time_us | One character at the applied baud + framing, µs — lets the Brain re-derive absolute gaps from the buckets |
| 14 | 2 | bursts | Bursts closed in the window (a burst ends on a gap ≥ 3.5 character times) |
| 16 | 16 | gap_hist[8] | u16 each — inter-byte gaps in character times |
| 32 | 16 | burst_hist[8] | u16 each — burst lengths in bytes |
| 48 | 2 | burst_gap_min_ms | Shortest inter-burst gap |
| 50 | 2 | burst_gap_mean_ms | Mean inter-burst gap — the poll period, on a polled bus |
| 52 | 2 | burst_gap_max_ms | Longest inter-burst gap |
| 54 | 1 | burst_len_min | Shortest burst, bytes (saturating at 255) |
| 55 | 1 | burst_len_mean | Mean burst, bytes (saturating) |
| 56 | 2 | burst_len_max | Longest burst, bytes |
Gap bucket edges, in character times:
| Bucket | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|---|
| Range | <1.5 | 1.5–3.5 | 3.5–8 | 8–20 | 20–100 | 100–1k | 1k–10k | ≥10k |
Burst-length bucket edges, in bytes:
| Bucket | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|---|
| Range | 1–3 | 4–8 | 9–16 | 17–40 | 41–88 | 89–200 | 201–512 | >512 |
The edges are not arbitrary, and they are the whole content of the message. 1.5 and 3.5 character times are the Modbus RTU intra-frame and inter-frame limits, so a histogram dense in bucket 0, empty in bucket 1 and populated beyond says the traffic arrives as bursts separated by a proper silent interval. The burst-length edges are anchored on an 8-byte Modbus request, a 5+2N response, and a 41–88 byte NMEA 0183 sentence.
That gap shape alone does not identify Modbus. A device that emits a block then falls silent makes two modes too — a Victron MPPT on a plain serial line, unpolled, with no master anywhere on the segment, produces exactly it. Reading the two histograms together is what separates them: short bursts plus a populated 3.5–20 turnaround band mean an exchange; long bursts with an empty middle mean one device talking to nobody. That is why both histograms travel in the same message. See
04-monitoring.md.
flags bit 0 (the existing overflow bit) means a histogram counter saturated: the distribution's shape is still valid, the absolute counts are floors. Counters saturate rather than wrap — a clipped histogram keeps the shape being read, a wrapped one inverts it.
Transport negotiation
- The control plane (
HELLO,OPEN,OPEN_ACK,PING,STATUS,CLOSE,NAK,READY,SET_INSTANCE,FD_PROBE,FD_PROBE_ACK,FLOW,TERM) always runs on ISO-TP over classic frames — keep-alive, teardown and the FD probe itself never depend on FD working. HELLOadvertises CAN-FD capability. The Brain attempts FD only if the gateway advertised it, the Brain's own controller is FD-capable, andallow-fdpermits the attempt (08-reference.md) — FD tolerance of the other nodes cannot be assumed, and even probing disturbs an intolerant bus for a few ms, so the attempt stays an explicit deployment decision.- FD probe — verify before switching. A successful CAN frame has, by protocol, been accepted by every node on the bus; a classic (error-active) node destroys any FDF frame with an error flag, visibly to the transmitter. So between
HELLOandOPENthe Brain sendsFD_PROBE(count); the gateway emitscountFD test frames on the data pair, each with retransmission bounded to ~3 ms by a completion-callback timeout (Zephyr has no per-frame one-shot, and the controller-wide one-shot mode would disturb the shared control plane) — a few ms of error frames at worst — and reportstx_ok/tx_errinFD_PROBE_ACK. The Brain requeststransport = 1only if it also received at least one test frame andtx_err≈ 0 — end-to-end double confirmation. Otherwise it opens withtransport = 0and logs the reason. The Brain's owncan0must be up in CAN-FD mode (MTU =CANFD_MTU) to probe or run FD; a classiccan0is detected via its MTU and treated as one more clean probe-failure gate, not an error. - Runtime fallback — cause-agnostic. The probe covers bring-up, not a classic node hot-plugged mid-session nor a bus whose 2 Mbit/s data phase turns out marginal. Trigger: the FD data plane under-delivers while control stays healthy, cross-checked over 3 consecutive
STATUSvia the per-direction tunnel message counters ("sent N, peer counted ~0" — a silent remote bus is not a trigger). Reaction: re-OPENwithtransport = 0(a secondOPENalready reconfigures): session, mirror and apps continue over ISO-TP at reduced throughput; data in flight during the detection window is lost (counted). No FD retry within the session. An FD session lost outright recovers on ISO-TP (loss recovery re-opens withtransport = 0); re-open with--fdto probe FD again. If errors hit classic frames too (genuinely sick bus), falling back cannot help — the same wires carry both; the keep-alive then decides, and a lost session reverts the gateway as always. - After
OPEN_ACKaccepts FD,DATA/SDATAmessages flow as raw FD frames (FDF flag set) on the data pair, replacing its ISO-TP encoding; the control pair is untouched. Planes demultiplex by CAN ID; on the data pair the FDF flag still separates ISO-TP data messages in flight from before the switch. Raw FD frames are padded to the next valid FD DLC; receivers recover the logical message length from the header (countfor DATA,lenfor SDATA). - Current-fleet gateways (STM32L4 bxCAN) never advertise FD; STM32U5-based products will, and the i.MX8MP Brain can request it.
- Only the Brain and the gateway need to be FD-capable; every other node needs to be FD-tolerant — and tolerance is a controller configuration, not just silicon: an FDCAN (STM32U5) configured classic-only destroys FD frames exactly like a bxCAN. Future U5-based product firmwares should enable the controller's FD mode even when speaking classic only, or they will keep the segment FD-intolerant. The L4 fleet can not be made tolerant in firmware — bxCAN is a pure CAN 2.0 engine with no FD-tolerance mode, and silent mode would mute the product entirely; the only hardware route (an FD-passive filtering transceiver) means a board re-spin. FD therefore remains a 100 %-U5-segment feature.
STATUS payload
CAN: bus state (error-active / error-passive / bus-off), TEC/REC, RX/TX frame counters, RX overflow drops, tunnel-side drops, uptime. Serial: RX/TX byte counters, framing/parity/overrun error counters, TX underruns (RS-485, 10-firmware.md), drops, uptime. LIN: applied mode and checksum rule, RX frame and TX header counters, no-response / checksum / PID-parity / framing / overrun counters, drops, uptime, applied baud. Every variant includes per-direction tunnel message counters (sent/received on the data pair) — the runtime fallback of Transport negotiation cross-checks them against local activity. On multi-channel sessions — serial and LIN — the gateway answers each PING with one STATUS per channel the board has, open or not (channel index in flags) — a closed channel reports open = 0 with zeroed counters, so the Brain always sees the full port inventory (04-monitoring.md); receiving any of them feeds the Brain's liveness check. Where the board advertises termination control, each serial STATUS also carries the termination status: a driven-by-portal flag plus the effective on/off state — the firmware reads the pin back when it is not driving it, so the state is always known. Exact layout is frozen in libportal/include/portal/tunnel.h (the shared definition, vendored byte-identical into linux/src/tunnel.h); kept deliberately compact (≲ 60 bytes per message) so the 1 Hz poll costs a handful of frames.
Error and flow behaviour
- Bus errors on a remote CAN side are reported both via STATUS counters and as SocketCAN error frames (
CAN_ERR_FLAGset incan_id) in DATA records, socandump portal0,#FFFFFFFFshows them natively. Serial line errors (framing, parity, overrun) are reported via STATUS counters only. LIN errors go both ways like CAN: an ERR-flagged DATA record per event (class indata[0], see LIN frames on the DATA record) plus the STATUS counters. - Overflow: if either side's queue fills, oldest frames are dropped, the drop counter increments, and the next header sets the overflow flag. The tunnel never blocks the MUXEN bus.
- Loss detection:
seqgaps are counted and exposed inmuxen-portal status; the tunnel itself is not retransmitting (same best-effort semantics as a real CAN interface). - Unexpected FD frames: an FDF frame received on the data pair when FD was not negotiated (misconfiguration, stray sender) is dropped and counted; it never reaches the mirror.
ISO-TP parameters
Both ends run identical ISO-TP settings on the tunnel pairs: BS = 0 (single flow control per transfer — the muxen-uds convention), TX padding on (pad to 8, fixed DLC), and STmin from the stmin session parameter — daemon-configurable, default 0 = back-to-back. muxen-uds itself runs STmin = 1 on the DIAG band; these are per-connection parameters, so the tunnel choice does not interfere. The bandwidth figures below assume STmin = 0: at 1 ms, useful throughput caps at ~56 kbit/s and mirroring a loaded remote CAN degrades to lossy listening — raising stmin is a deliberate bus-gentleness trade-off. stmin is bounded to 10 ms because the gateway's sender is a single sequencer: a data transfer's consecutive frames are paced by STmin, and beyond ~10 ms one full 48-CF transfer would hold STATUS long enough (≥ 0.5 s per 10 ms step) to false-trip the Brain's 3-missed-STATUS loss detector; at 10 ms the worst delay stays ≈ 0.5 s — 6× margin.
Bandwidth sanity check (ISO-TP baseline)
ISO-TP on the shared 500 k MUXEN bus yields at best ~400 kbit/s of tunnel payload on an otherwise idle bus (7 useful bytes per consecutive frame), less in practice since normal MUXEN traffic continues.
- Remote CAN at 125 k: fully mirrorable (worst case ≈ 22 kB/s of records).
- Real 250 k profile — an NMEA 2000 segment with ~10 nav devices (GNSS, pilot, router, displays): a few hundred frames/s, ~15–25 % bus load ≈ 6–10 kB/s of records — comfortably mirrorable at STmin = 0 (4–6× margin), while busy moments already exceed the ~56 kbit/s STmin = 1 ceiling — which is why 0 is the default.
- Remote CAN at 250 k saturated: marginal (≈ 44 kB/s ≈ 360 kbit/s) — expect drops on bursts; drop counters and the overflow flag make this visible.
- Remote CAN at 500 k saturated: not mirrorable losslessly over ISO-TP; acceptable for diagnostics (listen with drops), the negotiated FD transport removes the limit (~3× headroom at 2 Mbit/s data phase).
- Serial: actual profile is 19200 baud with ≤ 300-byte bursts at 1 Hz (~300 B/s average) — negligible, ≈ 3 tunnel messages per burst; even four channels saturated at 19200 stay < 10 kB/s.
In practice the fleet's remote buses never approach saturation — the "saturated" rows above are engineering bounds, not operating points; the real profiles sit comfortably inside the ISO-TP budget, and the drop counters + overflow flag exist for the unlikely day a bound is hit.
Note: MUXEN devices accept all extended frames and filter in software (match-all hardware filter, dispatch in the app), so sustained tunnel traffic adds interrupt/dispatch load on every node of the MUXEN bus — the same class of load as a UDS firmware upload, acceptable for a diagnostic session but worth keeping in mind for long mirroring sessions on busy remote buses.
