Skip to content

Gateway firmware ​

Minimal Zephyr applications for the STM32 gateways, reusing each product's existing DTS, MCUboot bootloader and signing key, delivered as srec for muxen-uds upload.

The electrical realities differ enough that one bridge implementation cannot cover them: CAN is frame-oriented with a controller doing the framing; TTL VE.Direct is a plain inverted-level byte stream; RS-485 is half-duplex and needs the driver-enable keyed around every transmission; LIN is a master-scheduled frame protocol carried on a UART with hardware break generation. So: one shared library (libportal), four bridge backends, six boards, seven images.

BoardSoCAppBackendRemote interface
muxen_cancanSTM32L496, 2× bxCANapp_portal_canCANcan2 (canhost alias)
muxen_cancan_v2STM32L496, 2× bxCANapp_portal_canCANcan2 (canhost alias)
muxen_can-serial_v2STM32U575, 1× FDCAN, 4 UARTapp_portal_vedirect / app_portal_modbusRS-232 + RS-485usart2/3/1 + uart4; TTL VE.Direct (tx/rx-invert, 19200) or RS-485 Modbus (DE/RE/termination GPIOs), by wiring — both hardwares on one board
muxen_can-linSTM32L431, 1× bxCAN, 3 UARTapp_portal_linLINlin01/lin02/lin03 = usart2/usart3/usart1, each with its own transceiver enable en-linN; 19200 nominal
muxen_can-mdv_v2STM32U575, 1× FDCAN, 1 LIN UARTapp_portal_linLINlin01 = usart2 with enable en-lin1; the board DTS names them lin / en-lin, renamed in the board overlay
muxen_can-enoceanSTM32U575, 1× FDCAN, 1 UARTapp_portal_vedirectRS-232vedirect-ch1 = usart2 (board alias uartenocean), 57600 8N1, no tx/rx inversion — a point-to-point link to an on-board EnOcean TCM 515Z, not a customer bus

The board reference, image selection and product tables are in Supported gateways.

app_portal_can builds for both cancan boards and app_portal_lin for both LIN boards. muxen_can-serial_v2 is a combined serial board: it carries both the TTL VE.Direct and the RS-485 Modbus hardware and selects between them by wiring, so it builds under both serial apps — two serial images from one board. muxen_can-enocean is serial but not combined: one UART, one backend, one image. Each app shares its backend across boards, differing only by boards/<board>.conf — the same one-app-many-boards pattern the product repos already use.

Retired from the matrix. The legacy L431 muxen_can-modbus (990010053) and muxen_can-ve-direct (990010012) are no longer built and no longer selectable. Both ship the v5 bootloader only, and its revert-on-reset behaviour is unverified. A portal image is only safe because MCUboot runs it padded-and-unconfirmed and reverts to the product firmware on any reset (non-negotiable property 1); on a bootloader where that has not been demonstrated the correct number of images to ship is zero. Re-adding either board means proving the revert on real hardware first, then restoring the job in ci/portal-firmware.yml and the entry in .fw_optional_needs and the row in the daemon's board table.

Channel numbering ​

Channel numbering is 1-based and follows the product firmware: portal channel n is the same physical UART and connector as the product's channel n, taken from the same DTS aliases the product apps use. The portal never redefines the order.

On muxen_can-serial_v2 that is usart2, usart3, usart1, uart4 for channels 1–4 — note the order is not usart1..4. The portal image runs with no console at all, so no UART is ever reserved away from a channel.

The same rule, and the same trap, applies on muxen_can-lin: its aliases are lin01 = usart2, lin02 = usart3, lin03 = usart1, each paired with en-lin1/2/3 by index, not by any devicetree link. Assuming lin0n = usartn swaps channels 1 and 3 while the enables keep following the index — the board would transmit on one UART with another channel's transceiver keyed, and the symptom is a dead bus, not a build error. backend_lin.c pins the map with BUILD_ASSERTs so a reshuffle of the board DTS breaks the build instead of the bus.

Firmware identity ​

Every MUXEN image carries a muxen,identification DTS node with four fields, read back over muxen-uds readconfig: HardwareId (990010xxx, the board), ProductId (010010xxx, the product variant), SoftwareId (990050xxx, the firmware build identity) and SoftwareVersion.

The portal is a mode, not a product. It is a transient firmware the gateway runs during a session, never a shippable device, so a portal image sets ProductId 000000000 while keeping the board's real HardwareId. It builds board-level with no ProductId overlay, and open selects it by the underlying gateway product's identity — never by any portal identity field.

Where a board DTS declares muxen,identification only in its product overlays — muxen_can-lin, muxen_can-mdv_v2 and muxen_can-enocean do — the portal app's overlay must create the node with the board's real HardwareId rather than merely override fields.

The four SoftwareIds are listed in Supported gateways. A SoftwareId names a backend, not a board.

libportal ​

Everything that is not hardware-shaped lives in libportal, an in-tree Zephyr module beside the apps, pulled into each app through EXTRA_ZEPHYR_MODULES before find_package(Zephyr). One repo and one tag cover the library and the apps.

ConcernWhere
Tunnel wire codec (header, DATA/SDATA records, batching, seq, flags)libportal
Control-plane state machine (every opcode: bootstrap, session, FD probe, FLOW, NAK)libportal
ISO-TP transport endpoints, FD negotiationlibportal
Keep-alive timer, software supervision, reboot-to-revertlibportal
UDS identity (function 0x3E, SET_INSTANCE, reset + routines)libportal
RX rings, TX queues, drop and error counterslibportal
Wire-timing collector (TIMING)libportal
Remote interface bring-up, TX, RX timestamping, line errorsbackend (per app)

The seam is a small vtable in backend.h: libportal owns the session and calls down, the backend owns the wire and calls back up.

c
struct portal_backend {
    uint8_t  type;          /* PORTAL_IFACE_CAN | _SERIAL | _LIN      */
    uint8_t  nb_channels;   /* derived from the aliases the board declares */
    uint32_t max_rate;      /* bit/s resp. baud cap advertised in HELLO */
    bool     listen_only;   /* listen-only (CAN) / monitor (LIN)      */
    bool     termination_ctrl; /* line termination drivable (RS-485)  */
    bool     supports_fd;   /* false on every L4 board                */

    int (*open)(const struct portal_open *cfg);          /* per channel   */
    int (*close)(void);                                  /* whole session */
    int (*tx)(uint8_t ch, const void *d, size_t len, uint8_t flags);
    int (*set_term)(uint8_t ch, bool on);                /* NULL if unsupported */
    int (*stats)(uint8_t ch, struct portal_status *out); /* PING → STATUS */
};

RX runs the other way: the backend timestamps on reception and calls portal_rx_push(), and libportal's sender thread does the batching, the flush deadline and the transport encoding. Control and data ride distinct TX queues onto their respective ID pairs, control drained first, so a STATUS or CLOSE never waits behind queued DATA. Open serial channels share the data queue round-robin.

A backend therefore contains no protocol knowledge, and libportal contains no #ifdef CONFIG_BOARD_*. HELLO capabilities are filled straight from the backend fields, so an OPEN that does not match the image's hardware is NAKed by libportal without the backend being involved.

nb_channels is derived from which channel aliases the board declares, so a board gets exactly the channels it wires.

The backends ​

  • CAN (app_portal_can): the second bxCAN controller (classic), bridged as DATA records via the Zephyr CAN API.
  • RS-232 (app_portal_vedirect): up to four UART channels at once, each opened by its own OPEN, TTL levels with tx-invert/rx-invert from the board DTS, bridged as SDATA chunks. Full duplex, no line turnaround. It also serves the point-to-point EnOcean link, which differs only in baud and in not inverting the levels — both board-DTS facts, which is why that board needs no app of its own.
  • RS-485 (app_portal_modbus): the same four channels plus half-duplex line control — DE/RE/termination GPIOs, DE keyed by queue state and EOB. This is the only backend with a TX/RX turnaround to get wrong, which is why it is its own app rather than a flag on the RS-232 one.
  • LIN (app_portal_lin): each channel a UART in the STM32 hardware LIN mode (USART_CR2 LINEN: 13-bit break generation and break detection) plus a transceiver enable held asserted for as long as the channel is open — it is an enable, not an RS-485 turnaround, so it is never keyed per frame. Master or monitor per channel. The backend computes and checks PID parity and the classic/enhanced checksum, so the Brain never sees either.

Non-negotiable properties ​

1. Never confirm the image ​

This is where the safety property is won or lost. The UDS update path writes the uploaded artifact verbatim into the secondary slot, trailer included, so the artifact's trailer alone decides test versus permanent.

  • Sign the portal image padded but unconfirmed: imgtool --padwithout --confirm. Never enable CONFIG_MCUBOOT_GENERATE_CONFIRMED_IMAGE — the products' own OTA artifacts are pad+confirm, which is why their swaps are permanent.
  • No call to boot_write_img_confirmed() at runtime. The cancan product apps call it at boot; the portal app must not.

MCUboot then runs the image in test mode and any reset reverts to the original firmware. A portal session can always be ended by power-cycling the product.

2. Self-terminating ​

The firmware reboots the MCU — and therefore reverts — when a CLOSE arrives, when no PING arrives within the negotiated timeout (default 30 s), or on an unrecoverable internal error.

Supervision is a software watchdog: a kernel timer fed by the main loop. The STM32 IWDG must NOT be enabled — not by the app, not through CONFIG_WDT_*, not in a board .conf. The reason is specific and counter-intuitive enough to be worth recording, because the design reads like it wants a hardware watchdog:

On STM32 the IWDG survives a system reset. Only a power-on reset clears it; sys_reboot(), NVIC_SystemReset(), the NRST pin and an IWDG reset itself all leave it armed and counting. The WDG_SW option byte does not change this — it decides only who starts the watchdog, never whether it survives one. (This is the same trap MCUboot's CONFIG_BOOT_WATCHDOG_FEED exists to work around on nRF52.)

An IWDG started by the portal would still be running after the teardown reset, and would keep running:

  • through MCUboot's revert swap, several seconds on the scratch-swap boards, which no product bootloader feeds — resetting the swap mid-flight with an inherited, non-deterministic count;
  • and then indefinitely, because the reverted product app does not feed it either. The gateway would reset-loop every ≤ 32.7 s until someone power-cycled it.

That second consequence is the disqualifying one: it turns every session teardown into a physical trip to the hardware, which is precisely what this project exists to avoid. CONFIG_BOOT_WATCHDOG_FEED cannot rescue it — the bootloaders are in production and out of scope, and changing them means re-flashing the fleet by hand.

The same property was checked in the other direction, at session start: an IWDG started by the product firmware would equally survive the swap reset and run, unfed, inside the portal app, ending every session on that product within the IWDG period. No product app starts the IWDG, so the portal always boots watchdog-free.

Accepted cost. The software watchdog catches a wedged main loop, since the timer ISR still runs, and any Zephyr fatal error — HardFault, k_panic(), failed assert, stack-sentinel trip — is turned into an immediate sys_reboot(COLD) by the firmware's own k_sys_fatal_error_handler. (CONFIG_MUXEN_PORTAL_FATAL_HALT=y restores Zephyr's default halt for JTAG bench debugging only.) Only an MCU hung hard enough to run no ISRs at all is not auto-reverted; it stops bridging, the Brain sees ping loss, and the gateway waits in portal firmware for a power-cycle. Safe by the invariant, just not self-healing — and cheap next to a guaranteed reset-loop.

3. Passive until opened ​

Before OPEN the remote interface stays down. On the MUXEN bus the firmware speaks only on the tunnel IDs — READY, HELLO, control replies — and as its own UDS device. It must not disturb the bus beyond that.

4. First-class UDS device ​

The portal firmware registers on the MUXEN bus with function code 0x3E, the instance assigned by the Brain via SET_INSTANCE at tunnel bootstrap — the portal registers directly with that instance, with no dependency on the UDS parking assignment. It broadcasts life frames and answers UID requests with the muxen UID, derived from the STM32 unique ID exactly as the product apps derive it. That shared derivation is the invariant linking portal and product identities for revert verification.

The UDS surface is reset (0x11) and routines (0x31) only — never readconfig/writeconfig. Config storage is shared with the product firmware and changes apply on restart, which reverts the portal image, so the portal must not touch the product's stored settings at all. Session and bridge parameters stay on the tunnel control plane.

Two always-compiled self-test routines ride on 0x31 for bench use: 0xF001 wedges the main loop, proving the supervision revert on stuck firmware, and 0xF002 calls k_panic(), proving the fatal-to-reboot path.

5. Read the MUXEN transceiver enable where CAN is started ​

Every gateway DTS declares inh-can-gpios on PA1 as GPIO_ACTIVE_HIGH, and getting it wrong produces a firmware that boots and never reaches the bus — no READY, no tunnel, a failure that reads as a hang and takes its own evidence away when the unconfirmed image reverts.

It is an inhibit on every board, and every portal app drives it INACTIVE (low). The trap is that reading a product's GPIO initialisation in isolation suggests the opposite on some boards: InitGpio() only establishes the safe boot state — transceiver inhibited — and the product releases the pin later, inside the routine that actually starts CAN:

c
InitCommCan(): INH_CAN(1); ... can_start(); ... INH_CAN(0);

The rule is therefore: find where the product starts CAN, and read the pin there — not where it initialises its GPIOs.

Boot sequence ​

  1. MCUboot swaps to the portal image (test mode) and the portal boots.
  2. Arm the software supervision timer; start the keep-alive timer.
  3. Bring up the MUXEN-bus interface at the product's standard bitrate (classic 500 k) and the ISO-TP endpoint on the parking control pair.
  4. Transmit READY (protocol version, UID) at 1 Hz until SET_INSTANCE assigns the instance. A non-matching SET_INSTANCE is ignored.
  5. Register the UDS identity, re-bind the tunnel endpoints on the instance-derived pairs, and transmit HELLO at 1 Hz until the first PING/OPEN.

Runtime ​

Lines marked (backend) are what each app implements behind the vtable.

  • OPEN → backend->open() (backend): CAN — bitrate and mode via the Zephyr CAN API; serial — baudrate and framing via uart_configure(). libportal replies OPEN_ACK with the applied settings. OPEN is per-channel and additive on serial and LIN; an OPEN on an already open channel reconfigures it, briefly restarting that channel only.

  • Bridging GW→Brain, CAN (backend): the RX callback timestamps the frame from the µs cycle counter and calls portal_rx_push(). libportal's sender thread batches records into tunnel messages. On ring overflow: drop oldest, count, set the overflow flag.

  • Bridging GW→Brain, serial (backend): UART RX is interrupt-driven, into a byte ring; libportal emits SDATA chunks at the transport's chunk cap or after the idle gap. The product board DTS carry no DMA channels for these UARTs, and the product apps run interrupt-driven too.

  • Bridging Brain→GW: libportal decapsulates records and chunks and calls backend->tx(), which queues them on the remote interface. TX queue overflow: drop, count.

  • Flow pause (serial): libportal gates the GW→Brain stream on FLOW; the backend keeps receiving and counting, so STATUS still reflects remote activity while paused.

  • RS-485 line turnaround (backend, app_portal_modbus only): driver-enable, receiver-enable, termination and duplex are GPIOs declared in the DTS. Unlike the product app, which holds DE statically enabled, the portal keys DE by queue state: assert on the first byte, emit contiguously, deassert once the final byte of an EOB-flagged chunk has left the wire. RE is disabled during TX — a Modbus master must not see its own request.

    Gap-free TX within a burst is guaranteed by throughput, not by buffering ahead: the tunnel outruns the wire, so the backend caps the max rate it advertises in HELLO with comfortable margin (the real profile is 19200 and the tunnel is roughly ten times faster; the invariant erodes near 115200). If an underrun still occurs before EOB, or a seq gap shows a chunk was lost mid-burst, the backend deasserts DE, counts tx_underrun and discards everything up to and including the next EOB — resynchronising on a frame boundary, never resuming a cut frame. A resumed fragment would start with a data byte read as a slave address, and a lucky CRC could forge a command; and if the first fragment happened to be CRC-valid, resuming would collide with the slave's reply. The cut frame dies on CRC and the master retries, which is standard degraded behaviour.

  • Termination (backend, boards advertising termination_ctrl): the termination GPIOs stay undriven (hardware reset state) until the first explicit TERM — no silent change to a terminated installation. Control is opt-in per session, hot, and the driven state survives re-OPENs. When not driving, the backend reads the pin back, so STATUS always reports the effective state plus a driven-by-portal flag.

  • PING → STATUS via backend->stats(); this also feeds libportal's keep-alive timer. Serial and LIN backends answer with one STATUS per channel the board has, open or not: a closed channel reports open = 0 with zeroed counters, giving the Brain the full port inventory. Its UART is off, so those counters cannot move until the channel is opened — the Brain's monitor-open is what makes an unused port countable.

  • LIN turnaround and errors (backend): in master mode a Brain→GW record emits break + sync + PID, then either the data and checksum, or nothing while the response window is open (RTR). A window that closes empty is a NO_RESPONSE error record, not a silent drop — an absent slave must be visible, and it is the normal outcome of scanning an ID range. In monitor mode the backend never transmits: a Brain→GW record is dropped and counted, the same contract as a full TX queue. The response timeout is derived from the applied baud (LIN 2.x: 1.4 × nominal frame time), never a flat constant.

    Frame end is an inter-byte idle gap, and the last byte is the checksum. Without a schedule table the response length is not knowable in advance, so this is the only available rule, and it is the one sllin applies. RX is muted while transmitting a header, because the bus is single-wire and the transceiver echoes our own bytes back into our receiver. Master transactions run on their own worker thread: a LIN transaction takes milliseconds on the wire, and blocking libportal's delivery path would stall the tunnel.

  • Remote CAN bus-off: report via STATUS and error records, attempt automatic recovery. Line errors of any kind: count, keep running. Never reboot for line errors alone.

The wire-timing collector ​

libportal/src/timing.c implements the TIMING / TIMING_REPORT pair. It sits entirely in libportal and needs nothing from the backends: it is fed from portal_rx_push_serial(), which every serial backend already calls, so both serial apps gain the feature without a line of change and the vtable is untouched. The CAN and LIN apps answer NAK 0x0D — the collector is gated on the backend type.

A backend hands over a run of n bytes drained from the UART FIFO, so now is, within a character time, the arrival of the last of them. Bytes inside one drain are contiguous by construction — a gap wide enough to matter would have fired the ISR first — so they contribute n−1 sub-character gaps to bucket 0, and only the gap between drains can be a frame boundary.

Three implementation rules carry the design:

  1. The hook sits before the FLOW pause gate, deliberately. A monitor channel ships no bytes north and is exactly where the feature earns its keep.
  2. Unarmed cost is one boolean test, and in particular no clock read. portal_rx_push_serial() timestamps itself after the open/paused guard, so a paused channel never pays for a timestamp it would throw away — the cycle read takes a spinlock on SysTick and is one of the costlier single items in a per-character ISR. Arming is one-shot and bounded, so the steady state is unchanged.
  3. The ISR only accumulates. Bucket thresholds are precomputed in µs at arm time from the applied baud and framing, so there is no division in the hot path. The report is built in thread context from a snapshot taken under irq_lock(), and the window expires on the sender thread's existing 1 ms tick rather than on a timer of its own.

Counters saturate rather than wrap, and saturation is reported through the overflow flag: a clipped histogram keeps the shape being read, a wrapped one inverts it.

Sizing ​

  • RX ring GW→Brain: LIN ≥ 64 records per channel; CAN ≥ 256 frames (≈ 90 ms at 100 % load on a 250 k bus); serial ≥ 2 kB per open channel, which holds several 300 B bursts of the 19200-baud profile.
  • TX queue Brain→GW: CAN ≥ 64 frames; serial ≥ 1 kB per open channel.
  • The image must fit the MCUboot secondary slot, which is trivially true.

Footprint freeze. The serial images sat at roughly 65 % RAM on the 63 K L431 muxen_can-modbus when that board was still built. That figure still governs, because app_portal_lin lands on the same 63 K L431 in muxen_can-lin: the firmware is not meant to grow further — the U5 boards have ample headroom, the L431 does not. Raising ring or buffer sizes, or adding features to the portal apps, needs a RAM-budget review first; prefer daemon-side solutions. Each app records its measured footprint in its CHANGELOG. Ring gating by Kconfig keeps each image free of the rings its backend never uses.

Build ​

Standard MUXEN Zephyr workspace: west build from the app folder, docker zephyr-build image via the project Makefile, one CI job per image in a (BOARD, APPFOLDER) matrix.

JobBOARDAPPFOLDERArtifact
portal_fw_cancanmuxen_cancanapp_portal_canportal-muxen_cancan-<v>.srec
portal_fw_cancan_v2muxen_cancan_v2app_portal_canportal-muxen_cancan_v2-<v>.srec
portal_fw_serial_v2_vedirectmuxen_can-serial_v2app_portal_vedirectportal-muxen_can-serial_v2-vedirect-<v>.srec
portal_fw_serial_v2_modbusmuxen_can-serial_v2app_portal_modbusportal-muxen_can-serial_v2-modbus-<v>.srec
portal_fw_can_linmuxen_can-linapp_portal_linportal-muxen_can-lin-<v>.srec
portal_fw_can_mdv_v2muxen_can-mdv_v2app_portal_linportal-muxen_can-mdv_v2-<v>.srec
portal_fw_enoceanmuxen_can-enoceanapp_portal_vedirectportal-muxen_can-enocean-<v>.srec

Only the combined serial board carries an app suffix, because the board alone does not disambiguate its two images; the rest keep the plain form.

Three build-time rules exist because breaking any of them is silent:

  • The consuming side must name the job too. .fw_optional_needs in the parent .gitlab-ci.yml is what pulls an artifact into the srec cache and thence into the .deb. A job present in the firmware matrix and missing there builds green and ships nothing, because an absent optional need is a satisfied one.
  • The product jobs are not a template to copy. They set -DCONFIG_MCUBOOT_GENERATE_CONFIRMED_IMAGE=y and ship zephyr.signed.confirmed.hex, and some build a bootloader and emit a -usine artifact. The portal jobs must not set it and must ship zephyr.signed.hex, and build neither a bootloader nor a -usine artifact. Copying a product job verbatim is the single most likely way to break the revert guarantee — property 1 above.
  • A swap-using-offset board needs its --slot-size. Its board .conf must carry both CONFIG_MCUBOOT_BOOTLOADER_MODE_SWAP_USING_OFFSET=y and CONFIG_MCUBOOT_EXTRA_IMGTOOL_ARGS="--pad --slot-size <exec slot>". On a board whose secondary slot is the primary plus one sector, imgtool must be told the primary size, or it pads to the larger secondary size and the --pad trailer lands past the end of the executable partition — which the device refuses to write, failing the upload at the last block. Writing a board conf that says "no board-specific tuning" is the mistake to avoid.

No product overlay. Product apps build with a per-ProductId overlay; the portal is a board-level image and not a product, so it builds against the plain board DTS with no overlay. Per-board tuning lives in the app's boards/<board>.conf, picked up automatically by Zephyr.

All families pin Zephyr v4.4.0 with HWMv2 board files. The portal manifest is the one that applies — a product repo pinning an older Zephyr in its own manifest does not affect the portal build, and the bootloaders are untouched and do not care what Zephyr the app uses.

Bootloader modes in the fleet, taken from the product bootloader configs: scratch swap on muxen_cancan and muxen_can-lin; swap-using-offset on muxen_cancan_v2, muxen_can-serial_v2, muxen_can-mdv_v2 and muxen_can-enocean. There is no downgrade prevention and no rollback counter, and image header versions are 0.0.0 fleet-wide, so version checks never block a portal upload.

Note that a board's swap algorithm is a property of the MCUboot binary programmed at manufacture, not of the uploaded image: it lives below the primary slot and no firmware upload replaces it, because uploads go to the secondary slot. A .conf line therefore matches the current product build, and units flashed with an earlier bootloader generation can disagree — with a silent failure mode described in Troubleshooting.

Product firmwares stay bootable after revert: the cancan apps self-confirm at boot, and the serial and LIN families ship pre-confirmed images.

Extending ​

  • A new product on an existing board needs nothing.
  • A new board on an existing backend adds a boards/<board>.conf to the matching app, one HardwareId mapping in the daemon, a firmware CI job and its entry in .fw_optional_needs — with FD advertised in HELLO if the silicon supports it.
  • A genuinely new electrical interface adds a fifth app implementing struct portal_backend. If it is only a new wire, it touches nothing in libportal; if it is a new protocol shape, as LIN was, libportal learns the family too.

Integration of multiplexed solutions
MUXEN and the MUXEN logo are trademarks of MUXEN SAS.