Appearance
Monitoring a session
Three surfaces report what a session is doing: muxen-portal stats (a table, --watch to refresh), muxen-portal status (per-channel detail) and the retained MQTT topic portal/session. All three are built from one builder in the daemon, so they cannot disagree.
Per-port statistics
During a session the Brain reports TX/RX statistics for every port of the gateway, from the device's point of view — what the gateway sees on the wire, not what reached the Brain. Ports nobody has opened locally are included.
That matters because on a serial gateway a line carries no traffic toward the Brain unless something reads its PTY. Statistics that depended on a local reader would be blind exactly where an operator needs them: is anything talking on port 3?
$ muxen-portal stats --watch
port state rate RX bytes RX B/s TX bytes TX B/s errors err/s
1 in use 19200 184320 1920.0 512 4.0 0 0.0
2 monitor 19200 9600 96.0 0 0.0 1 0.0
3 monitor 19200 40000 400.0 0 0.0 31940 310.0
4 closed - 0 0.0 0 0.0 0 0.0
note: a port is receiving but with a high error rate — the configured baud is probably wrong.state is the word that matters operationally:
| State | Meaning |
|---|---|
in use | a mirror exists and something is reading it |
monitor | the gateway is counting, the stream is paused, there is no PTY |
idle | open, mirror exists, no reader |
closed | the UART is off — counters physically cannot move |
CAN sessions render bus state, TEC/REC and frame counters instead. LIN sessions render frames rather than bytes, plus the per-channel role and a column of their own for headers that drew no answer:
port state mode baud RX frames RX/s TX hdrs TX/s no-resp errors err/s
1 in use master 19200 1840 18.0 1860 18.0 2 0 0.0
2 monitor monitor 19200 960 9.6 0 0.0 0 1 0.0
3 closed - - 0 0.0 0 0.0 0 0 0.0no-resp is deliberately not folded into errors: on a master channel it is how an absent slave reads, and a normal scan of an empty ID range would otherwise look like a broken bus.
Three layers, kept distinct
Every surface separates them, because conflating them is how a healthy session gets misread as a broken one:
| Layer | Source | Answers |
|---|---|---|
| Wire-side | gateway STATUS counters | what the remote device is actually transmitting |
| Tunnel-side | tunnelTxMsgs, seqGapLost | whether the tunnel is carrying it |
| Mirror-side | daemon ring counters | whether a local process is consuming it |
rxBytes climbing while noReaderDropped climbs and flowPaused is true is the normal state of a monitored port — not a fault.
Monitor channels
The gateway counts serial RX in the UART interrupt, and that interrupt is enabled only by OPEN. A channel that was never opened has its UART off and physically cannot count — there is no baud-agnostic byte counting. So statistics on an unrequested port require opening it.
On a serial or LIN session the daemon therefore monitor-opens every channel the gateway advertises beyond the requested one:
- opened at the session's rate, so the gateway enables the UART and starts counting;
- no PTY is created — a monitor channel never claims
<name>nor a<name>-chNlink; - the stream is paused immediately, so the gateway ships no data. Paused, it keeps listening and counting, which is the whole mechanism this feature rests on.
Cost is one STATUS per second per port on the tunnel's lowest-priority ID band, and nothing else.
Monitor opens are best effort: a channel that refuses — nothing wired, a rate the board will not take — is logged, left closed, and never retried. They also run after the session is active rather than during bring-up, so an unused port can never fail the whole session.
The same holds on LIN, where it bites harder: on a master channel nothing on the bus moves at all unless the gateway sends headers. Monitor-opening a LIN channel therefore uses monitor mode, so inventorying an unused port cannot disturb a bus somebody else owns.
open-channel N promotes a monitor channel: the re-OPEN creates the PTY and hands flow control back to the reader-presence logic. There is no demotion.
Recovery after a session loss re-opens only the real channels; monitor channels are re-established best-effort once the session is active again.
Disable the whole behaviour with monitor-channels = false (or --no-monitor-channels); ports then stay dark unless explicitly opened and report open: false.
Wrong baud is a readable signal
A live line read at the wrong baud still raises RX interrupts. Nonzero rxRate together with a high errRate therefore means something is transmitting here, but the baud is wrong — a genuinely useful answer, and the reason monitor-opening at a possibly-wrong default is worth doing at all.
The daemon exposes this as baudSuspect (error rate above 10 % of the byte rate), which stats renders as the one-line note above.
Wire timing — beyond "how much"
Counters answer how much traffic a port carries. They cannot answer what shape it has, and shape is what identifies a protocol. A line carrying 1840 bytes in five seconds is a Modbus RTU segment under a polling master, or an NMEA 0183 talker, or a VE.Direct block stream, and the byte totals are identical in all three.
Shape lives in the inter-byte gaps, which is to say in the gateway's UART interrupt: the tunnel preserves byte order and contiguity, not silences, so every gap observable on the Brain is an artefact of tunnel batching. muxen-portal timing closes that gap by measuring where the timestamps already are and returning a histogram rather than a stream.
sh
muxen-portal timing 2 [--window-ms N] [--arm|--collect] [--json]It is the same idea as monitor channels, one level up: monitor-opening a port makes the gateway count it, timing makes the gateway characterise it — and it costs almost as little:
- No PTY, no reader, no data plane. The reply is 54 bytes on the control pair, once. Nothing is streamed in order to be measured.
- It works on a paused monitor channel, which is the state of every port nobody explicitly opened. The shape of all four ports is obtainable without creating a single mirror.
- Gaps are bucketed in character times, computed on the gateway from the applied line settings, so the 3.5-character Modbus inter-frame boundary is an exact bucket edge whatever the baud.
Serial gateways only. CAN framing is inherent to the bus and LIN is not covered; both refuse the command with a plain message rather than a session fault.
Reading the histogram
| Shape | Reading |
|---|---|
Dense <1.5, empty 1.5–3.5, populated beyond | Bursts separated by silence — and nothing more than that |
…with short bursts (≤ 40 B) and a populated 3.5–20 band | Request/response: the turnaround between a master and its slave. On RS-485, the Modbus RTU shape |
…with long bursts (≥ 89 B) and an empty 3.5–20 band | A block streamer talking unprompted — VE.Direct, or NMEA 0183 at a low rate. Nothing is answering anything |
burst_gap min ≈ mean ≈ max, on a framed shape | A master polling on a fixed cycle; the mean is the poll period |
Gaps spread smoothly across 1.5–20 | A continuously streaming protocol |
bursts == 0 with the channel open | Nobody is on the wire — not a dead port: a Modbus slave with no master reads exactly like this |
Gaps spread with baudSuspect set | Read the baud first; shape means nothing at the wrong rate |
Bimodality alone is not a request/response signature. A dense-<1.5 / empty-1.5–3.5 / populated-beyond histogram only says "bursts separated by silence". Any device that emits a block then falls silent produces two modes — a Victron MPPT on a plain serial line, with no master anywhere on the segment, produces exactly that shape. Two further tests separate them, and this is why both histograms travel in the same message:
- Burst length. A request/response exchange trades short frames — an 8-byte Modbus request, a
5+2Nreply. A block streamer emits one long run per cycle. - The middle of the histogram. A master and its slave are separated by a turnaround of a few character times, landing in
3.5–8and8–20. A block streamer leaves that band empty and jumps straight from "contiguous" to "a second away".
The same corollary applies to regularity: gap regularity alone is not a poll period. A 1 Hz block streamer's blocks are as evenly spaced as any master's cycle. The period is only meaningful once the shape is already known to be an exchange.
It answers a different question on each serial backend
Worth knowing before reading a report, because the expected result differs — and the reason is the portal itself: during a session the gateway runs portal firmware, not its product firmware.
RS-232 / VE.Direct — characterisation. These devices stream unprompted and do not care which firmware the gateway is running: a Victron MPPT keeps emitting its ~1 Hz text block throughout the session. The histogram is full and does what it was built for: confirms the line is alive, confirms the baud is right, and separates a block-streaming protocol from a framed one. This is where timing pays for itself as an identification tool, and it covers NMEA 0183 on a serial channel for the same reason.
RS-485 / Modbus — safety interlock. A Modbus slave says nothing until polled, and on a Muxen-native segment the poller is the gateway's own firmware, which the session just replaced. The expected reading is bursts == 0: the bus is quiet because the master was removed, not because nothing is connected. That is the measurement doing the one job that matters on this backend — telling you whether a foreign master (a third-party PLC, another vendor's controller) is on the segment. Nothing else can: the Brain cannot hear the wire, and byte counters cannot tell a polled bus from a chatty one. bursts > 0 here means do not transmit.
So: on RS-232 the histogram identifies the device; on RS-485 it decides whether transmitting is safe. An empty histogram on a Modbus channel is normal and must not be read as a broken port.
Measuring several channels at once
The collector is per channel and so is its deadline, so all channels can measure in the same window. Only a second request on a channel already measuring is refused.
| Mode | Behaviour |
|---|---|
| default | arm, hold the connection, reply with the report |
--arm | arm and reply at once; the report is held until collected |
--collect | take the held report, or wait if the window is still open |
--arm/--collect exist so one single-threaded client can have every channel measuring at once and spend the window doing something else. muxen-detect uses this: it arms all channels, waits one window, and has the whole gateway's wire shape before it identifies anything — instead of four sequential windows per gateway.
A window is abandoned, with an error to a waiting client, if the session leaves active; the gateway drops its own collector on re-OPEN or CLOSE, so neither side can be left armed.
Counter semantics
The gateway's STATUS fields are monotonic since its boot and are not reset by a re-OPEN, so set-rate preserves them. Two facts make the raw snapshot unusable on its own, and the daemon folds every field into 64-bit session totals accordingly:
- Width. Byte counters are 32-bit, but the line-error counters are 16-bit and wrap in seconds on a wrong-baud line — precisely the case above. Deltas are taken modulo the wire width, which is correct as long as less than one full period elapses between polls; at 1 Hz a 16-bit counter would need more than 65535 events per second to alias, more than a 115200-baud line can produce.
- Reboots. A recovery re-uploads the portal firmware, so the gateway's counters restart at zero mid-session. An
uptimethat goes backwards marks the reset, and the raw value is then the delta.
Rates are per-second, derived from consecutive 1 Hz polls. Totals are per session, zeroed at open and 60 s after a session ends.
Field reference
Per serial channel (status, stats and MQTT share one builder):
| Field | Layer | Note |
|---|---|---|
channel, open, monitored, rate | — | open: false ⇒ every counter is 0 by construction |
rxBytes, txBytes, rxRate, txRate | wire | session totals + B/s |
framingErr, parityErr, overrunErr, errRate | wire | 16-bit on the wire, accumulated here |
baudSuspect | wire | receiving with a high error rate |
txUnderrun, rxDropped | wire | RS-485 underruns; gateway ring drops |
seqGapLost, gwUptime | tunnel | |
flowPaused | tunnel | always true on a monitor channel |
pty, reader, ringDropped, noReaderDropped, termiosRejected | mirror | present only when a mirror exists |
termDriven, termOn | wire | only where the board advertises termination control |
CAN sessions expose busState, tec, rec, rxFrames, txFrames, rxRate, txRate, rxOverflow, txDrops, tunnelTxMsgs, tunnelRxMsgs, errWritesDropped and gwUptime.
Per LIN channel:
| Field | Layer | Note |
|---|---|---|
channel, open, monitored, rate, mode, checksum | — | mode is master / monitor; checksum the applied rule |
rxFrames, txHeaders, rxRate, txRate | wire | responses received / headers emitted |
noResponse | wire | headers that drew no answer — not an error |
checksumErr, framingErr, parityErr, overrunErr, errRate | wire | parityErr is PID parity |
baudSuspect | wire | receiving with a high error rate |
rxDropped, txDrops | wire | gateway ring / TX queue drops |
seqGapLost, gwUptime, flowPaused | tunnel | |
mirror, errWritesDropped | mirror | present only when a vxcan exists |
MQTT session state
Status only, on a single retained topic portal/session, following the standard MUXEN envelope. data is a list of sessions — the schema allows several in parallel, one per gateway, and the daemon owns them all so there is no concurrent-writer issue.
json
{
"data": [
{
"state": "opening | active | recovering | closing | closed | lost | failed",
"uid": "3100410003504257",
"deviceId": 256,
"board": "muxen_cancan",
"mirror": "portal0",
"offline": [ {"function": 4, "instance": 0} ],
"since": 1767800000
}
],
"metadata": {
"rxdate": "2026-01-07T15:33:32.000Z",
"rxTimestamp": 1767800012,
"expireAfterSec": 30
}
}The daemon republishes on every state change and refreshes the topic every 10 s while running.
metadata.expireAfterSec is what makes the retained object trustworthy: an object older than rxTimestamp + expireAfterSec is expired, which means no session information, which means assume no session. An unclean daemon death can therefore never leave a stale active behind — and by the time the object expires, an orphaned gateway is already reverting on its own keep-alive.
offline lists the product identities that stop broadcasting during the session, recorded at the pre-upload scan. Interfaces and alarm logic can show "gateway under maintenance" instead of raw data-loss alarms — but must honour the suppression only while the object is fresh. Devices behind a bus-extender gateway are not listed, because they cannot be attributed to the segment, and will show as plain timeouts.
Serial and LIN sessions add a channels array carrying one object per port the gateway has, open or not, with the port state and the wire-side counters above. Because the topic is refreshed on the 10 s tick rather than on each STATUS, the rates are the useful field there, not the totals.
data: [] means no session; a client connecting mid-session gets the current picture from the retained message. failed marks an open that aborted before becoming active. During recovering the entry and its offline suppression persist — the segment is still down while the daemon re-establishes the session. A finished session stays in the list as closed/lost/failed for 60 s, then is dropped; the daemon republishes a clean list at startup.
