Skip to content

Troubleshooting ​

This chapter is organised by what you see. The recorders have no status command, no control socket and no MQTT topic: everything they know they say at startup, and after that the evidence is the directory listing. Those two are the whole toolbox.

sh
systemctl status muxen-can-datalogger@can0.service
journalctl -u muxen-can-datalogger@can0.service -n 50
ls -l /var/lib/muxen/datalogger/
df -h /var/lib
candump can0                       # what the recorder is seeing

One thing to know before reading any of the entries below: a broken recorder does not show as failed. Every failure the recorders handle exits with status 0, so Restart=always brings the process back and systemctl shows activating (auto-restart) or active throughout. Anything that watches for failed units will not see it. Read the journal, not the unit state.

The service restarts every 30 seconds ​

Restart=always with RestartSec=30, so a recorder that fails at startup loops on that period. The journal names the reason each time.

MessageCauseFix
config: Ensure logFolder exists: <path>the log folder is not therethe package creates /var/lib/muxen/datalogger; if you passed --logFolder, create the directory yourself — the recorder does not
config: Check logFolder permissions: <path>the folder is not both readable and writable for this processcheck ownership and mode; on the CSV unit see the CSV unit cannot write below
busmaster: failed to init the file counter, check log folder permissionsthe folder could not be opened for readingsame causes, caught one step later
busmaster: failed to create file: <path>the new capture could not be createdthe filesystem is full or read-only. df -h /var/lib
csv: failed to create file: <path>the same, for a decoded fileas above
usage block, then exitan unparseable option, or an empty --interfacecheck the drop-in you added; note --gzip and --no-gzip need a value, see below
nothing at all in the journalthe unit's start condition is not metsee the unit does nothing when started

The unit does nothing when started ​

systemctl start muxen-can-datalogger@can9.service returns without error, nothing runs, and the journal is empty.

Both units carry ConditionPathIsDirectory=/sys/class/net/%i. If the interface named after the @ does not exist on this Brain, systemd skips the unit rather than starting and failing it. That is deliberate: an instance for a bus this Brain does not have stays quiet.

sh
systemctl status muxen-can-datalogger@can9.service   # "Condition: start condition failed"
ls /sys/class/net/                                   # what does exist
ip link show can0

The instance name is the interface name and nothing else. There is no other meaning to it in either unit.

The CSV unit cannot write, or starts the wrong program ​

Two separate problems in one unit, up to and including 2.3.3. Both are fixed in later versions; this section is for a boat still running an older package.

It started the CAN recorder. ExecStart named /usr/bin/muxen-can-datalogger, so enabling the unit gave a second raw recorder, not a decoded one — no CSV files appeared, however long you waited.

Its hardening blocked the log folder. ProtectSystem=strict with no ReadWritePaths= and no StateDirectory= left the filesystem read-only for it. The recorder's startup write test failed and it exited with config: Check logFolder permissions: /var/lib/muxen/datalogger, every 30 seconds, without ever being marked failed.

On an older package, run the decoded recorder by hand and do not enable its unit:

sh
sudo /usr/bin/muxen-csv-datalogger --interface can0

On a current package both units start their own binary, run as muxen, and get /var/lib/muxen/datalogger created and owned for them by StateDirectory=. If a current unit still reports the permissions error, check that the muxen user exists (id muxen) and that muxen-systemd is installed — that package creates it.

The capture is corrupt — interleaved or truncated frames ​

Two recorder processes are writing the same file. They both scan the folder at startup, both compute the same next index, and both open it with fopen(…, "w+").

The usual cause is enabling both units on the same interface, since both start muxen-can-datalogger (above). Check:

sh
pgrep -a muxen-can-datalogger
systemctl list-units 'muxen-*datalogger@*'

Expect exactly one process per interface. Stop the extra instance; the next rotation starts a clean file.

No files appear at all ​

Likely causeWhat to checkFix
The recorder is not runningsystemctl status, pgrep -a muxen-can-dataloggersee the restart-loop entries above
It is running but pointed elsewhereconfig: logFolder = in the journala drop-in with --logFolder
It is running on a different interfaceconfig: interface = in the journalthe instance name is the interface
The unit was never enabledsystemctl is-enabled muxen-can-datalogger@can0.servicesystemctl enable --now …

Note that the raw recorder creates its first file at startup, before any frame arrives. An empty folder therefore means the recorder is not running, not that the bus is quiet. The decoded recorder is the other way round: it creates a file only when a frame for that device arrives, so an empty folder there does mean silence.

The file exists but does not grow ​

The bus is silent as far as this Brain is concerned.

sh
candump can0            # nothing here means nothing for the recorder either
ip -details link show can0

The recorder and candump read the same socket. If candump is quiet, the problem is upstream: the interface is down, the bitrate is wrong, the wiring is wrong, or the bus really is idle.

One case where candump shows traffic and the capture does not: error frames are dropped. The raw recorder discards any frame carrying the error flag before writing anything, so a bus in an error storm — which candump will show — produces a capture that looks empty. That is the one traffic pattern the capture cannot represent.

A decoded file has timestamps and nothing else ​

The header line is datetime, and every row is a timestamp followed by an empty field.

That combination of function code and broadcast identifier is not one the decoder knows. It still gets a file, because a file is created for every distinct combination seen on the bus; there are simply no columns to fill. The list of frame kinds that do decode is in File formats.

The file is not useless — it still records that the frame arrived, and when. For the payload, use the raw capture.

There are far more files than expected ​

The decoded recorder creates one file series per distinct combination of function code, instance and broadcast identifier seen on the bus, not one per device. A device that sends four different frames produces four series, and every series rotates on its own timer. Several dozen files per rotation period is normal on a real boat.

sh
ls /var/lib/muxen/datalogger/ | wc -l
ls /var/lib/muxen/datalogger/ | sed 's/\.[0-9]*\.log.*//' | sort -u | wc -l

The second command counts the series. If that number is a surprise, the bus has more distinct broadcast frames on it than you thought — which is itself worth knowing.

Files are not being compressed ​

Check that /usr/bin/gzip exists and is executable. Compression is not optional and cannot be switched off, so an uncompressed rotated file means the exec failed.

That failure has a second symptom that matters more: the process that was supposed to become gzip does not exit. It returns into the recorder's code and runs on as a duplicate recorder, one more per rotation.

sh
ls -l /usr/bin/gzip
pgrep -c muxen-can-datalogger      # should equal the number of instances
dpkg -l gzip

If the count is climbing, stop the service, kill the strays, restore gzip, and start again.

--no-gzip has no effect ​

It has none. The option is parsed and the resulting setting is never read anywhere, so compression always happens.

There is a second trap in the same pair of options: both --gzip and --no-gzip are declared as taking a required argument. Writing either of them bare makes the option parser fail, which prints the usage block and exits 0 — under systemd, a restart loop with a usage block in the journal. If you must pass them, pass a value (--no-gzip=1); it changes nothing either way.

--maxRecordTime did not take the value I gave ​

Values above 86400 are silently clamped to 86400 — one day is the hard ceiling on the age of a file. The startup summary prints the value actually in force:

config: maxRecordTime = 86400

--maxFileSizeBeforeCompression behaves the same way in the other direction on the raw recorder: values below 1048576 are raised to 1 MiB. The decoded recorder does not apply that floor, so a very small value there really does produce very small files.

Both values are read with a plain integer conversion. A non-numeric argument becomes 0 rather than an error: --maxRecordTime abc gives an age limit of zero, and the file then rotates roughly once a second for as long as frames keep arriving.

The disk filled up ​

See Disk budget — the whole chapter is about preventing this. In the moment:

sh
df -h /var/lib
du -sh /var/lib/muxen/datalogger
ls -1t /var/lib/muxen/datalogger/ | tail -5     # oldest

Stop the recorders, copy off anything worth keeping, delete the oldest compressed files, then put a retention rule in place before restarting. Deleting every file resets the file index to 1, which is harmless but makes the numbering discontinuous with what you copied off.

The timestamps are wrong ​

Three separate things people mean by this:

  • They are not local time. They are UTC, in both formats, always. There is no timezone setting.
  • They are absurd, or they jump. The recorder uses the Brain's clock. A Brain that boots without a valid clock and is corrected by NTP a minute later produces a capture with a jump in it — and the jump also triggers an immediate rotation, because the age of the file is measured against the arriving frame's timestamp.
  • The end date is a month early. The closing line of a raw capture prints the month one lower than the real month. The start date in the header is correct. It is a formatting fault in the footer only; see File formats.

Check the clock before trusting a capture for anything time-sensitive:

sh
timedatectl

The capture header says Kvaser and 500000 bps ​

It always does. That line is a fixed literal in the header, written identically whatever interface is being recorded and whatever bitrate it is running at. It names hardware that is not present.

It is there because the format is BUSMASTER's and the field is mandatory. Do not read a bus configuration out of it, and do not conclude from it that the capture came from a Kvaser adapter. The recorder does not know the bitrate and records it nowhere.

The last capture has no end marker and is not compressed ​

The recorder was killed, or the Brain lost power. The closing lines and the compression both happen on a clean shutdown; neither happens on SIGKILL or a power cut.

The file is still readable. Every row was flushed as it was written, so the content is intact up to whatever the kernel had managed to write to storage — expect the last seconds to be missing after a power cut. There is no fsync() anywhere in the write path.

Compress it by hand if you want it out of the way:

sh
gzip --best /var/lib/muxen/datalogger/can0.7.log

The next start picks the next free index and does not touch it.

Bash completion does nothing ​

Neither completion file registers anything for these commands. Both are copies of the muxen-boat completion: they guard on muxen-boat being present and they complete muxen-boat. Typing muxen-can-datalogger --<TAB> produces nothing on any Brain.

The options are in Reference.

Replay produces no traffic ​

See Replay. The short version: a missing or unreadable capture file is not reported, and the process sits idle instead of exiting. Check the path first.

FAQ ​

Is the boat recording right now?systemctl is-active muxen-can-datalogger@can0.service answers it, and ls -l /var/lib/muxen/datalogger/ shows the file that is currently being written — it is the one whose name ends in .log rather than .log.gz, and its size grows.

How long is the history kept? As long as the disk lasts, unless somebody set up a deletion rule. Nothing in this package removes a file. If the boat was commissioned by MUXEN there is usually a rule under /etc/tmpfiles.d/; if there is none, the history goes back to the day recording was switched on.

Can I turn it off?sudo systemctl disable --now muxen-can-datalogger@can0.service. The existing files stay where they are. Nothing else on the boat depends on the recording, so switching it off changes nothing visible.

Does recording slow the boat's systems down? It listens to the bus and writes text to disk; it transmits nothing and talks to no other service. The cost is disk space, not responsiveness — and disk space is the thing to watch, because a full disk does affect everything.

Can I read the files on my laptop? Yes. Copy them off with scp, expand them with any unzip tool, and open a .log in a text editor, a BUSMASTER installation, or — for the decoded ones — a spreadsheet. Nothing MUXEN-specific is needed to read them.

Do I have to stop the service to copy the files off? No. The compressed files are finished and will not change. The one open .log file can be copied too; you get everything up to that moment.

Why are there so many files? Two reasons. Files rotate on time or size, so a long recording is always many files. And the decoded recorder makes a separate file for every kind of frame every device sends, which multiplies quickly on a well-equipped boat.

The times in the file do not match my watch. They are UTC. Add your own offset, or convert them when you import.

We lost power — is the recording ruined? No. The capture that was open at the time is missing its last few seconds and its closing lines, and it was not compressed. Everything already compressed is untouched, and the recorder starts a fresh file on the next boot without overwriting anything.

Does it record anything private? It records CAN traffic: device states, currents, voltages, navigation values. There is no audio, no video, no position beyond what the boat's own instruments broadcast on the bus, and nothing leaves the Brain — the package uploads nothing anywhere.

A device stopped working yesterday. Is it in the capture? If the recorder was running and the device speaks on the recorded bus, yes. Find the file by date, expand it, and search for the device's identifier. Note that a bus error storm is not recorded — error frames are discarded — so "the capture goes quiet" is itself evidence.

Tips ​

At commissioning, read the seven startup lines once. The journal prints the interface, the log folder, both limits and the next file index. Confirming those match what you intended takes ten seconds and prevents most of the entries above.

Enable one recorder, not both. The raw capture answers more questions, costs less disk, and — as shipped in 2.3.3 — enabling the second unit gives you a duplicate raw recorder rather than a decoded one.

Measure the daily volume on the commissioned boat. A bus at the dock with the systems off is not the bus under way. Ten minutes of du -sb before and after is enough; see Disk budget.

Write the expected steady-state folder size on the handover sheet. It turns "the disk is filling" into a five-second check for whoever comes next.

Put a retention rule in place and watch it fire once. A rule that was written but never ran is the same as no rule. systemd-tmpfiles --clean runs it on demand.

Give the folder its own filesystem on a boat that records permanently. It is the only arrangement in which a runaway recorder inconveniences the recording instead of the boat.

Put an end date on a diagnostic recording. Recording enabled for one investigation and never disabled is the ordinary way this package fills a disk.

Check the clock before trusting a capture. Timestamps come from the Brain, in UTC, with no correction and no record of what the clock was doing.

Copy captures off before deleting them, and keep them with the fault report. They are the only history that exists; nothing else in the MUXEN stack keeps one.

Search compressed captures directly. zgrep, zcat and zless work on the rotated files, so there is no need to expand a season of recordings to find one identifier:

sh
zgrep -h ' 0x1040 ' /var/lib/muxen/datalogger/can0.*.log.gz | head

Do not use file indices to order a long history. Deleting old files lowers the next index, so numbering restarts. Sort by modification time.

Keep a virtual interface handy for replay. vcan0 costs nothing and removes any chance of injecting a recorded frame onto a live bus.

Integration of multiplexed solutions
MUXEN and the MUXEN logo are trademarks of MUXEN SAS.