Appearance
How the engine runs automations
Everything on board runs on one small computer, and automations are code the boat's owner or integrator can change. The engine is built so that a badly written automation costs you that automation and nothing else: not the other automations, not the daemon, not the boat.
This chapter is what the engine guarantees, what it does when something goes wrong, and the behaviour you need to predict when you deploy or diagnose an automation. It is behaviour, not implementation — the design record is in internal/architecture.md.
One process, many isolated automations
muxen-synapsed is a single daemon. Inside it, each automation gets its own Lua interpreter. There is no shared one.
That means a script's globals, functions and helpers are completely invisible to every other script. Two automations can both use a variable called pump_on, or redefine the same helper, with no interference. Tearing one down frees only its interpreter and touches nothing else.
Execution is single-threaded. Every callback — MQTT messages, timers, lifecycle hooks, popup answers, state-machine transitions — is dispatched one at a time from a single loop. No two automations ever run at the same instant, and no automation ever runs re-entrantly. The isolation is about namespaces, not parallelism.
The practical consequence: a callback that does not return blocks everything. Which is why there is a watchdog.
The watchdog
Every entry into Lua gets a hard budget of 50 ms of CPU time. That covers the top-level chunk at load, onStart and onStop, every MQTT and timer handler, every state-machine transition, every popup answer, and every validation run.
The cap is engine-wide and a script cannot raise its own budget. It is not configurable per automation.
Three details matter operationally.
It is CPU time, not wall-clock. The Brain also runs the kiosk and the frontend, which can spike the load. Measuring wall-clock would abort a perfectly lean script that merely got descheduled during a spike, so the budget is charged against the script's own CPU time. Only genuinely CPU-bound Lua hits the cap. (Snoozes and heartbeats still use wall-clock — those are real-time concepts.)
An overrun is treated like any other script error. The callback is aborted, the failure is logged to the journal and recorded in the automation's last_error, and it counts toward auto-disable. The daemon keeps running throughout.
A script cannot escape it. pcall, xpcall and coroutine are absent from the sandbox precisely because they are the only ways a script could catch an error — including this one. With them gone, the abort always unwinds to the engine. And there is nothing else to block on: every synapse.* binding is non-blocking (publish-and-return, register-and-return), and the blocking standard-library surfaces — os.execute, io, sockets — are not present at all. A long Lua computation is the only remaining way to overrun, and that is exactly what the watchdog bounds.
Memory is capped separately. The instruction-counting mechanism cannot interrupt a single large allocation, so each automation also has a 16 MB Lua heap ceiling. Past it, the script gets an out-of-memory error rather than the daemon getting killed. The ceiling is adjustable through the environment (SYNAPSE_LUA_MEM_MAX, in bytes; 0 disables it).
Faults and auto-disable
The engine counts consecutive faults per automation — script errors and watchdog timeouts alike. A successful callback resets the count. At five, the engine calls onStop, tears the automation down, and stops running it.
This is a runtime fault state, not a change to what you asked for: enabled stays true, running goes false, and last_error says why. It is not persisted, so a restart with fixed code retries automatically; an explicit enable command also clears it and retries immediately. Full detail in Deploying automations.
Other error handling that keeps one failure local:
- A script that throws or times out during
onStartis marked errored and is not left half-subscribed — the engine rolls back any partial registration. - A script that throws during
onStopis logged, and the engine still tears down its wiring and destroys its interpreter. - A syntactically invalid script does not stop the other instances from loading.
Dispatching to many automations
Several automations may subscribe to overlapping topics and register independent timers. The engine tracks, per automation, exactly the patterns and timer handles it owns, so:
- an incoming message is delivered to every automation whose pattern matches;
- disabling one automation removes only its subscriptions and timers;
- the broker subscription for a topic is reference-counted across automations, and is dropped only when the last one stops needing it.
Only automations in the running state receive anything. Stopped and errored ones are skipped.
Status and the heartbeat
Each automation publishes its state on synapse/<id>, retained, so a frontend connecting later immediately sees the truth. It is republished on two triggers:
- On change — enabled, disabled, errored, popup raised or cleared.
- On a 30 s heartbeat — every instance that owns a status topic, running and runtime-disabled alike, even when nothing changed.
The heartbeat exists because of the standard MUXEN freshness check. Each status carries a metadata block with a publish time and a 60 s expiry window. Without a heartbeat, an automation whose state machine sat in one state for hours would look expired to any consumer applying that check — indistinguishable from a dead daemon. With it, a consumer can always tell "disabled but the daemon is healthy" from "no update from any automation, so the daemon is gone".
Two kinds of automation own no status topic, and their retained message is cleared instead of republished: one deleted from deploy.json, and one forced off by deploy enabled: false. Both drop off the frontend list entirely. A runtime-disabled automation keeps its topic, published with enabled: false, so it stays listed as present-but-off. The engine reconciles this at connect time against the retained messages the broker replays.
Popups
Some automations must ask before acting. The whole flow is asynchronous and runs over MQTT:
script calls synapse.ui.prompt{ title, content, actions, key, onAnswer }
│
▼ engine assigns a requestId and attaches the popup to the status
synapse/<id> (retained) = { …, popup: { requestId, title, content, actions } }
│
▼ frontend renders it; the crew taps a button
synapse/<id>/feedback = { requestId, action, duration? }
│
▼ engine matches the requestId and invokes onAnswer(action, duration)
│
▼ engine clears the popup from the status; if the action snoozes, arms itThe rules that shape what a script can do:
- One popup pending per automation. A second
promptwhile one is pending is rejected — it returnsfalseand does nothing. It neither replaces nor queues. This keeps the statuspopupfield a single object and makes correlating the answer trivial. - The
requestIdexists only to reject stale answers. Since at most one popup is ever pending, an answer whose id does not match the currently pending popup — already answered, cleared, or timed out — is ignored. - A snooze suppresses re-prompting. An ignore-type action carries a duration; the engine arms and persists the snooze itself. While it is active, a
promptwith the samekeyresolves immediately as"snoozed"rather than raising anything. Snoozes live instate.json, so they survive a restart and an explicit disable. - Every popup times out. Default 300 s unless the prompt sets its own
timeout. On expiryonAnsweris invoked with"timeout"and the popup is cleared. - Disabling clears a pending popup and
onAnsweris not called.
Timers and the clock
Script timers and state-machine after/every triggers are monotonic. A GPS clock correction does not distort an interval — "30 seconds from now" stays 30 seconds.
Snoozes, by contrast, are wall-clock: they expire at an absolute instant. A clock jump can therefore shorten or lengthen a snooze window. A snooze and a state-machine after modelling "the same" window can diverge across a jump; the timer is the one to trust for safety logic.
Astronomical events are edge triggers scheduled from muxen-sextant data, and they re-arm each day as new data arrives. A forward clock jump that lands past an armed instant still fires it exactly once — no double fire, no skip — and the next arming uses the corrected times.
Where synapse takes "now" from depends on where it runs. On the Brain it uses the local clock, which is the very clock the publishers stamp with, so freshness maths is exact. Run against a remote broker and it instead tracks the boat's own clock from the system/time topic, extrapolated between ticks, so a laptop with a skewed clock does not read every value as stale. The default is automatic — a loopback broker host means local, any other host means boat time — and is overridable (--prefer-mqtt-time / --prefer-local-time). The resolved choice is printed in the startup banner.
Validation of submitted scripts
raken can check an inline script before saving it. The engine loads the submitted source in a throwaway interpreter built with the same sandbox and the same strict globals as a real instance, but with every synapse.* I/O binding stubbed out and onStart never called. Running the top-level chunk therefore has no side effects — no subscriptions, no commands — and only surfaces load, syntax and strict-global errors while capturing the config schema.
The validate topic is open on the local broker like every other MUXEN topic, so it is hardened against abuse: the chunk runs under the same 50 ms CPU watchdog (a submitted while true do end aborts with an error result rather than stalling the loop), and at most one validation runs at a time — a concurrent request is rejected outright.
The same watchdog covers the load-time chunk of a real instance, so a pathological template cannot hang the daemon at load either.
Edge cases, by category
The guiding rule throughout: one bad input or instance never takes down the daemon or the others.
Resolving instances from deploy.json
| Case | Behaviour |
|---|---|
| Duplicate instance id | First wins; the duplicate is rejected and logged |
Instance id not matching [a-zA-Z0-9-]+ | Rejected and logged |
template name not in the library | That instance is errored; the others load normally |
Both template and script, or neither | Invalid — errored and skipped, logged |
Inline script fails to compile | That instance is errored, isolated from the rest |
Inline script over ~64 KB | Warned, but still runs |
| Config value out of range, or wrong type | Clamped or coerced (schema default if uncoercible), warned. A key absent from the schema is kept as-is with a warning |
Deploy enabled: false | Hard force-off — not loaded, runtime enable refused, retained status cleared |
No synapse section at all | Nothing runs; the template library alone starts nothing |
| Template with no config schema | Legal; its manifest entry has an empty config array |
At runtime
| Case | Behaviour |
|---|---|
| Broker disconnects | The client reconnects; on reconnect every active subscription is re-issued and all retained statuses are republished |
| A device command issued while disconnected | QoS 0, fire-and-forget — may be silently dropped. Not queued. Make commands idempotent and re-assert them |
| Two enabled instances commanding the same channel | No arbitration — last-write-wins at the device. An advisory warning is logged at boot when two enabled instances declare commands to the same target |
synapse.variable.get on an unseen or stale value | Returns nil — the script must handle it |
| Device command with an out-of-range function code, instance or channel | Rejected and logged; nothing is published. functionCode 0–63, instance 0–63, channel ≥ 1 with the upper bound per device |
| Timer interval ≤ 0 | Rejected and logged; no timer armed |
fsm:fire of an event invalid in the current state | No-op, returns false |
| Event fired during a transition | Queued, processed in order after it completes — no re-entrancy |
| FSM resume onto a state the redeployed code no longer defines | Falls back to initial (firing its enter), warns, sets last_error |
Time and persistence
| Case | Behaviour |
|---|---|
state.json missing or corrupt at boot | Start with empty state, log it, do not crash |
state.json write fails | Log it; keep state in memory. StateDirectory is the only writable path under ProtectSystem=strict |
| System clock jump | Timers are monotonic and unaffected. Snoozes are wall-clock and can shorten or lengthen |
Stored state for an instance no longer in deploy.json | Dropped during boot reconciliation; its retained status topic is cleared too |
| SIGTERM | state.json is persisted first, then onStop runs for each running instance within a 10 s aggregate budget. If the budget is exceeded, remaining hooks are skipped and the daemon exits cleanly |
Lifecycle and popups
| Case | Behaviour |
|---|---|
onStart throws or times out | Partial registration rolled back; instance errored |
onStop throws or times out | Logged; wiring still torn down and the interpreter still destroyed |
| Instance disabled or removed with a popup pending | Popup cleared; onAnswer is not called |
Feedback with a non-matching or expired requestId, or for an unknown instance | Ignored |
| Two instances of one template both prompt | Independent — each has its own state, config and single popup slot |
