Skip to content

How the engine runs automations ​

Everything on board runs on one small computer, and automations are code the boat's owner or integrator can change. The engine is built so that a badly written automation costs you that automation and nothing else: not the other automations, not the daemon, not the boat.

This chapter is what the engine guarantees, what it does when something goes wrong, and the behaviour you need to predict when you deploy or diagnose an automation. It is behaviour, not implementation — the design record is in internal/architecture.md.

One process, many isolated automations ​

muxen-synapsed is a single daemon. Inside it, each automation gets its own Lua interpreter. There is no shared one.

That means a script's globals, functions and helpers are completely invisible to every other script. Two automations can both use a variable called pump_on, or redefine the same helper, with no interference. Tearing one down frees only its interpreter and touches nothing else.

Execution is single-threaded. Every callback — MQTT messages, timers, lifecycle hooks, popup answers, state-machine transitions — is dispatched one at a time from a single loop. No two automations ever run at the same instant, and no automation ever runs re-entrantly. The isolation is about namespaces, not parallelism.

The practical consequence: a callback that does not return blocks everything. Which is why there is a watchdog.

The watchdog ​

Every entry into Lua gets a hard budget of 50 ms of CPU time. That covers the top-level chunk at load, onStart and onStop, every MQTT and timer handler, every state-machine transition, every popup answer, and every validation run.

The cap is engine-wide and a script cannot raise its own budget. It is not configurable per automation.

Three details matter operationally.

It is CPU time, not wall-clock. The Brain also runs the kiosk and the frontend, which can spike the load. Measuring wall-clock would abort a perfectly lean script that merely got descheduled during a spike, so the budget is charged against the script's own CPU time. Only genuinely CPU-bound Lua hits the cap. (Snoozes and heartbeats still use wall-clock — those are real-time concepts.)

An overrun is treated like any other script error. The callback is aborted, the failure is logged to the journal and recorded in the automation's last_error, and it counts toward auto-disable. The daemon keeps running throughout.

A script cannot escape it. pcall, xpcall and coroutine are absent from the sandbox precisely because they are the only ways a script could catch an error — including this one. With them gone, the abort always unwinds to the engine. And there is nothing else to block on: every synapse.* binding is non-blocking (publish-and-return, register-and-return), and the blocking standard-library surfaces — os.execute, io, sockets — are not present at all. A long Lua computation is the only remaining way to overrun, and that is exactly what the watchdog bounds.

Memory is capped separately. The instruction-counting mechanism cannot interrupt a single large allocation, so each automation also has a 16 MB Lua heap ceiling. Past it, the script gets an out-of-memory error rather than the daemon getting killed. The ceiling is adjustable through the environment (SYNAPSE_LUA_MEM_MAX, in bytes; 0 disables it).

Faults and auto-disable ​

The engine counts consecutive faults per automation — script errors and watchdog timeouts alike. A successful callback resets the count. At five, the engine calls onStop, tears the automation down, and stops running it.

This is a runtime fault state, not a change to what you asked for: enabled stays true, running goes false, and last_error says why. It is not persisted, so a restart with fixed code retries automatically; an explicit enable command also clears it and retries immediately. Full detail in Deploying automations.

Other error handling that keeps one failure local:

  • A script that throws or times out during onStart is marked errored and is not left half-subscribed — the engine rolls back any partial registration.
  • A script that throws during onStop is logged, and the engine still tears down its wiring and destroys its interpreter.
  • A syntactically invalid script does not stop the other instances from loading.

Dispatching to many automations ​

Several automations may subscribe to overlapping topics and register independent timers. The engine tracks, per automation, exactly the patterns and timer handles it owns, so:

  • an incoming message is delivered to every automation whose pattern matches;
  • disabling one automation removes only its subscriptions and timers;
  • the broker subscription for a topic is reference-counted across automations, and is dropped only when the last one stops needing it.

Only automations in the running state receive anything. Stopped and errored ones are skipped.

Status and the heartbeat ​

Each automation publishes its state on synapse/<id>, retained, so a frontend connecting later immediately sees the truth. It is republished on two triggers:

  • On change — enabled, disabled, errored, popup raised or cleared.
  • On a 30 s heartbeat — every instance that owns a status topic, running and runtime-disabled alike, even when nothing changed.

The heartbeat exists because of the standard MUXEN freshness check. Each status carries a metadata block with a publish time and a 60 s expiry window. Without a heartbeat, an automation whose state machine sat in one state for hours would look expired to any consumer applying that check — indistinguishable from a dead daemon. With it, a consumer can always tell "disabled but the daemon is healthy" from "no update from any automation, so the daemon is gone".

Two kinds of automation own no status topic, and their retained message is cleared instead of republished: one deleted from deploy.json, and one forced off by deploy enabled: false. Both drop off the frontend list entirely. A runtime-disabled automation keeps its topic, published with enabled: false, so it stays listed as present-but-off. The engine reconciles this at connect time against the retained messages the broker replays.

Popups ​

Some automations must ask before acting. The whole flow is asynchronous and runs over MQTT:

 script calls synapse.ui.prompt{ title, content, actions, key, onAnswer }
    │
    ▼  engine assigns a requestId and attaches the popup to the status
 synapse/<id>  (retained) = { …, popup: { requestId, title, content, actions } }
    │
    ▼  frontend renders it; the crew taps a button
 synapse/<id>/feedback = { requestId, action, duration? }
    │
    ▼  engine matches the requestId and invokes onAnswer(action, duration)
    │
    ▼  engine clears the popup from the status; if the action snoozes, arms it

The rules that shape what a script can do:

  • One popup pending per automation. A second prompt while one is pending is rejected — it returns false and does nothing. It neither replaces nor queues. This keeps the status popup field a single object and makes correlating the answer trivial.
  • The requestId exists only to reject stale answers. Since at most one popup is ever pending, an answer whose id does not match the currently pending popup — already answered, cleared, or timed out — is ignored.
  • A snooze suppresses re-prompting. An ignore-type action carries a duration; the engine arms and persists the snooze itself. While it is active, a prompt with the same key resolves immediately as "snoozed" rather than raising anything. Snoozes live in state.json, so they survive a restart and an explicit disable.
  • Every popup times out. Default 300 s unless the prompt sets its own timeout. On expiry onAnswer is invoked with "timeout" and the popup is cleared.
  • Disabling clears a pending popup and onAnswer is not called.

Timers and the clock ​

Script timers and state-machine after/every triggers are monotonic. A GPS clock correction does not distort an interval — "30 seconds from now" stays 30 seconds.

Snoozes, by contrast, are wall-clock: they expire at an absolute instant. A clock jump can therefore shorten or lengthen a snooze window. A snooze and a state-machine after modelling "the same" window can diverge across a jump; the timer is the one to trust for safety logic.

Astronomical events are edge triggers scheduled from muxen-sextant data, and they re-arm each day as new data arrives. A forward clock jump that lands past an armed instant still fires it exactly once — no double fire, no skip — and the next arming uses the corrected times.

Where synapse takes "now" from depends on where it runs. On the Brain it uses the local clock, which is the very clock the publishers stamp with, so freshness maths is exact. Run against a remote broker and it instead tracks the boat's own clock from the system/time topic, extrapolated between ticks, so a laptop with a skewed clock does not read every value as stale. The default is automatic — a loopback broker host means local, any other host means boat time — and is overridable (--prefer-mqtt-time / --prefer-local-time). The resolved choice is printed in the startup banner.

Validation of submitted scripts ​

raken can check an inline script before saving it. The engine loads the submitted source in a throwaway interpreter built with the same sandbox and the same strict globals as a real instance, but with every synapse.* I/O binding stubbed out and onStart never called. Running the top-level chunk therefore has no side effects — no subscriptions, no commands — and only surfaces load, syntax and strict-global errors while capturing the config schema.

The validate topic is open on the local broker like every other MUXEN topic, so it is hardened against abuse: the chunk runs under the same 50 ms CPU watchdog (a submitted while true do end aborts with an error result rather than stalling the loop), and at most one validation runs at a time — a concurrent request is rejected outright.

The same watchdog covers the load-time chunk of a real instance, so a pathological template cannot hang the daemon at load either.

Edge cases, by category ​

The guiding rule throughout: one bad input or instance never takes down the daemon or the others.

Resolving instances from deploy.json ​

CaseBehaviour
Duplicate instance idFirst wins; the duplicate is rejected and logged
Instance id not matching [a-zA-Z0-9-]+Rejected and logged
template name not in the libraryThat instance is errored; the others load normally
Both template and script, or neitherInvalid — errored and skipped, logged
Inline script fails to compileThat instance is errored, isolated from the rest
Inline script over ~64 KBWarned, but still runs
Config value out of range, or wrong typeClamped or coerced (schema default if uncoercible), warned. A key absent from the schema is kept as-is with a warning
Deploy enabled: falseHard force-off — not loaded, runtime enable refused, retained status cleared
No synapse section at allNothing runs; the template library alone starts nothing
Template with no config schemaLegal; its manifest entry has an empty config array

At runtime ​

CaseBehaviour
Broker disconnectsThe client reconnects; on reconnect every active subscription is re-issued and all retained statuses are republished
A device command issued while disconnectedQoS 0, fire-and-forget — may be silently dropped. Not queued. Make commands idempotent and re-assert them
Two enabled instances commanding the same channelNo arbitration — last-write-wins at the device. An advisory warning is logged at boot when two enabled instances declare commands to the same target
synapse.variable.get on an unseen or stale valueReturns nil — the script must handle it
Device command with an out-of-range function code, instance or channelRejected and logged; nothing is published. functionCode 0–63, instance 0–63, channel ≥ 1 with the upper bound per device
Timer interval ≤ 0Rejected and logged; no timer armed
fsm:fire of an event invalid in the current stateNo-op, returns false
Event fired during a transitionQueued, processed in order after it completes — no re-entrancy
FSM resume onto a state the redeployed code no longer definesFalls back to initial (firing its enter), warns, sets last_error

Time and persistence ​

CaseBehaviour
state.json missing or corrupt at bootStart with empty state, log it, do not crash
state.json write failsLog it; keep state in memory. StateDirectory is the only writable path under ProtectSystem=strict
System clock jumpTimers are monotonic and unaffected. Snoozes are wall-clock and can shorten or lengthen
Stored state for an instance no longer in deploy.jsonDropped during boot reconciliation; its retained status topic is cleared too
SIGTERMstate.json is persisted first, then onStop runs for each running instance within a 10 s aggregate budget. If the budget is exceeded, remaining hooks are skipped and the daemon exits cleanly

Lifecycle and popups ​

CaseBehaviour
onStart throws or times outPartial registration rolled back; instance errored
onStop throws or times outLogged; wiring still torn down and the interpreter still destroyed
Instance disabled or removed with a popup pendingPopup cleared; onAnswer is not called
Feedback with a non-matching or expired requestId, or for an unknown instanceIgnored
Two instances of one template both promptIndependent — each has its own state, config and single popup slot

Integration of multiplexed solutions
MUXEN and the MUXEN logo are trademarks of MUXEN SAS.