Updated 2026-08-11. The stopgap this issue described is gone. HAM-182 shipped as a
level-triggered alert instead of a spooled one-shot event, which removes its need for
store-and-forward entirely. The general problem below still stands — it just no longer
has a live consumer, so it is lower priority than when filed.
Problem
Notifications are fire-and-forget. StateMonitor.send_message() calls self.notifier.send(msg) and nothing records whether it landed. If the link is down at that moment, the message is gone.
That is tolerable for a recurring or level-triggered check — the next poll re-raises it. It is not tolerable for a genuine one-shot event, where the notification is the only signal that something happened.
What the failure signal actually is
Verified against the sender implementations, not assumed:
- Network failures raise.
urllib.error.URLError and HTTPError are OSError subclasses, so DNS failure, refused connections and 4xx/5xx all propagate out of sender.send(). A caller that wants to know can find out.
Notifier.send() is a silent no-op when no sender is configured. notifiers/notify.py leaves self.sender = None on method=None, an unknown method, or a missing key file, and send() then returns normally without doing anything. Any caller inferring delivery from "did not raise" is wrong in this case — and a missing .googlechat key file is a real, observed condition on this fleet.
- A non-200-but-non-error response is logged, not raised. Both
notifiers/slack.py and notifiers/google_chat.py treat resp.status != 200 as a warning when a logger is supplied, and StateMonitor.__init__ always supplies one. urlopen auto-raises for 4xx/5xx, so this only bites on unusual 2xx/3xx responses.
- No timeout anywhere. Neither sender passes
timeout= to urlopen(), and sends happen synchronously inside the single-threaded telemetry loop (run_periodic). On a blackholed route this can stall the whole pipeline — ping, power, battery, CSV logging — for as long as the OS-level TCP/DNS retry takes.
Why the original stopgap is gone
HAM-182 originally wrote a per-boot JSON event to /var/lib/hamma/ and had a state_monitor sweep deliver it once connectivity returned. That whole mechanism was deleted along with the remediation it recorded.
The replacement is a level-triggered check: it re-compares the relay against brokkr's mode every hour and re-alerts for as long as they disagree. An alert raised with no link simply lands on the next cycle after the link returns. Repetition does the job store-and-forward was there for, with no spool, no ownership problem and no replay logic.
That pattern is worth reaching for first — it only works when the condition is still observable later, but when it is, it is much cheaper than durable queuing.
What to look into
Only if a genuinely one-shot event needs reliable delivery. Nothing in the tree needs it today.
- Add an explicit
timeout= to both senders' urlopen() calls. Smallest change here, widest benefit, and independent of everything else — it protects the telemetry loop regardless of what is decided about queuing.
- Have
send() report status, or raise on any non-200, so "did not raise" becomes a truthful delivery signal.
- Make
Notifier.send()'s no-sender case distinguishable from a successful send — right now they are identical from the caller's side.
- If a durable queue is ever built: retention/ageing, what to do with events that can never be delivered, ordering and de-duplication behind an outage, and whether it belongs in
notifiers (so brokkr and sindri both benefit) or in mjolnir-hamma.
- Note
journald Storage=volatile is set on mj03/41/42/43, so the journal is not a durable fallback on those units.
Context
- HAM-182 — the check that originally motivated this; now level-triggered and no longer dependent on it
- HAM-189 — field-log update path; the most likely future consumer of real one-shot delivery
- Per-array notification channels already exist (e.g.
panama_status for PAMMA), so any queue must preserve channel routing
Problem
Notifications are fire-and-forget.
StateMonitor.send_message()callsself.notifier.send(msg)and nothing records whether it landed. If the link is down at that moment, the message is gone.That is tolerable for a recurring or level-triggered check — the next poll re-raises it. It is not tolerable for a genuine one-shot event, where the notification is the only signal that something happened.
What the failure signal actually is
Verified against the sender implementations, not assumed:
urllib.error.URLErrorandHTTPErrorareOSErrorsubclasses, so DNS failure, refused connections and 4xx/5xx all propagate out ofsender.send(). A caller that wants to know can find out.Notifier.send()is a silent no-op when no sender is configured.notifiers/notify.pyleavesself.sender = Noneonmethod=None, an unknown method, or a missing key file, andsend()then returns normally without doing anything. Any caller inferring delivery from "did not raise" is wrong in this case — and a missing.googlechatkey file is a real, observed condition on this fleet.notifiers/slack.pyandnotifiers/google_chat.pytreatresp.status != 200as a warning when a logger is supplied, andStateMonitor.__init__always supplies one.urlopenauto-raises for 4xx/5xx, so this only bites on unusual 2xx/3xx responses.timeout=tourlopen(), and sends happen synchronously inside the single-threaded telemetry loop (run_periodic). On a blackholed route this can stall the whole pipeline — ping, power, battery, CSV logging — for as long as the OS-level TCP/DNS retry takes.Why the original stopgap is gone
HAM-182 originally wrote a per-boot JSON event to
/var/lib/hamma/and had astate_monitorsweep deliver it once connectivity returned. That whole mechanism was deleted along with the remediation it recorded.The replacement is a level-triggered check: it re-compares the relay against brokkr's mode every hour and re-alerts for as long as they disagree. An alert raised with no link simply lands on the next cycle after the link returns. Repetition does the job store-and-forward was there for, with no spool, no ownership problem and no replay logic.
That pattern is worth reaching for first — it only works when the condition is still observable later, but when it is, it is much cheaper than durable queuing.
What to look into
Only if a genuinely one-shot event needs reliable delivery. Nothing in the tree needs it today.
timeout=to both senders'urlopen()calls. Smallest change here, widest benefit, and independent of everything else — it protects the telemetry loop regardless of what is decided about queuing.send()report status, or raise on any non-200, so "did not raise" becomes a truthful delivery signal.Notifier.send()'s no-sender case distinguishable from a successful send — right now they are identical from the caller's side.notifiers(so brokkr and sindri both benefit) or inmjolnir-hamma.journaldStorage=volatileis set on mj03/41/42/43, so the journal is not a durable fallback on those units.Context
panama_statusfor PAMMA), so any queue must preserve channel routing