Skip to content

Notifications are fire-and-forget: add durable store-and-forward for one-shot events #85

Description

@pbitzer

Updated 2026-08-11. The stopgap this issue described is gone. HAM-182 shipped as a
level-triggered alert instead of a spooled one-shot event, which removes its need for
store-and-forward entirely. The general problem below still stands — it just no longer
has a live consumer, so it is lower priority than when filed.

Problem

Notifications are fire-and-forget. StateMonitor.send_message() calls self.notifier.send(msg) and nothing records whether it landed. If the link is down at that moment, the message is gone.

That is tolerable for a recurring or level-triggered check — the next poll re-raises it. It is not tolerable for a genuine one-shot event, where the notification is the only signal that something happened.

What the failure signal actually is

Verified against the sender implementations, not assumed:

  • Network failures raise. urllib.error.URLError and HTTPError are OSError subclasses, so DNS failure, refused connections and 4xx/5xx all propagate out of sender.send(). A caller that wants to know can find out.
  • Notifier.send() is a silent no-op when no sender is configured. notifiers/notify.py leaves self.sender = None on method=None, an unknown method, or a missing key file, and send() then returns normally without doing anything. Any caller inferring delivery from "did not raise" is wrong in this case — and a missing .googlechat key file is a real, observed condition on this fleet.
  • A non-200-but-non-error response is logged, not raised. Both notifiers/slack.py and notifiers/google_chat.py treat resp.status != 200 as a warning when a logger is supplied, and StateMonitor.__init__ always supplies one. urlopen auto-raises for 4xx/5xx, so this only bites on unusual 2xx/3xx responses.
  • No timeout anywhere. Neither sender passes timeout= to urlopen(), and sends happen synchronously inside the single-threaded telemetry loop (run_periodic). On a blackholed route this can stall the whole pipeline — ping, power, battery, CSV logging — for as long as the OS-level TCP/DNS retry takes.

Why the original stopgap is gone

HAM-182 originally wrote a per-boot JSON event to /var/lib/hamma/ and had a state_monitor sweep deliver it once connectivity returned. That whole mechanism was deleted along with the remediation it recorded.

The replacement is a level-triggered check: it re-compares the relay against brokkr's mode every hour and re-alerts for as long as they disagree. An alert raised with no link simply lands on the next cycle after the link returns. Repetition does the job store-and-forward was there for, with no spool, no ownership problem and no replay logic.

That pattern is worth reaching for first — it only works when the condition is still observable later, but when it is, it is much cheaper than durable queuing.

What to look into

Only if a genuinely one-shot event needs reliable delivery. Nothing in the tree needs it today.

  • Add an explicit timeout= to both senders' urlopen() calls. Smallest change here, widest benefit, and independent of everything else — it protects the telemetry loop regardless of what is decided about queuing.
  • Have send() report status, or raise on any non-200, so "did not raise" becomes a truthful delivery signal.
  • Make Notifier.send()'s no-sender case distinguishable from a successful send — right now they are identical from the caller's side.
  • If a durable queue is ever built: retention/ageing, what to do with events that can never be delivered, ordering and de-duplication behind an outage, and whether it belongs in notifiers (so brokkr and sindri both benefit) or in mjolnir-hamma.
  • Note journald Storage=volatile is set on mj03/41/42/43, so the journal is not a durable fallback on those units.

Context

  • HAM-182 — the check that originally motivated this; now level-triggered and no longer dependent on it
  • HAM-189 — field-log update path; the most likely future consumer of real one-shot delivery
  • Per-array notification channels already exist (e.g. panama_status for PAMMA), so any queue must preserve channel routing

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions