Skip to content

mjol_array --status reports a powered-off front end as Up (ping leaks to wwan0) #86

Description

@pbitzer

Summary

mjol_array.py --status reports "FCM is Up" for a front end that is definitively powered off. The check is a ping, and on a unit with eth0 down the ping is routed out over the cellular modem where something carrier-side answers it.

This fails in the dangerous direction: a downed sensor reads as healthy, and the condition that produces the false positive (eth0 down) is exactly the state a downed sensor is in.

Observed

mj03, 2026-08-09, immediately after mjol_array.py -p 3 --down completed successfully.

Ground truth — the front end is off:

raspi-gpio get 4   →  GPIO 4: level=1 fsel=1 func=OUTPUT pull=NONE   (de-energised, active_high=true)
Load Current       →  0.505 A   (was ~1.4 A)
/sys/class/net/eth0/operstate  →  down

But:

$ python3 /home/monitor/mjol_array.py -p 3 --status
Mjolnir03 is Up; FCM is Up; Brokkr Up; Sindri Up; ...

Root cause

status_fcm() (server/mjol_array.py:262-280) derives the result from a ping return code:

ping_code = next((x for x in retval if 'Ping Retcode' in x))
ping_code = int(ping_code.split(':')[-1])
...
return not ping_code

The AGS Pi is on the switched (relay) leg, so when the front end is powered down its Pi loses power and eth0 goes down with it. There is then no local route to 10.10.10.1, and the packet falls through to the default route:

$ ip route get 10.10.10.1
10.10.10.1 via 100.119.73.150 dev wwan0 src 100.119.73.149

Replies come back at 38–49 ms, versus the sub-millisecond a local link gives. The reply is not from the AGS.

Reproduced on mj03; expected on any unit where eth0 is down and the default route is via wwan0, which is the normal cellular configuration.

Why this matters beyond one status line

  • A sensor that was powered off, or that lost its front end unexpectedly, reads as "Up" in every fleet status sweep.
  • It undermines --status as the tool for confirming that an --up/--down actually took effect.
  • It is the same looks-healthy-but-isn't class as HAM-182, where a partial-capture state passed every check that was being run.

It also breaks a rule that has been relied on operationally: "a ping reply proves the sensor is on." It does not, when the route leaks to the WAN. Silence proves nothing and a reply now proves nothing either — the check needs to be something other than reachability.

Suggested fix

Options, roughly in order of preference:

  1. Bind the ping to the local interface — ping -I eth0 (or the configured sensor interface) so it cannot fall through to wwan0. Smallest change; turns the false positive into a correct negative.
  2. Check the link first — read /sys/class/net/eth0/operstate and report down without pinging at all. Cheap and unambiguous.
  3. Use load current instead — adc_il_f (telemetry field 7) is the actual power-state measurement, ~1.4 A on versus ~0.4 A off. Most truthful, but per-unit and time-varying, so it needs a per-unit baseline rather than a fixed threshold.

(1) and (2) together are probably right: link state as the primary signal, an interface-bound ping to confirm the AGS is actually answering.

There is a test file at tests/python/test_mjol_array.py that should get a case for "eth0 down, default route via wwan0 → reports down".

Related

  • HAM-182 — the cold-loss reconciler; same failure class
  • sensor-log entry for mj03's 2026-08-09 power-down, which carries a warning about this
  • docs/sensors-usage.md already warns that ping is not a valid indicator of front-end power; status_fcm() does exactly what the docs say not to do

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions