Both Clyde North observers stopped answering over WiFi after days of
uptime while LoRa and MQTT kept working. `memory` showed heap_min ~1KB
and a largest internal block of 19KB: the HTTPS listener was alive but
no mbedTLS handshake (~40KB contiguous) could be allocated.
The trigger was the Aug 9 lru_purge change: with a 2-socket pool, every
browser page-load burst evicted a live session and forced a fresh TLS
handshake, and the repeated 40KB alloc/free cycles fragmented internal
RAM. Before that change the pool simply jammed, so nothing churned.
- Send `Connection: close` and trigger a session close on one-shot
responses (pages, favicon, redirect, login, 401s) so page-load bursts
release their sockets immediately. The authenticated /api/* polling
connection keeps keep-alive so it does not pay a handshake per poll.
- Gate WebPanelServer::start() on internal heap headroom (56KB free /
32KB largest), retrying every 15s instead of every loop tick.
- Self-heal: when the panel is idle and the largest internal block drops
below 24KB, stop and re-create the server to return its pools.
- Count the MQTT teardown path that deliberately abandons a client on a
heap-integrity failure.
- Expose the above as `heals:`/`deferred:` in `get web.status` and
`leaked:` in `get mqtt.status`.
- Fix the stats page Channel/Gateway cells showing `--`: the /api/stats
summary JSON never carried channel, gateway health, or the watchdog
count (only the `get wifi.status` string did).
- Add eastmesh-tools/web-heap-check.sh to read heap/service health over
the API and optionally stress the panel with browser-style bursts.
- Release notes 2026.8.3 for both observer tracks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Both Clyde North observers sat WL_CONNECTED on an AP that kept beaconing
after its bridge to the wired LAN died: association up, good RSSI, stale
DHCP lease, but ARP-invisible and both MQTT brokers stuck in backoff.
The firmware equated "associated" with "online", so neither node ever
re-scanned and the outage held until a manual reassociation.
Add a connectivity watchdog to NetworkService: while connected, ARP-probe
the gateway every 30s (posted to the lwIP tcpip thread, safe on both the
IDF4 and IDF5 cores). If the gateway stays silent for 3 minutes, clear
the channel hint and force a full disconnect/rescan so the node can roam
to a healthy AP, backing off exponentially (up to 48min) while the
outage persists. Nodes without a gateway skip the probe.
Add a `wifi reconnect` repeater command that triggers the same forced
reassociation on demand, and report `gw:ok|lost wd:<count>` in
`get wifi.status`. Teach the web panel parser the new fields (also fixes
the IP metric rendering as "x.x.x.x channel:n") and show Gateway and
Channel in the Wi-Fi card.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
SPI pins that are set to -1 cause OOB reads on NRF52. The unused Serial2 pin definitions were removed to avoid potential issues with the Uart framework.
Setting SPI pins to values that are OOB of the g_ADigitalPinMap[] array causes OOB reads. This variant.h has pin 0 as 0xFF which passes through as NRFX_SPIM_PIN_NOT_USED.
Introduced consistent preferences for external LoRa FEM RX and TX gain settings in NodePrefs. Updated companion MyMesh to apply these settings during initialization and transmission. Added unit tests to verify the round-trip serialization of these new preferences.
With saved credentials present (seeded from the WIFI_SSID build flag on
first boot), the recovery AP came up in AP+STA mode while auto-reconnect
and the 10s manual retry kept the station scanning. STA scans drag the
shared radio across channels, so joining clients' WPA2 handshakes timed
out - reported by phones as a wrong password.
Disable auto-reconnect and pause the 10s retry loop while the recovery
AP is active; retry the configured network once a minute instead, so the
AP stays stable between attempts and the node still self-recovers (and
shuts the AP down) when its network returns.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The upstream merge switched SH1106Display::begin() to probe-then-init.
On the T-Beam S3 Supreme the OLED bus can NAK a cold address probe, so
begin() bailed before display.begin() ever ran, and UITask::begin()'s
unconditional turnOn() then drove Adafruit_GrayOLED with a null bus
device: bogus gpio_set_level(227) followed by a LoadProhibited boot loop.
Restore the proven S3 Supreme bring-up order (Adafruit init with reset
sequence first, then settle the bus and verify), and track begin()
success so turnOn/turnOff/clear are safe no-ops when the display never
initialised - a failed or absent display now boots headless instead of
crash-looping.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
With max_open_sockets=2 and no LRU purge, clients that vanish without
closing (lid shut, out of range) permanently occupy the pool on both the
HTTPS panel and the port-80 redirect listener; after enough uptime every
new connect is reset (Firefox PR_CONNECT_RESET_ERROR) until reboot.
Enable lru_purge_enable on both servers so a new connection evicts the
least-recently-used one instead of being refused.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>