Both Clyde North observers stopped answering over WiFi after days of
uptime while LoRa and MQTT kept working. `memory` showed heap_min ~1KB
and a largest internal block of 19KB: the HTTPS listener was alive but
no mbedTLS handshake (~40KB contiguous) could be allocated.
The trigger was the Aug 9 lru_purge change: with a 2-socket pool, every
browser page-load burst evicted a live session and forced a fresh TLS
handshake, and the repeated 40KB alloc/free cycles fragmented internal
RAM. Before that change the pool simply jammed, so nothing churned.
- Send `Connection: close` and trigger a session close on one-shot
responses (pages, favicon, redirect, login, 401s) so page-load bursts
release their sockets immediately. The authenticated /api/* polling
connection keeps keep-alive so it does not pay a handshake per poll.
- Gate WebPanelServer::start() on internal heap headroom (56KB free /
32KB largest), retrying every 15s instead of every loop tick.
- Self-heal: when the panel is idle and the largest internal block drops
below 24KB, stop and re-create the server to return its pools.
- Count the MQTT teardown path that deliberately abandons a client on a
heap-integrity failure.
- Expose the above as `heals:`/`deferred:` in `get web.status` and
`leaked:` in `get mqtt.status`.
- Fix the stats page Channel/Gateway cells showing `--`: the /api/stats
summary JSON never carried channel, gateway health, or the watchdog
count (only the `get wifi.status` string did).
- Add eastmesh-tools/web-heap-check.sh to read heap/service health over
the API and optionally stress the panel with browser-style bursts.
- Release notes 2026.8.3 for both observer tracks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Both Clyde North observers sat WL_CONNECTED on an AP that kept beaconing
after its bridge to the wired LAN died: association up, good RSSI, stale
DHCP lease, but ARP-invisible and both MQTT brokers stuck in backoff.
The firmware equated "associated" with "online", so neither node ever
re-scanned and the outage held until a manual reassociation.
Add a connectivity watchdog to NetworkService: while connected, ARP-probe
the gateway every 30s (posted to the lwIP tcpip thread, safe on both the
IDF4 and IDF5 cores). If the gateway stays silent for 3 minutes, clear
the channel hint and force a full disconnect/rescan so the node can roam
to a healthy AP, backing off exponentially (up to 48min) while the
outage persists. Nodes without a gateway skip the probe.
Add a `wifi reconnect` repeater command that triggers the same forced
reassociation on demand, and report `gw:ok|lost wd:<count>` in
`get wifi.status`. Teach the web panel parser the new fields (also fixes
the IP metric rendering as "x.x.x.x channel:n") and show Gateway and
Channel in the Wi-Fi card.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
With saved credentials present (seeded from the WIFI_SSID build flag on
first boot), the recovery AP came up in AP+STA mode while auto-reconnect
and the 10s manual retry kept the station scanning. STA scans drag the
shared radio across channels, so joining clients' WPA2 handshakes timed
out - reported by phones as a wrong password.
Disable auto-reconnect and pause the 10s retry loop while the recovery
AP is active; retry the configured network once a minute instead, so the
AP stays stable between attempts and the node still self-recovers (and
shuts the AP down) when its network returns.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The upstream merge switched SH1106Display::begin() to probe-then-init.
On the T-Beam S3 Supreme the OLED bus can NAK a cold address probe, so
begin() bailed before display.begin() ever ran, and UITask::begin()'s
unconditional turnOn() then drove Adafruit_GrayOLED with a null bus
device: bogus gpio_set_level(227) followed by a LoadProhibited boot loop.
Restore the proven S3 Supreme bring-up order (Adafruit init with reset
sequence first, then settle the bus and verify), and track begin()
success so turnOn/turnOff/clear are safe no-ops when the display never
initialised - a failed or absent display now boots headless instead of
crash-looping.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
With max_open_sockets=2 and no LRU purge, clients that vanish without
closing (lid shut, out of range) permanently occupy the pool on both the
HTTPS panel and the port-80 redirect listener; after enough uptime every
new connect is reset (Firefox PR_CONNECT_RESET_ERROR) until reboot.
Enable lru_purge_enable on both servers so a new connection evicts the
least-recently-used one instead of being refused.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Mirror the BLE pin rules on WiFi builds so the recovery AP always has a
knowable password: devices with a display get a random per-session pin
shown as Pin:NNNNNN on the home screen, headless devices default to
123456, and 'set pin' overrides both persistently. The AP password is
the active pin zero-padded to 8 digits (WPA2 minimum), so a fresh
device is never an open AP.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
When a companion WiFi device has no WiFi credentials or can't connect for
60s, it broadcasts an 'EastMesh-WiFi' AP (WPA2 password = device pin
zero-padded to 8 digits; open if no pin set) with the rescue CLI on TCP
port 23, so wifi.ssid/pwd can be fixed without a serial cable.
- rescue CLI command handler now writes to any Stream (Serial or the TCP
rescue client)
- while the AP is up, STA keeps retrying; the AP shuts down automatically
once the configured network connects
- 'set wifi.ssid/pwd' keeps AP+STA mode during a rescue session so the
session isn't dropped mid-config
- guard ESP32-only SD/SPI code in ArchiveStorage so non-ESP32 builds of
the shared helpers get further
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Integrates upstream v1.17.0. Key adaptations:
- companion + repeater NodePrefs moved to upstream's new ConfigSerializer
(prefs.json) format; EastMesh fields (companion wifi, repeater fan and
MQTT bridge peer) appended as serializer groups/keys
- legacy binary prefs loaders retain EastMesh tail fields, so existing
devices migrate their settings to the new format
- companion main.cpp adopts upstream InterfaceManager while keeping
runtime WiFi credentials from prefs
- kept EastMesh release-tag versioning in setup-build-environment action
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
mqtt2.eastmesh.au switched from a Let's Encrypt cert to a Google Trust
Services WE1-issued *.eastmesh.au wildcard (July 2026), so devices on the
pinned-PEM fallback path failed the handshake with mbedTLS -9984. Pin the
combined ISRG Root X1 + GTS WE1 bundle for these brokers so either chain
verifies.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
loadPrefsInt guarded each EastMesh field with its own available() check
and no bail-out. On a short prefs file that skipped a large field but
still satisfied a later, smaller one: with 8 bytes left, bridge_peer_host
(64) was correctly skipped, then bridge_peer_port (2) passed its guard
and read two misaligned bytes. It has no constrain(), so it loaded 0
instead of its 1883 default.
Read sequentially and stop at the first field that isn't fully present,
leaving the rest at their constructor defaults.
This does not make a differently-laid-out prefs file safe to read, only a
shorter one -- fan_mode and fan_timeout_secs still take whatever bytes sit
at 512..514. A format marker on the EastMesh block would close that.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Upstream's persisted fields end at 295. EastMesh's own fields sat
immediately after, so any future upstream append would shift them and
conflict. Base the EastMesh block at 512 instead, with the 295..511 gap
reserved for upstream growth.
Upstream can now append fields with no change to EastMesh offsets; a
static_assert fails the build if upstream ever overruns the gap.
EastMesh fields now occupy 512..741 (was 295..524).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adopts upstream's 3.13 bump. Also renames the uv cache key from py311
to py313 so a 3.11-built .venv is not restored under the 3.13
interpreter, and updates .python-version so local uv runs match CI.
pyproject requires-python stays >=3.11; 3.13 satisfies it and the
existing uv.lock resolves unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Keep upstream's persisted layout verbatim (rx_boosted_gain 290 ..
cad_enabled 294) and append the EastMesh-only fields after it, rather
than interleaving fan_mode/fan_timeout_secs into upstream's range.
Minimises CommonCLI.cpp conflicts on every upstream sync.
EastMesh fields now occupy 295..524: fan_mode, fan_timeout_secs, and
bridge_peer_*. Existing prefs files written by prior EastMesh firmware
will be misread from offset 292 onward and need reconfiguring.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
De-emphasize the observer-eastmesh-bridge-mqtt feature in public-facing
docs (the community prefers not to surface another MQTT-backed mesh):
- eastmesh-docs: drop the track from index/releases/boards track lists,
the custom-cli "MQTT Bridge Settings" section, and the web-panel peer
broker fields; keep only a vague mention in local-builds.md.
- release-notes.yml: remove the track from the list and drop the
2026.7.0 observer-eastmesh-bridge-mqtt entry.
- README: remove the "bidirectional MQTT mesh bridge" bullet and add the
missing repeater-bridge-espnow track to "What This Repo Adds".
The feature, build envs, and CI are unchanged; this is visibility only.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
MQTT status payload:
- Add packets_sent and packets_received (cumulative radio TX/RX totals)
to the status `stats` object, aligning with the Waev/MeshMapper schema.
Variants (variants/eastmesh_mqtt -> variants/eastmesh):
- Rename the variant folder to variants/eastmesh.
- Add the remaining *_repeater_observer_mqtt_bridge envs (2 -> 36), one
per observer board.
- Regroup envs by board (observer -> espnow -> mqtt_bridge) and add
per-vendor section banners.
Docs & release notes:
- Update folder path references in README, AGENTS, and boards docs.
- Add observer-eastmesh and observer-eastmesh-bridge-espnow 2026.6.6
release notes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Resolve leftover merge conflict markers in release-notes.yml from PR #73,
keeping both the observer-eastmesh-bridge-espnow 2026.6.5 and the new
observer-eastmesh-bridge-mqtt 2026.7.0 entries in track order
- Normalize the bridge-mqtt release entry `area:` values to the file's
conventions (bridge -> mqtt, boards -> board)
- Remove dead bridge_peer_port > 65535 check in loadPrefsInt (uint16_t can
never exceed 65535; the 0 -> 1883 default is handled at use)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add Waev (mqtt.waev.app) as a curated MQTT broker, selectable from the
web-panel Primary/Secondary dropdowns and the CLI.
- Broker spec on bit 0x20, audience mqtt.waev.app, Google Trust Services
WE1 CA (matches Waev's cert chain); widen broker mask/count to 0x3F/6
- Per-broker JWT lifetime via BrokerSpec.token_ttl_secs: Waev enforces a
1-hour max token TTL and rejects the 6h default with MQTT CONNACK 5,
so it mints a 3600s token while other brokers keep the 6h default
- CLI get/set mqtt.waev, web-panel dropdown entry + hidden toggle, docs
- Log token TTL in the "token ready" line for broker auth debugging
fix(web): make Custom MQTT settings full width when Custom is selected
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>