fix(supervisor): retry appservice login at startup — a homeserver-unreachable boot race permanently kills every bridge until manual restart #126

Closed
opened 2026-07-26 01:41:52 +00:00 by robocub · 2 comments
Member

Failure mode

If the host running the bridge reboots and the bridge process comes up before the homeserver is reachable (or the homeserver's front proxy briefly returns 5xx), every bridge's startup m.login.application_service fails:

ERROR bridge{name=<bridge>}: appservice login failed — ... status=502 Bad Gateway
ERROR bridge{name=<bridge>}: supervisor: bridge failed error=appservice login failed: 502
INFO  nvb_discord::puppet: channel unregistered from puppet pool bridge="<bridge>"

Every configured bridge plus the management bot dies within ~1s of startup and never retries — the outage lasts until an operator manually restarts the service. (Observed live 2026-07-22 after a host reboot.)

The failure mode is nasty because it's masked: the Discord puppet bots reconnect fine, so systemctl is-active reports active and the process looks healthy while zero bridging is happening in either direction.

Root cause

The supervisor treats a failed startup appservice login as a terminal bridge error. That's the wrong call for a network-shaped error: a 502/timeout at boot means "homeserver not up yet", not "config is wrong". The same applies to the management bot, which exits with management-bot appservice login.

Proposed fix: login retry/backoff in the supervisor

On appservice login failure with a retryable status (5xx, connect error, timeout), retry with capped exponential backoff (e.g. 5s → 10s → 20s → ... cap 60s, indefinitely or for a generous window), keeping the bridge in a "starting" state instead of failing it. Non-retryable statuses (401/403 = bad token, M_FORBIDDEN, 404 wrong URL) should still fail fast and loudly — those are config errors.

Same treatment for the management bot's login.

Why not After=/Wants= systemd ordering instead?

  • The homeserver generally doesn't run on the same host as the bridge, so there may be no local unit to order against.
  • A proxy-fronted homeserver can intermittently 502 even in steady state. A bridge that dies permanently on one 502 at startup is fragile regardless of boot ordering; the retry fixes the whole class.
  • A Restart=on-failure band-aid also doesn't fit: the process as a whole stays up (puppets alive), so systemd never sees a failure.

Acceptance

  • Start the bridge with the homeserver unreachable, restore reachability ~2 min later → all bridges + management bot come up on their own, no manual restart.
  • Startup with a wrong as_token still fails fast with a clear error (no infinite silent retry).
## Failure mode If the host running the bridge reboots and the bridge process comes up before the homeserver is reachable (or the homeserver's front proxy briefly returns 5xx), every bridge's startup `m.login.application_service` fails: ``` ERROR bridge{name=<bridge>}: appservice login failed — ... status=502 Bad Gateway ERROR bridge{name=<bridge>}: supervisor: bridge failed error=appservice login failed: 502 INFO nvb_discord::puppet: channel unregistered from puppet pool bridge="<bridge>" ``` Every configured bridge plus the management bot dies within ~1s of startup and **never retries** — the outage lasts until an operator manually restarts the service. (Observed live 2026-07-22 after a host reboot.) The failure mode is nasty because it's **masked**: the Discord puppet bots reconnect fine, so `systemctl is-active` reports `active` and the process looks healthy while zero bridging is happening in either direction. ## Root cause The supervisor treats a failed startup appservice login as a **terminal** bridge error. That's the wrong call for a *network-shaped* error: a 502/timeout at boot means "homeserver not up yet", not "config is wrong". The same applies to the management bot, which exits with `management-bot appservice login`. ## Proposed fix: login retry/backoff in the supervisor On appservice login failure with a **retryable** status (5xx, connect error, timeout), retry with capped exponential backoff (e.g. 5s → 10s → 20s → ... cap 60s, indefinitely or for a generous window), keeping the bridge in a "starting" state instead of failing it. Non-retryable statuses (401/403 = bad token, `M_FORBIDDEN`, 404 wrong URL) should still fail fast and loudly — those *are* config errors. Same treatment for the management bot's login. ### Why not `After=`/`Wants=` systemd ordering instead? - The homeserver generally doesn't run on the same host as the bridge, so there may be no local unit to order against. - A proxy-fronted homeserver can intermittently 502 even in steady state. A bridge that dies permanently on one 502 at startup is fragile regardless of boot ordering; the retry fixes the whole class. - A `Restart=on-failure` band-aid also doesn't fit: the process as a whole stays up (puppets alive), so systemd never sees a failure. ## Acceptance - Start the bridge with the homeserver unreachable, restore reachability ~2 min later → all bridges + management bot come up on their own, no manual restart. - Startup with a wrong `as_token` still fails fast with a clear error (no infinite silent retry).
Author
Member

Fix on branch fix/126-mgmt-login-retry (bd0ec42), awaiting review/merge.

Scope note: the bridge half of this issue was already closed by #130's supervisor restart backoff (a failed startup login → BridgeOutcome::Failed → backoff restart → self-heals). The remaining single-shot was the management bot (main.rs spawned it once; "bot that fails stays down"). The fix retries inside mgmt::run:

  • appservice_login_as now returns the typed AppserviceHttpError, so retryability is decided structurally (no text matching — same discipline as #130's auth_error)
  • login: 5xx/408/429/transport/timeout → capped backoff 5s→10s→20s→40s→60s, indefinitely, WARN per attempt; 400/401/403/404 → fail fast and loud (config error)
  • initial sync: same tolerance; structurally-fatal auth stays terminal; both loops honor shutdown
  • unit tests: login_retryability_splits_on_status_class, login_error_terminal_only_with_typed_status, startup_backoff_doubles_to_a_60s_cap

Acceptance mapping: HS unreachable at boot → mgmt bot self-heals when it returns (bridges already did via #130); wrong as_token → immediate loud failure, no silent retry. The homeserver-down acceptance leg still deserves a live staging check before release.

Fix on branch `fix/126-mgmt-login-retry` (`bd0ec42`), awaiting review/merge. Scope note: the **bridge** half of this issue was already closed by #130's supervisor restart backoff (a failed startup login → BridgeOutcome::Failed → backoff restart → self-heals). The remaining single-shot was the **management bot** (`main.rs` spawned it once; "bot that fails stays down"). The fix retries inside `mgmt::run`: - `appservice_login_as` now returns the typed `AppserviceHttpError`, so retryability is decided structurally (no text matching — same discipline as #130's `auth_error`) - login: 5xx/408/429/transport/timeout → capped backoff 5s→10s→20s→40s→60s, indefinitely, WARN per attempt; 400/401/403/404 → fail fast and loud (config error) - initial sync: same tolerance; structurally-fatal auth stays terminal; both loops honor shutdown - unit tests: `login_retryability_splits_on_status_class`, `login_error_terminal_only_with_typed_status`, `startup_backoff_doubles_to_a_60s_cap` Acceptance mapping: HS unreachable at boot → mgmt bot self-heals when it returns (bridges already did via #130); wrong `as_token` → immediate loud failure, no silent retry. The homeserver-down acceptance leg still deserves a live staging check before release.
Author
Member

Merged to master (32ba33a) and released in v0.3.3. The management bot now retries network-shaped login/initial-sync failures with capped backoff (5s→60s, indefinitely) and fails fast on config errors; bridges were already covered by #130's restart backoff. Closing.

Merged to master (32ba33a) and released in **v0.3.3**. The management bot now retries network-shaped login/initial-sync failures with capped backoff (5s→60s, indefinitely) and fails fast on config errors; bridges were already covered by #130's restart backoff. Closing.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
nether/nether-voicebridge#126
No description provided.