rtc: ICE gathers host candidates on tailscale0 too — stuck SYN-SENT to LiveKit :7881; add a client-side interface exclusion (rtc.interfaces twin of #91) #138

Open
opened 2026-09-28 13:54:13 +00:00 by robocub · 0 comments
Member

Development note (filed 2026-09-28 at the operator's request). Not urgent; cosmetic today, but it is the client-side twin of #91.

Observation (prod, mautrix, v0.3.7)

During and shortly after calls the bridge process shows sockets stuck in SYN-SENT toward LiveKit's ICE-TCP port on nether-red (97.107.138.76:7881) — six at once on 2026-09-28 03:36, unchanged across ~10 min, gone once the calls ended. The port itself is reachable from mautrix (a plain /dev/tcp connect succeeds), so the stuck SYNs are not a firewall problem on the server.

The bridge's ICE-TCP listeners explain it. ss -tlnp on the process shows passive ICE-TCP sockets bound to every non-loopback address on the box:

eth0        10.34.154.179     (incus bridge, NATed to the public IP — the only one that can reach LiveKit)
tailscale0  100.64.0.33       (tailnet v4)
tailscale0  fd7a:115c:a1e0::21 (tailnet v6)

libwebrtc gathers host candidates on all of them and pairs each with the SFU's public candidates. Pairs sourced from the tailscale addresses can never complete (no route from that interface to a public IP), so their TCP connectivity checks sit in SYN-SENT until they time out. Media is unaffected: the eth0 UDP pair wins. Costs are noise in the socket table, wasted connectivity checks and marginally slower ICE, and one more thing for anyone reading ss to explain.

Server side we already solved the mirror problem in #91 (rtc.interfaces.includes: [eth0] on LiveKit, which had been advertising tailscale/incus candidates → ~2/3 zero-media). This is the same thing from the client's chair.

What the SDK exposes today

livekit 0.7.45 / libwebrtc 0.3.36: RoomOptions.rtc_config carries only ice_servers, continual_gathering_policy and ice_transport_type (All / Relay / NoHost). There is no network_ignore_mask, interface allow-list or adapter-type filter. NoHost is not an option for us — we have no TURN and rely on the host candidate on eth0.

Options

  1. Upstream (preferred long-term): ask/PR livekit/rust-sdks (webrtc-sys) to surface libwebrtc's PortAllocator::set_network_ignore_mask() (or an RTCConfiguration-level interface allow-list) through RtcConfiguration; then add livekit.ice_interfaces = ["eth0"] / ice_ignore_interfaces to our [bridge.livekit] config and thread it into RoomOptions. Pin bump implications: constraint #7 (livekit is pinned) — do it as a deliberate bump.
  2. Deployment workaround: run the bridge in a network namespace / container that does not see tailscale0 (the systemd unit could use RestrictNetworkInterfaces=eth0 lo — worth a try first: it's one line in a drop-in and needs no code). Caveat: anything the bridge legitimately does over the tailnet would break; today it should not need the tailnet at all (homeserver, Discord, LiveKit are all public).
  3. Accept: keep as-is. nvb-watch counts ESTABLISHED only, so these never trip the connection envelope (#135-adjacent alert of 2026-09-28 was real sockets from concurrent calls, envelope raised to 90 in 7e9fe17).

Verification recipe

pid=$(systemctl show nether-voicebridge -p MainPID --value)
ss -tanp state syn-sent | grep "pid=$pid,"     # during a call: should be empty after the fix
ss -tlnp | grep "pid=$pid," | awk '{print $4}' # listeners: only eth0 addresses after the fix

Cheapest first experiment: option 2 (RestrictNetworkInterfaces=) on staging, confirm audio + empty SYN-SENT, then decide whether option 1 is worth pursuing.

Development note (filed 2026-09-28 at the operator's request). Not urgent; cosmetic today, but it is the client-side twin of #91. ## Observation (prod, mautrix, v0.3.7) During and shortly after calls the bridge process shows sockets stuck in `SYN-SENT` toward LiveKit's ICE-TCP port on nether-red (`97.107.138.76:7881`) — six at once on 2026-09-28 03:36, unchanged across ~10 min, gone once the calls ended. The port itself is reachable from mautrix (a plain `/dev/tcp` connect succeeds), so the stuck SYNs are not a firewall problem on the server. The bridge's ICE-TCP *listeners* explain it. `ss -tlnp` on the process shows passive ICE-TCP sockets bound to every non-loopback address on the box: ``` eth0 10.34.154.179 (incus bridge, NATed to the public IP — the only one that can reach LiveKit) tailscale0 100.64.0.33 (tailnet v4) tailscale0 fd7a:115c:a1e0::21 (tailnet v6) ``` libwebrtc gathers host candidates on all of them and pairs each with the SFU's public candidates. Pairs sourced from the tailscale addresses can never complete (no route from that interface to a public IP), so their TCP connectivity checks sit in SYN-SENT until they time out. Media is unaffected: the eth0 UDP pair wins. Costs are noise in the socket table, wasted connectivity checks and marginally slower ICE, and one more thing for anyone reading `ss` to explain. Server side we already solved the mirror problem in #91 (`rtc.interfaces.includes: [eth0]` on LiveKit, which had been advertising tailscale/incus candidates → ~2/3 zero-media). This is the same thing from the client's chair. ## What the SDK exposes today `livekit` 0.7.45 / `libwebrtc` 0.3.36: `RoomOptions.rtc_config` carries only `ice_servers`, `continual_gathering_policy` and `ice_transport_type` (`All` / `Relay` / `NoHost`). There is **no** `network_ignore_mask`, interface allow-list or adapter-type filter. `NoHost` is not an option for us — we have no TURN and rely on the host candidate on eth0. ## Options 1. **Upstream (preferred long-term):** ask/PR livekit/rust-sdks (`webrtc-sys`) to surface libwebrtc's `PortAllocator::set_network_ignore_mask()` (or an `RTCConfiguration`-level interface allow-list) through `RtcConfiguration`; then add `livekit.ice_interfaces = ["eth0"]` / `ice_ignore_interfaces` to our `[bridge.livekit]` config and thread it into `RoomOptions`. Pin bump implications: constraint #7 (livekit is pinned) — do it as a deliberate bump. 2. **Deployment workaround:** run the bridge in a network namespace / container that does not see `tailscale0` (the systemd unit could use `RestrictNetworkInterfaces=eth0 lo` — worth a try first: it's one line in a drop-in and needs no code). Caveat: anything the bridge legitimately does over the tailnet would break; today it should not need the tailnet at all (homeserver, Discord, LiveKit are all public). 3. **Accept:** keep as-is. nvb-watch counts `ESTABLISHED` only, so these never trip the connection envelope (#135-adjacent alert of 2026-09-28 was real sockets from concurrent calls, envelope raised to 90 in `7e9fe17`). ## Verification recipe ``` pid=$(systemctl show nether-voicebridge -p MainPID --value) ss -tanp state syn-sent | grep "pid=$pid," # during a call: should be empty after the fix ss -tlnp | grep "pid=$pid," | awk '{print $4}' # listeners: only eth0 addresses after the fix ``` Cheapest first experiment: option 2 (`RestrictNetworkInterfaces=`) on staging, confirm audio + empty SYN-SENT, then decide whether option 1 is worth pursuing.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
nether/nether-voicebridge#138
No description provided.