LxBox|科学之家 kexuezj.com
LxBox · 科学之家
LeadaxeLxBox / sing-box-lx 开发者
GitHub·2026.09.24 06:26:27(UTC+8)

Leadaxe:XHTTP 连接池断路器误判主动中断,健康连接被逐出

在 LxBox #148 中,Leadaxe 分析 connection download closed: http2: response body closed 常为中继关闭哨兵被误判为 ERROR;实节点成功请求也可能刷日志,WS 更干净因走另一关闭路径。他们将修分类,并讨论断路器把主动中断当失败的问题。详见原文。

作者原文@Leadaxe

Thanks for the detailed report — the transport matrix and the urltest/selector layout were genuinely useful. Some findings from reading the code and running your setup (VLESS+XHTTP+REALITY under urltest inside a selector, interrupt_exist_connections: true) against a live node.

1. connection download closed: http2: response body closed is not a failure. This line comes from the generic connection relay, not from the XHTTP transport. It is logged at ERROR only because the error-classification helper does not recognise the HTTP/2 response body closed sentinel as a normal close. On a live XHTTP node here, every successful request (HTTP 204, six out of six) produced exactly one such ERROR line, with zero connection evictions. WS closes its body through a different path that is recognised, which is why it looks clean by comparison. So a large part of the "XHTTP is broken, WS is fine" impression is a logging artifact. We will fix the classification.

2. lx_idle_suspend does not apply to XHTTP. That mechanism only acts on WireGuard/AWG endpoints; an XHTTP outbound is never touched by the idle tick. Setting lx_idle_suspend: 30s has no effect on your XHTTP connections, so this can be ruled out.

3. "Resumed from background — re-syncing tunnel (heartbeat/streams were paused)" does not touch the tunnel. The "heartbeat/streams" in that message are the app's own UI statistics streams, not HTTP/2 streams. On resume the app re-reads the VPN status, restarts a 20-second UI timer and re-subscribes its telemetry clients. It does not reinstall the TUN, does not reload the config and does not close any connection. The wording is misleading and we will look at it. The correlation you saw is most likely you returning to the app and generating traffic again.

4. There is a real defect, and it is in the connection pool. An XHTTP connection torn down by us (which is exactly what interrupt_exist_connections: true does on every urltest/selector switch) cancels its request context. An upload request still in flight then returns a cancellation error, and the pool's circuit breaker currently counts that as a stream failure. Three of those in a row mark a perfectly healthy pooled connection as failing and evict it, plus arm a backoff before a new one may be opened. We confirmed this in isolation. Note the breaker itself is ours (upstream Xray has no equivalent) — so this is our bug, not inherited behaviour.

Your layout makes this much easier to hit than a normal one: six XHTTP variants of the same server sitting under one urltest have near-identical latency, so the auto-selection can flap, and every flap fires a batch of interrupts.

Things worth trying on your side, to confirm:

  • set interrupt_exist_connections: false on both the selector and the urltest group;
  • keep a single XHTTP variant in the urltest group instead of six, or raise tolerance (say 200–300 ms);
  • test one XHTTP node directly, with no selector and no urltest.

What would help us most:

  • the full dump, specifically every xhttp: xmux: line — opened connection, evicted connection (cause=...) and any breaker tripped. The cause= value is what separates our bug from a genuine reset by the server or the network;
  • the XHTTP node config without secrets: mode, downloadSettings, the xmux section, and the urltest tolerance/interval;
  • whether you are actually losing traffic, or mainly seeing these ERROR lines in the log — that tells us which of the two problems above you are hitting;
  • whether it reproduces on core lx.8 (you are on lx.4).

Note that nothing in the XHTTP/xmux code changed between lx.4 and lx.8, so upgrading alone is not expected to fix this — it just rules out other differences.