From cba357068dc586e7540971eb40bb56c5e62c4de1 Mon Sep 17 00:00:00 2001 From: Aris Konstantoulas Date: Mon, 28 Sep 2026 13:31:20 +0300 Subject: tcp: Don't fast re-transmit if only our FIN is outstanding In the TAP_FIN_RCVD path of tcp_tap_handler(), a bare segment from the guest acknowledging exactly seq_ack_from_tap, with an unchanged window, is taken as a duplicate ACK and triggers a fast re-transmit. If the only unacknowledged sequence number is our own FIN, that's harmful: tcp_rewind_seq() rewinds seq_to_tap and clears TAP_FIN_SENT, so tcp_data_from_sock() immediately sends the FIN again. If the guest answers that FIN with the same bare ACK, as a socket in TIME-WAIT will, we loop at packet rate: - conn->retries is never incremented on this path, so we never reach TCP_MAX_RETRIES and tcp_rst() - ACK_FROM_TAP_DUE is re-armed on every iteration, so the backed-off re-transmission in tcp_timer_handler() never fires - TAP_FIN_ACKED can't be set, as it requires TAP_FIN_SENT, which the rewind just cleared On an idle Podman host (rootless, pasta), this showed up as a single flow exchanging ~45,000 54-byte segments per second between pasta and a container whose socket was in TIME-WAIT, with pasta using ~75% of one core, until the socket was killed by hand. It recurred on the idle teardown of an HTTP/2 connection to an ACME server. With a raw-socket peer driving the same sequence against pasta at f8df3f1, pasta re-sent the FIN 727,509 times in 10 seconds. With this change it's re-transmitted by the timer at 1, 3, 7, 15, 31, 63 and 127 seconds, and the connection is reset once TCP_MAX_RETRIES is reached. Don't consider a duplicate ACK as a fast re-transmit trigger if the only outstanding sequence number is the FIN, and leave it to the timer. Fast re-transmit of data is unaffected, with or without a FIN queued after it. Fixes: bde1847960cf ("tcp: Fast re-transmit if half-closed, make TAP_FIN_RCVD path consistent") Link: https://bugs.passt.top/show_bug.cgi?id=125 Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Aris Konstantoulas Signed-off-by: Stefano Brivio --- tcp.c | 11 +++++++++-- 1 file changed, 9 insertions(+), 2 deletions(-) diff --git a/tcp.c b/tcp.c index 3b78d2e..e279f8a 100644 --- a/tcp.c +++ b/tcp.c @@ -2417,15 +2417,22 @@ int tcp_tap_handler(const struct ctx *c, uint8_t pif, sa_family_t af, /* Established connections not accepting data from tap */ if (conn->events & TAP_FIN_RCVD) { + bool fin_only, retr; size_t dlen; - bool retr; if ((dlen = tcp_packet_data_len(th, l4len))) { flow_dbg(conn, "data segment in CLOSE-WAIT (%zu B)", dlen); } - retr = th->ack && !th->fin && + /* If only our FIN is outstanding, rewinding on a duplicate ACK + * would re-send it on every ACK, bypassing the backoff and + * retry limit of tcp_timer_handler(): let the timer do it. + */ + fin_only = (conn->events & TAP_FIN_SENT) && + conn->seq_to_tap == conn->seq_ack_from_tap + 1; + + retr = th->ack && !th->fin && !fin_only && ntohl(th->ack_seq) == conn->seq_ack_from_tap && ntohs(th->window) == conn->wnd_from_tap; -- cgit v1.2.3