diff options
| author | Aris Konstantoulas <aris@ariscodes.com> | 2026-09-28 13:31:20 +0300 |
|---|---|---|
| committer | Stefano Brivio <sbrivio@redhat.com> | 2026-10-02 23:05:26 +0200 |
| commit | cba357068dc586e7540971eb40bb56c5e62c4de1 (patch) | |
| tree | 0664923eb945f7c31bf78a79cdbeda0b8c69450e | |
| parent | 032f082ffad094649066fa94822c5a78a246a766 (diff) | |
| download | passt-cba357068dc586e7540971eb40bb56c5e62c4de1.tar passt-cba357068dc586e7540971eb40bb56c5e62c4de1.tar.gz passt-cba357068dc586e7540971eb40bb56c5e62c4de1.tar.bz2 passt-cba357068dc586e7540971eb40bb56c5e62c4de1.tar.lz passt-cba357068dc586e7540971eb40bb56c5e62c4de1.tar.xz passt-cba357068dc586e7540971eb40bb56c5e62c4de1.tar.zst passt-cba357068dc586e7540971eb40bb56c5e62c4de1.zip | |
tcp: Don't fast re-transmit if only our FIN is outstandingHEAD2026_10_02.cba3570master
In the TAP_FIN_RCVD path of tcp_tap_handler(), a bare segment from the
guest acknowledging exactly seq_ack_from_tap, with an unchanged window,
is taken as a duplicate ACK and triggers a fast re-transmit.
If the only unacknowledged sequence number is our own FIN, that's
harmful: tcp_rewind_seq() rewinds seq_to_tap and clears TAP_FIN_SENT,
so tcp_data_from_sock() immediately sends the FIN again. If the guest
answers that FIN with the same bare ACK, as a socket in TIME-WAIT will,
we loop at packet rate:
- conn->retries is never incremented on this path, so we never reach
TCP_MAX_RETRIES and tcp_rst()
- ACK_FROM_TAP_DUE is re-armed on every iteration, so the backed-off
re-transmission in tcp_timer_handler() never fires
- TAP_FIN_ACKED can't be set, as it requires TAP_FIN_SENT, which the
rewind just cleared
On an idle Podman host (rootless, pasta), this showed up as a single
flow exchanging ~45,000 54-byte segments per second between pasta and
a container whose socket was in TIME-WAIT, with pasta using ~75% of one
core, until the socket was killed by hand. It recurred on the idle
teardown of an HTTP/2 connection to an ACME server.
With a raw-socket peer driving the same sequence against pasta at
f8df3f1, pasta re-sent the FIN 727,509 times in 10 seconds. With this
change it's re-transmitted by the timer at 1, 3, 7, 15, 31, 63 and 127
seconds, and the connection is reset once TCP_MAX_RETRIES is reached.
Don't consider a duplicate ACK as a fast re-transmit trigger if the
only outstanding sequence number is the FIN, and leave it to the timer.
Fast re-transmit of data is unaffected, with or without a FIN queued
after it.
Fixes: bde1847960cf ("tcp: Fast re-transmit if half-closed, make TAP_FIN_RCVD path consistent")
Link: https://bugs.passt.top/show_bug.cgi?id=125
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Aris Konstantoulas <aris@ariscodes.com>
Signed-off-by: Stefano Brivio <sbrivio@redhat.com>
| -rw-r--r-- | tcp.c | 11 |
1 files changed, 9 insertions, 2 deletions
@@ -2417,15 +2417,22 @@ int tcp_tap_handler(const struct ctx *c, uint8_t pif, sa_family_t af, /* Established connections not accepting data from tap */ if (conn->events & TAP_FIN_RCVD) { + bool fin_only, retr; size_t dlen; - bool retr; if ((dlen = tcp_packet_data_len(th, l4len))) { flow_dbg(conn, "data segment in CLOSE-WAIT (%zu B)", dlen); } - retr = th->ack && !th->fin && + /* If only our FIN is outstanding, rewinding on a duplicate ACK + * would re-send it on every ACK, bypassing the backoff and + * retry limit of tcp_timer_handler(): let the timer do it. + */ + fin_only = (conn->events & TAP_FIN_SENT) && + conn->seq_to_tap == conn->seq_ack_from_tap + 1; + + retr = th->ack && !th->fin && !fin_only && ntohl(th->ack_seq) == conn->seq_ack_from_tap && ntohs(th->window) == conn->wnd_from_tap; |
