Skip to content

feat(proxy): retry another backend when the picked one drains mid-dial - #179

Merged
lexfrei merged 1 commit into
masterfrom
feat/retry-on-drain-refusal
Oct 8, 2026
Merged

lexfrei merged 1 commit into
masterfrom
feat/retry-on-drain-refusal

Conversation

@lexfrei

@lexfrei lexfrei commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

A client connection now gets one more backend when the first one starts draining while the proxy is still dialing it. The proxy picks a backend, dials it, and only then registers the connection. If the endpoint was removed during that dial, registration was refused and the client got EOF, even with healthy backends left.

Now the proxy closes only the upstream side it just dialed and picks once more from the pickable backends. There is one retry. If the second attempt is refused too, its dial fails, or nothing is left to pick, the client is closed with the same accounting as before. An empty pick counts under none. The retried refusal itself does not count as a connection error, because the client did not fail. The retry is logged at info level. In the worst case the client now waits for two dials instead of one.

Two cases are deliberately not retried. A refusal during shutdown closes the client as before, and so does a failed dial: retrying failed dials is a different change. Tests pin both cases.

The connection handler is split into pick, dial and serve steps, so the retry reuses them and tests can enter each step with prepared backends.

Closes #169

@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

Container image available

Multi-arch image (amd64 + arm64):

podman pull ttl.sh/extractedprism:pr-179-1d

Note: Image is rebuilt on every commit and expires after 24 hours (hosted on ttl.sh).

@lexfrei
lexfrei force-pushed the feat/retry-on-drain-refusal branch from 228f0c3 to 93d9611 Compare October 8, 2026 20:53
The proxy picks a backend, dials it, and only then registers the
connection. If the backend starts draining while the dial is in
flight, registration is refused and the client was closed, even when
other healthy backends were available.

Such a refusal now closes only the dialed upstream side and serves
the client from one more pick among the pickable backends. There is
a single retry: a second refusal, a failed retry dial or an empty
pick fails the client as before, and an empty pick is counted under
the "none" label. The retried refusal itself is not counted as a
connection error, since the client was not failed by it.

A refusal during shutdown and a failed dial are not retried.

handleConn is split into pick, dial and serve stages so the retry can
reuse them and tests can enter each stage with prepared state.

Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
@lexfrei
lexfrei force-pushed the feat/retry-on-drain-refusal branch from 93d9611 to 15c5662 Compare October 8, 2026 20:59
@lexfrei
lexfrei merged commit f65a6bb into master Oct 8, 2026
8 checks passed
@lexfrei
lexfrei deleted the feat/retry-on-drain-refusal branch October 8, 2026 21:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(proxy): retry another backend when the picked one starts draining mid-dial

1 participant