Repository navigation
fix(discovery): re-watch without backoff after a routine stream end - #177
Merged
Merged
Conversation
lexfrei
force-pushed
the
fix/rewatch-backoff
branch
2 times, most recently
from
October 8, 2026 20:21
f365397 to
fa9bf6d
Compare
Container image availableMulti-arch image (amd64 + arm64): podman pull ttl.sh/extractedprism:pr-177-1d
|
lexfrei
force-pushed
the
fix/rewatch-backoff
branch
5 times, most recently
from
October 8, 2026 20:52
37cc4a8 to
d4d1fda
Compare
This was referenced Oct 8, 2026
lexfrei
force-pushed
the
fix/rewatch-backoff
branch
from
October 8, 2026 21:05
d4d1fda to
15b1584
Compare
When a 410 Gone re-list failed, the provider reported it to the error hook without checking why it failed. A shutdown that cancels the context during the re-list was counted as a discovery error, so every such shutdown raised the error metric. A re-list that fails because the context is done now returns quietly with the cached endpoints restored. Fixes: 16e2185 ("feat(discovery): name providers and record discovery metrics") Assisted-by: LLM Signed-off-by: Aleksei Sviridkin <f@lex.la>
The API server ends every watch after its own timeout, and the provider treated that like a failure: it logged a warning, bumped the attempt counter and slept before the next watch. The counter never reset on this path, so after a few hours each routine re-watch waited the full 30 second backoff, and EndpointSlice changes in that window arrived late. A watch that stayed open for at least the backoff ceiling now counts as healthy. A stream that ends after a healthy lifetime is re-opened at once and logged at debug level, and any other error after a healthy watch restarts the backoff from the first attempt. The lifetime starts when the Watch call returns a stream, so a Watch call that fails never counts as healthy, however long it blocked. A watch that fails or ends sooner keeps the backoff, so a server or proxy that closes streams immediately cannot drive a tight re-watch loop. A 410 Gone answered by a successful re-list no longer logs a warning. The watch also asks for bookmarks. The kubernetes EndpointSlice rarely changes, so without them the stored resource version goes stale and the immediate re-watch can get 410 Gone. On servers with the watch cache, bookmarks keep the version fresh. Assisted-by: LLM Signed-off-by: Aleksei Sviridkin <f@lex.la>
After a 410 Gone the provider re-lists, and a successful re-list reset the attempt counter and opened the next watch at once. If the server answers every fresh watch with 410, for example a member with a broken watch cache behind the load balancer, the provider ran list and watch in a loop with no delay at all. A successful re-list now skips the backoff only when the watch that ended in 410 had stayed open for a healthy lifetime, which is the routine case of a long watch outliving etcd compaction. A 410 on a short-lived watch waits for the backoff like any other early failure. One more 410 skips the backoff: the one on the re-watch right after a healthy stream closed, since that re-watch can carry a resource version the server has already dropped. That allows at most one quick re-list per healthy stream. Fixes: a09adab ("fix(kubernetes): add watch backoff and handle 410 Gone with re-list (#17)") Assisted-by: LLM Signed-off-by: Aleksei Sviridkin <f@lex.la>
The provider recognized an expired resource version only as an error event inside the watch stream. When the server refused the version on the Watch call instead, the error was wrapped as a generic watch failure: no re-list ran, and every retry sent the same stale version after a backoff, so discovery never recovered. client-go's reflector checks the Watch call for this error too. An expired or gone error from the Watch call now takes the same path as a 410 event: re-list, then the usual 410 backoff rules. Fixes: a09adab ("fix(kubernetes): add watch backoff and handle 410 Gone with re-list (#17)") Assisted-by: LLM Signed-off-by: Aleksei Sviridkin <f@lex.la>
lexfrei
force-pushed
the
fix/rewatch-backoff
branch
from
October 8, 2026 21:17
15b1584 to
26669c4
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
After a few hours on a healthy cluster, the Kubernetes provider waits the full 30 second backoff every time the API server closes a watch on its own timeout. EndpointSlice changes in that gap reach the load balancer late. The provider treated a routine stream end as a failure: a warning in the log, one more backoff attempt, and a sleep before the next watch. Nothing reset the attempt counter on that path, so it climbed to the cap and stayed there.
A watch that stayed open for at least 30 seconds now counts as healthy. When such a stream ends, the provider opens a new watch right away and logs it at debug level. Any other error after a healthy watch, except 410 Gone, starts the backoff again from the first attempt. The watch now asks for bookmarks, so on servers with the watch cache the resource version stays fresh and the quick re-watch usually doesn't hit 410 Gone.
A watch that fails or ends sooner still waits for its backoff, so a server or proxy that closes streams at once can't cause a tight re-watch loop. The lifetime starts when the Watch call returns a stream, so a failed Watch call never looks healthy, even if it blocked for a long time. I took 30 seconds because it is the backoff ceiling: even a proxy that cuts every stream right at that mark makes the provider open watches no more often than the backoff already does.
A 410 Gone with a successful re-list doesn't log a warning anymore. The re-list already logs it at info level. A re-list that fails because the provider is shutting down doesn't count as a discovery error anymore. If a watch gets 410 Gone before it was open for 30 seconds, the provider now waits for the backoff after the re-list, so a server that answers every new watch with 410 can't keep it re-listing in a loop. A 410 on the first watch after a routine stream end still skips the backoff, so a re-watch with an outdated version doesn't bring the pause back. The same re-list now also runs when the server refuses an outdated version on the Watch call itself, not only inside the stream. Before, that case retried the same version forever.
The new tests inject a clock and a backoff function, so they run without sleeps. They check both sides of the 30 second boundary, a failed Watch call, the counter reset and the 410 Gone log level.
Closes #168
Closes #180
Closes #181
Closes #182