Skip to content

fix(discovery): re-watch without backoff after a routine stream end - #177

Merged
lexfrei merged 4 commits into
masterfrom
fix/rewatch-backoff
Oct 8, 2026
Merged

lexfrei merged 4 commits into
masterfrom
fix/rewatch-backoff

Conversation

@lexfrei

@lexfrei lexfrei commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

After a few hours on a healthy cluster, the Kubernetes provider waits the full 30 second backoff every time the API server closes a watch on its own timeout. EndpointSlice changes in that gap reach the load balancer late. The provider treated a routine stream end as a failure: a warning in the log, one more backoff attempt, and a sleep before the next watch. Nothing reset the attempt counter on that path, so it climbed to the cap and stayed there.

A watch that stayed open for at least 30 seconds now counts as healthy. When such a stream ends, the provider opens a new watch right away and logs it at debug level. Any other error after a healthy watch, except 410 Gone, starts the backoff again from the first attempt. The watch now asks for bookmarks, so on servers with the watch cache the resource version stays fresh and the quick re-watch usually doesn't hit 410 Gone.

A watch that fails or ends sooner still waits for its backoff, so a server or proxy that closes streams at once can't cause a tight re-watch loop. The lifetime starts when the Watch call returns a stream, so a failed Watch call never looks healthy, even if it blocked for a long time. I took 30 seconds because it is the backoff ceiling: even a proxy that cuts every stream right at that mark makes the provider open watches no more often than the backoff already does.

A 410 Gone with a successful re-list doesn't log a warning anymore. The re-list already logs it at info level. A re-list that fails because the provider is shutting down doesn't count as a discovery error anymore. If a watch gets 410 Gone before it was open for 30 seconds, the provider now waits for the backoff after the re-list, so a server that answers every new watch with 410 can't keep it re-listing in a loop. A 410 on the first watch after a routine stream end still skips the backoff, so a re-watch with an outdated version doesn't bring the pause back. The same re-list now also runs when the server refuses an outdated version on the Watch call itself, not only inside the stream. Before, that case retried the same version forever.

The new tests inject a clock and a backoff function, so they run without sleeps. They check both sides of the 30 second boundary, a failed Watch call, the counter reset and the 410 Gone log level.

Closes #168
Closes #180
Closes #181
Closes #182

@lexfrei
lexfrei force-pushed the fix/rewatch-backoff branch 2 times, most recently from f365397 to fa9bf6d Compare October 8, 2026 20:21
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

Container image available

Multi-arch image (amd64 + arm64):

podman pull ttl.sh/extractedprism:pr-177-1d

Note: Image is rebuilt on every commit and expires after 24 hours (hosted on ttl.sh).

When a 410 Gone re-list failed, the provider reported it to the error
hook without checking why it failed. A shutdown that cancels the
context during the re-list was counted as a discovery error, so every
such shutdown raised the error metric. A re-list that fails because the
context is done now returns quietly with the cached endpoints restored.

Fixes: 16e2185 ("feat(discovery): name providers and record discovery metrics")
Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
The API server ends every watch after its own timeout, and the provider
treated that like a failure: it logged a warning, bumped the attempt
counter and slept before the next watch. The counter never reset on
this path, so after a few hours each routine re-watch waited the full
30 second backoff, and EndpointSlice changes in that window arrived
late.

A watch that stayed open for at least the backoff ceiling now counts as
healthy. A stream that ends after a healthy lifetime is re-opened at
once and logged at debug level, and any other error after a healthy
watch restarts the backoff from the first attempt. The lifetime starts
when the Watch call returns a stream, so a Watch call that fails never
counts as healthy, however long it blocked. A watch that fails or ends
sooner keeps the backoff, so a server or proxy that closes streams
immediately cannot drive a tight re-watch loop. A 410 Gone answered by
a successful re-list no longer logs a warning.

The watch also asks for bookmarks. The kubernetes EndpointSlice rarely
changes, so without them the stored resource version goes stale and
the immediate re-watch can get 410 Gone. On servers with the watch
cache, bookmarks keep the version fresh.

Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
After a 410 Gone the provider re-lists, and a successful re-list reset
the attempt counter and opened the next watch at once. If the server
answers every fresh watch with 410, for example a member with a broken
watch cache behind the load balancer, the provider ran list and watch
in a loop with no delay at all.

A successful re-list now skips the backoff only when the watch that
ended in 410 had stayed open for a healthy lifetime, which is the
routine case of a long watch outliving etcd compaction. A 410 on a
short-lived watch waits for the backoff like any other early failure.

One more 410 skips the backoff: the one on the re-watch right after a
healthy stream closed, since that re-watch can carry a resource version
the server has already dropped. That allows at most one quick re-list
per healthy stream.

Fixes: a09adab ("fix(kubernetes): add watch backoff and handle 410 Gone with re-list (#17)")
Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
The provider recognized an expired resource version only as an error
event inside the watch stream. When the server refused the version on
the Watch call instead, the error was wrapped as a generic watch
failure: no re-list ran, and every retry sent the same stale version
after a backoff, so discovery never recovered. client-go's reflector
checks the Watch call for this error too.

An expired or gone error from the Watch call now takes the same path
as a 410 event: re-list, then the usual 410 backoff rules.

Fixes: a09adab ("fix(kubernetes): add watch backoff and handle 410 Gone with re-list (#17)")
Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
@lexfrei
lexfrei force-pushed the fix/rewatch-backoff branch from 15b1584 to 26669c4 Compare October 8, 2026 21:17
@lexfrei
lexfrei merged commit 7d42183 into master Oct 8, 2026
8 checks passed
@lexfrei
lexfrei deleted the fix/rewatch-backoff branch October 8, 2026 21:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment