Enable share groups by default and enhance protocol gates - #51
Merged
Merged
Conversation
…nest Share groups stopped being experimental in Apache Kafka 4.2, and krafka kept hiding them behind `unstable-protocol` for two more releases. The gate contradicted its own data: every share API is marked stable in the vendored 4.3 snapshot and the version table negotiated all of them unconditionally, so only the API that could reach the protocol was hidden. BREAKING CHANGE: the KIP-932 share consumer moved to a `share-groups` feature, on by default. `unstable-protocol` now means only what Kafka marks `latestVersionUnstable` — ApiVersions v5 and InitProducerId v6. Gates added: - protocol-parity check 5 fails when a row is gated on `unstable-protocol` but Kafka ships that version as stable, so this class cannot recur. - ci-job-parity asserts every `just ci` recipe has a CI job, that the new ci-success aggregator needs every job, and that it carries `if: always()` (GitHub reads a skipped required check as success). - bench-check gates the producer send path against a stored baseline. An end-to-end fake-broker benchmark was built and removed once before because it could not rank compression codecs; that is a within-run comparison, where the harness constant swamps the signal. A regression gate is an across-run comparison, where the same constant sits on both sides and cancels. - semver-check classifies API changes against the last published release. - docs-test and a publish-time verify job now actually run. Documentation is current state only, not history: - The README's 536-line upgrade section is gone; CHANGELOG.md already covered every version of it. - The protocol reference no longer flags the four share APIs as unstable, the share-consumer guide no longer recounts Kafka's own release history, and the configuration guide describes behaviour rather than what it used to be. - The performance guide and landing page reflect the regression gate while keeping the standard that no absolute figure is published. - fuzz/README.md documented three of six targets.
NOT_COORDINATOR, COORDINATOR_NOT_AVAILABLE and COORDINATOR_LOAD_IN_PROGRESS
all mean "re-run FindCoordinator and try again". A freshly started or
rebalancing cluster answers this way routinely, and the Java client retries
transparently.
krafka dropped the cached coordinator and then returned the error anyway.
`invalidate_coordinator_on_error` returns a bool documented as "retriable
after re-discovery" and both call sites discarded it, so nothing ever made the
next attempt — a routine coordinator election surfaced to the application as
`Failed to subscribe: Broker { code: NotCoordinator }`.
Both join paths now re-discover and retry with jittered backoff, bounded at
five attempts. The KIP-848 path had the same defect in a worse form: it
returned the error without invalidating the cached coordinator at all, so
nothing downstream could recover either.
Found by the Redpanda integration suite, which is where a real coordinator
election actually happens. Each fix carries a fake-broker regression test,
both verified against the defect they cover.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Share groups stopped being experimental in Apache Kafka 4.2, and krafka kept hiding them behind
unstable-protocolfor two more releases. The gate contradicted its own data: every share API is marked stable in the vendored 4.3 snapshot and the version table negotiated all of them unconditionally, so only the API that could reach the protocol was hidden.BREAKING CHANGE: the KIP-932 share consumer moved to a
share-groupsfeature, on by default.unstable-protocolnow means only what Kafka markslatestVersionUnstable— ApiVersions v5 and InitProducerId v6.Gates added:
unstable-protocolbut Kafka ships that version as stable, so this class cannot recur.just cirecipe has a CI job, that the new ci-success aggregator needs every job, and that it carriesif: always()(GitHub reads a skipped required check as success).Documentation is current state only, not history: