Repository navigation
Screen the opener vocabulary before folding it, and store what it settles - #363
Merged
Merged
Conversation
…tles The leaderboard folded all 14,855 candidate openers on every build, building a per-group state dict for each of ~1.4 million response groups to discover that 13,983 of them are still being searched. Measured on the production cache at 872 completed openers, that build took 25.3s against a client that polls every 2s. Screen first. ScoreCache.report_reusable_branch_facts decides the reusability gate in SQL and returns the qualifying keys with three columns instead of every row with six; _screen_and_fold_openers then reaches _candidate_erd_summary only for the openers a fold can settle. Same build, same 872 rows byte-identical, 3.0s. The screen visits every group rather than stopping at the first unsettled one: a candidate holding both an unsettled group and a proven loss is infeasible, not pending, and an early exit can return before reaching the loss that decides it. Groups of fewer than two answers hold no branch result and never will, so the screen skips them instead of asking the cache about them. opener_erd_by_policy stores each completed opener's fold. It is bounded at one row per candidate word rather than at every (branch, candidate) pair the dropped candidate_erd_by_policy was keyed by, which is what makes it affordable to rescreen the whole set on every build instead of trusting a stored row. _store_opener_folds writes the folds the screen settled and deletes the rows whose openers it no longer settles, so a repair or requeue that removes a branch result removes the folds that read it. A row is rewritten only when its value changed, because the build runs on a poll against the cache the swarm is writing into. The table is local to each machine and travels in neither EXPORT_TABLES nor TABLES. candidate_erd_by_policy is now dropped on every writable open rather than once behind a migration flag. A process running code from before the table was removed recreates it, and a migration already recorded as done never looks again -- which is what happened on the production cache, where the table was present again with its original schema months after its migration completed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpacbZRqT46b4Duhxq5QNQ
Owner
Author
|
Review. @codex |
|
Codex Review: Didn't find any major issues. Breezy! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The leaderboard folded all 14,855 candidate openers on every build, constructing
a per-group state dict for each of ~1.4 million response groups in order to
discover that 13,983 of them are still being searched. Measured on the
production cache at 872 completed openers, one build took 25.3s against a
client that polls every 2s. That is why
collect_leaderboard_once's stampedelock exists.
Screen before folding.
ScoreCache.report_reusable_branch_factsdecides thereusability gate in SQL and returns the qualifying keys with three columns,
instead of loading every row of the policy with six and filtering in Python
(4.5s to 0.9s on the live cache).
_screen_and_fold_openersthen reaches_candidate_erd_summaryonly for the openers a fold can settle. The result is3.0s, returning the same 872 rows byte-identical to the shipped collector.
The screen visits every group rather than stopping at the first unsettled one. A
candidate holding both an unsettled group and a proven loss is
infeasibleandnot
pending, because_candidate_erd_summarydecides infeasibility ahead ofpendency, and an early exit can return before reaching the loss that decides it.
Early exit measured 0.13s against 1.0s for the full scan; against a 0.9s load
the full scan is worth its cost. Groups of fewer than two answers hold no branch
result and never will, so the screen skips them instead of asking the cache.
opener_erd_by_policystores each completed opener's fold. The invalidationproblem that retired
candidate_erd_by_policydoes not apply at this key: thetable is bounded at one row per candidate word rather than at every (branch,
candidate) pair, which is what makes it affordable to rescreen the whole set on
every build instead of trusting a stored row.
_store_opener_foldswrites thefolds the screen settled and deletes the rows whose openers it no longer
settles, so a repair or requeue that removes a branch result removes the folds
that read it. Nothing reads a stored fold in place of folding. A row is
rewritten only when its value changed, because the build runs on a poll against
the cache the swarm is writing into. The table is local to each machine and
travels in neither
EXPORT_TABLESnorTABLES.candidate_erd_by_policyis now dropped on every writable open rather thanonce behind a migration flag. A process running code from before the table was
removed recreates it, and a migration already recorded as done never looks
again. That is not hypothetical: the production cache carried the table with its
original schema, empty, two weeks after
drop_candidate_erd_memocompleted. Thecheck is a
sqlite_masterlookup, so carrying it permanently costs one indexedread per open.
AGENTS.md's "a candidate's own ERD is derived, never stored" section is scopedrather than removed: the general rule still holds for a candidate's ERD at an
arbitrary branch, and the section now states what makes the opener fold
different and what keeps it honest.
Schema
opener_erd_by_policyis created by an idempotentCREATE TABLE IF NOT EXISTSin
_ensure_schema. It is additive, so the running swarm's pre-change workersare unaffected, and it is outside the five tables that cross to the phone, so it
needs no deploy sequencing.