Skip to content

Leaderboard ships every row's full response-group breakdown, projecting to 62.9 MB #367

Description

@ahernsean

Problem

The leaderboard ships the full response-group breakdown for every displayed
row. Measured at 872 complete openers: 3.69 MB of JSON and ~96,000 DOM
nodes
, of which 85,825 are answer-strip <span>s. Median 98 response groups
per card, max 161.

Projected to 14,855 openers: 62.9 MB and ~1,640,000 DOM nodes.

A 63 MB response is a wall that cannot be rendered past, on a phone least of
all. It is also the wall that virtualization does not touch — the browser still
downloads and parses every byte before the first row draws.

What to do

Ship a cheap glyph by default and the full breakdown only on expand or hover.
Concretely, response_groups leaves the default payload; a card carries enough
to draw a summary strip (a small fixed-size histogram, or a bitmap in the shape
of PR #365) and fetches its own detail when opened.

Why this one first

The code already commits to this pattern and stops halfway.
collect_leaderboard_report builds response_groups only for displayed_rows,
never for all ranked_rows — the expansion is already deferred past the limit.
What remains is that the client's default limit is "everything", so the
deferral never fires in practice.

This is the keystone of the large-item-count work: it is the only item that
attacks the payload wall at its source, and both the distribution overview and
virtualized rendering are decorating a page that still will not load until it
lands.

Also in scope: the server-side twin of the same defect

Measured 20 Sep 2026 at 889 completed openers, a leaderboard build is two
costs with different shapes
:

cost shape
_candidate_group_skeletons 18.1 s, ~0.25 GB resident once per process, whole vocabulary
screen + fold 3.5 s per build; 2.4 s of it reading 652,989 rows

The 18 s does not grow with completed openers — it already runs all 14,855
candidates. It is memoized in a module global, so it is paid once per process;
but the memo dies with the process, and the report cache is cold then too, so
the first leaderboard view after every swarm restart pays it.

What it produces is 1,389,596 response groups holding 238 MB of
encode_subset strings
(183 MB deduplicated across 637,811 distinct keys) —
a re-expansion of the PatternMatrix, which is 47.7 MB and already
memory-mapped. Persisting the skeletons is therefore the wrong fix: it writes
5x the matrix to disk to avoid recomputing something the matrix implies.

The same partition computes in 0.86 s with two np.bincount calls per
candidate against an order-independent hash of each group's member indices.

_screen_and_fold_openers — the only consumer of all 1.39 M groups — uses each
branch key for exactly two set-membership tests and never reads the words. So
hashing both sides makes the screen integer work, and exact branch_key strings
are needed only for the openers that screen complete (889 today, ~84k groups
rather than 1.39 M).

Hash equality is implied by key equality, so there are no false negatives: an
opener that should screen complete always does. The only error mode is a false
complete, and the exact fold that follows re-materializes real keys and
rejects it — the screen stays conservative in the safe direction, which is what
permits it to be approximate at all.

Expected: restart cliff 18 s -> ~1 s, memo 0.25 GB -> ~14 MB, no new file on
disk.

This belongs here rather than in its own issue because it is the same decision
as the payload split — which rows need their full breakdown materialized, and
which need only enough to be ranked and summarized.


Blocks #369 (distribution overview) and #371 (virtualized rendering).

Activity

  1. added
    enhancementNew feature or request
    P1Blocks other work, or produces a wrong answer that will mislead
    on Sep 20, 2026
  2. ahernsean commented on Sep 20, 2026

    @ahernsean
    OwnerAuthor

    Part of #376.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1Blocks other work, or produces a wrong answer that will misleadenhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions