Repository navigation
fix(storage): preserve FIFO eviction order across restart for recovered blocks - #23
Open
detail-app[bot] wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Detail bug report: View on Detail
Bug
After a restart, the disk-cache
FifoPickerevicted recovered blocks in a corrupted order — a newer-written recovered block could be reclaimed before an older-written one.Root cause:
BlockManager::init(the recovery path) rebuilt the evictable set as aHashSet<BlockId>and iterated it in arbitrary order when notifying eviction pickers.FifoPickerderives FIFO order entirely from the call order ofon_block_evictable(push_back), so the per-process-randomHashSetiteration permuted its queue. Immediately after recovery every block hasinvalid == 0, soInvalidRatioPickerreturnsNoneandFifoPickeris the sole deciding picker for the first post-restart eviction round — making the corruption directly observable. (Bug introduced by the recovery-path refactor in foyer-rs#1074, which batched picker notifications over aHashSet.)Fix
Derive the recovered blocks' recency from the persisted per-entry write
sequenceand feed them to the pickers oldest-written-first, instead of from aHashSet.foyer-storage/src/engine/block/recover.rs: while scanning recovered entries, aggregate each block's recency =max(addr.sequence)over its entries, then sort the evictable blocks ascending by that sequence via a neworder_evictable_blocks_by_sequencehelper before handing them toBlockManager::init.foyer-storage/src/engine/block/manager.rs:BlockManager::initnow takes an orderedevictable_blocks: &[BlockId]and iterates it in order, with a doc comment stating the contract (oldest-written-first). It no longer rebuilds/iterates aHashSet.Recency is keyed on
addr.sequence, not onBlockId, so the order stays correct even after block ids are recycled by reclamation (a reclaimed id returns to the back of the clean list and can be reused, so a low id may hold newer data than a high id). The live runtime path (on_writing_finish→push_backin real write order) is untouched and remains correct.Testing
recover::tests): the sort is keyed on sequence rather than block id (a low-id block holding the newest write is ordered last — covers the recycled-id case), handles empty input, and tie-breaks stably by ascending block id.engine::tests::test_store_fifo_eviction_order_after_recovery): writes one entry per block across 7 of 8 blocks, restarts, then drives 4 sequential evictions and asserts the recovered blocks drain in strict oldest→newest write order (keys 0→1→2→3). This passes deterministically on the fix; I verified it catches the regression by temporarily restoring the originalHashSet-basedinit(test failed) and then restoring the fix (test passed).foyer-storagesuite (30 tests) and the downstreamfoyerhybrid-cache suite (15 tests, includingtest_load_after_recoveryand the hybrid fuzzy test) all pass. Routinecargo check,clippy -- -D warnings, and stable + nightlycargo fmt --checkare clean.Automatic Fixes PRs can be configured here.