Repository navigation
Preserve cache prefixes and task revisions across long-running sessions - #30
Merged
Merged
Conversation
CormickKneey
marked this pull request as ready for review
October 1, 2026 16:47
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Long sessions could lose later user corrections after compaction, invalidate cache prefixes when Goal context changed, and spend seconds repeatedly serializing full thread snapshots. This PR preserves request identity within a context window, restores real user revisions independently of summaries, and reduces durable snapshot overhead without weakening recovery semantics.
Changes
limits.context_target_tokens,model.summary_reasoning_effort, andmodel.summary_max_output_tokens, including environment overrides, validation and documentation. Summary caps cannot exceed global or Goal output allowances. Existing defaults remain unchanged.tool_choice=none, and preserve full accounting APIs. Audit/reporting distinguish unknown usage from misses and correlate provider/gateway request IDs.End-to-end validation
Research, reproduction and limits · English · JSON results
Real gpt-6-sol tasks read/write files, receive five repair requirements followed by a separate revision, cross repeated compaction, restart Core, and finish a Goal. Independent assertions verify every final JSON field, nonce, sum, an unchanged accepted artifact and a real successful verification command. The requirements are not saved in a workspace sidecar.
Both completed, passed all artifact assertions and had zero unknown usage. The small test-only context window deliberately causes frequent cache-prefix rebuilds; these are continuity stress results, not steady-state cache benchmarks. Setup credential failures and an initial grader correction accepting successful
verify_commandare retained in the report, not hidden.Regression includes 8,027-byte valid summaries, omitted/invalid summaries, multiple revisions, restart, steering inside an automatic continuation, long steering during summarization, impossible reductions, target headroom, summary parameter isolation/Goal caps and cancelled usage. Engine/config tests, workspace/all-targets Clippy, formatting/docs checks and native Harness smoke passed locally. Native smoke exercises actual read/patch/command, hooks, Core SIGKILL, UNKNOWN inspection and no replay.
On the development
/datafilesystem, a 9,399,322-byte debug-profile snapshot benchmark produced identical bytes in all three pairs: median encoding/clone/write/fsync fell from 4495.5 ms to 847.2 ms (81.2%, 5.3×). This excludes admission and directory commit, and is not a whole-task speedup. Reproduction is an ignored benchmark, not a flaky CI timing assertion.All five Verify checks passed on ae78ffa: https://github.com/areal-project/AReaL-Harness/actions/runs/37113694208 . This includes macOS native Harness, Linux portable checks, Linux container Runtime, the combined Linux gate and dependency advisories.
Additional local validation boundary: shared-service PTY smoke on the development host hit its 20-second startup wait with full debug symbols; after rebuilding with the CI profile it reached a 5-second Ctrl-Q exit timeout. The same script passed in final macOS native CI. These local failures are not counted as passes, their logs are retained (
/tmp/areal-continuity-local-service.log,/tmp/areal-continuity-local-service-ci-profile.log), and no test timeout was relaxed. The TUI timing issue is not claimed fixed by this PR. The configuration-archive compatibility cases and all-targets Clippy passed locally.Cache measurements and boundaries
The earlier interleaved Chat task comparison verified four Goals: main input/cached/uncached totals 217012/161152/55860 (74.26%) versus candidate 164324/144384/19940 (87.87%); prefix preservation 0/18 versus 18/18. Candidate mean elapsed time was worse, 234.01s versus 120.70s. This is not a latency claim.
Across all eight successful pre-session-header Responses transport runs, HTTP used 1820237 wire bytes at 83.03% weighted cache rate; WebSocket used 1027194 bytes at 84.01%. A later interleaved set favored HTTP on cache rate. The repeatable transport benefit is roughly 44% fewer wire bytes, not reliable cache/cost/latency superiority. Matching Core thread identity in the handshake also passed a real Goal, but cannot ensure backend cache residency. No universal 99% or industry-best claim is made.
Pinned references: Codex compaction, OpenCode compaction, and Codex request continuation. We preserve actual user source data rather than trusting summary completeness. Very large user inputs can still prevent reduction; existing children require explicit delivery of relevant revisions.
No production tasks, budgets, ledger entries or deployment binaries are changed. Studio's review wrapper, diagnostic tail/cancellation attribution and archive scanning are external components; browser temporary-path/input adapters and other production candidate patches are not claimed fixed by this PR. The report records the operational follow-ups. These synthetic tests do not establish Golden game acceptance.