Skip to content

Preserve cache prefixes and task revisions across long-running sessions - #30

Merged
CormickKneey merged 14 commits into
mainfrom
fix/cache-aware-chat-context
Oct 3, 2026
Merged

CormickKneey merged 14 commits into
mainfrom
fix/cache-aware-chat-context

Conversation

@CormickKneey

@CormickKneey CormickKneey commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Long sessions could lose later user corrections after compaction, invalidate cache prefixes when Goal context changed, and spend seconds repeatedly serializing full thread snapshots. This PR preserves request identity within a context window, restores real user revisions independently of summaries, and reduces durable snapshot overhead without weakening recovery semantics.

Changes

  • Replay compacted real user messages verbatim and chronologically from the existing archive. Exclude the automatic Goal-continuation seed by persisted origin while keeping subsequent steering. Summaries are historical evidence, not authority to undo user corrections; no game-specific or inferred second task contract is introduced.
  • Check actual projected net saving before summarizing. Probe the recent-history and deepest complete-round boundaries, preserving tool/result and opaque-reasoning groups. Accept valid summaries up to 16 KiB when space permits, remove internal user-role summary controls, and prevent recursive degraded summaries. Steering received during summarization is included in both sides of the commit-time comparison.
  • Add optional limits.context_target_tokens, model.summary_reasoning_effort, and model.summary_max_output_tokens, including environment overrides, validation and documentation. Summary caps cannot exceed global or Goal output allowances. Existing defaults remain unchanged.
  • Encode each thread once for capacity checking and atomic writing. Remove repeated clones/serialization and unbuffered incremental JSON writes while retaining durable intent before execution, file/directory fsync, atomic replacement and UNKNOWN recovery. Add stage timing for tools and persistence. Cancelled open streams get a bounded one-second usage drain without executing output tools; genuinely missing usage remains UNKNOWN.
  • Preserve original encoded model-archive entries and validate their original digests across configuration upgrades. Optional defaults no longer invalidate historical revisions; four legacy layouts survive repeated reload while tampering is still rejected.
  • Keep append-only typed runtime context, deduplicate unchanged Goal projections, omit volatile prompt counters, retain final-round tool schemas with tool_choice=none, and preserve full accounting APIs. Audit/reporting distinguish unknown usage from misses and correlate provider/gateway request IDs.
  • Keep experimental Responses WebSocket continuation off by default. Reuse requires matching Thread/Turn ownership, unchanged non-input properties and an exact input-plus-complete-output baseline; otherwise open a fresh connection. TLS verification, bounded idle storage and close-on-failure remain enforced. Direct connections only; no HTTP-proxy environment support or automatic HTTP fallback.

End-to-end validation

Research, reproduction and limits · English · JSON results

Real gpt-6-sol tasks read/write files, receive five repair requirements followed by a separate revision, cross repeated compaction, restart Core, and finish a Goal. Independent assertions verify every final JSON field, nonce, sum, an unchanged accepted artifact and a real successful verification command. The requirements are not saved in a workspace sidecar.

Protocol Turns Compactions Tool calls Model requests Input / cached tokens Cache rate Seconds
Chat Completions 7 6 19 27 206935 / 119552 57.77% 165.9
Responses 7 7 18 28 212487 / 107264 50.48% 220.2

Both completed, passed all artifact assertions and had zero unknown usage. The small test-only context window deliberately causes frequent cache-prefix rebuilds; these are continuity stress results, not steady-state cache benchmarks. Setup credential failures and an initial grader correction accepting successful verify_command are retained in the report, not hidden.

Regression includes 8,027-byte valid summaries, omitted/invalid summaries, multiple revisions, restart, steering inside an automatic continuation, long steering during summarization, impossible reductions, target headroom, summary parameter isolation/Goal caps and cancelled usage. Engine/config tests, workspace/all-targets Clippy, formatting/docs checks and native Harness smoke passed locally. Native smoke exercises actual read/patch/command, hooks, Core SIGKILL, UNKNOWN inspection and no replay.

On the development /data filesystem, a 9,399,322-byte debug-profile snapshot benchmark produced identical bytes in all three pairs: median encoding/clone/write/fsync fell from 4495.5 ms to 847.2 ms (81.2%, 5.3×). This excludes admission and directory commit, and is not a whole-task speedup. Reproduction is an ignored benchmark, not a flaky CI timing assertion.

All five Verify checks passed on ae78ffa: https://github.com/areal-project/AReaL-Harness/actions/runs/37113694208 . This includes macOS native Harness, Linux portable checks, Linux container Runtime, the combined Linux gate and dependency advisories.

Additional local validation boundary: shared-service PTY smoke on the development host hit its 20-second startup wait with full debug symbols; after rebuilding with the CI profile it reached a 5-second Ctrl-Q exit timeout. The same script passed in final macOS native CI. These local failures are not counted as passes, their logs are retained (/tmp/areal-continuity-local-service.log, /tmp/areal-continuity-local-service-ci-profile.log), and no test timeout was relaxed. The TUI timing issue is not claimed fixed by this PR. The configuration-archive compatibility cases and all-targets Clippy passed locally.

Cache measurements and boundaries

The earlier interleaved Chat task comparison verified four Goals: main input/cached/uncached totals 217012/161152/55860 (74.26%) versus candidate 164324/144384/19940 (87.87%); prefix preservation 0/18 versus 18/18. Candidate mean elapsed time was worse, 234.01s versus 120.70s. This is not a latency claim.

Across all eight successful pre-session-header Responses transport runs, HTTP used 1820237 wire bytes at 83.03% weighted cache rate; WebSocket used 1027194 bytes at 84.01%. A later interleaved set favored HTTP on cache rate. The repeatable transport benefit is roughly 44% fewer wire bytes, not reliable cache/cost/latency superiority. Matching Core thread identity in the handshake also passed a real Goal, but cannot ensure backend cache residency. No universal 99% or industry-best claim is made.

Pinned references: Codex compaction, OpenCode compaction, and Codex request continuation. We preserve actual user source data rather than trusting summary completeness. Very large user inputs can still prevent reduction; existing children require explicit delivery of relevant revisions.

No production tasks, budgets, ledger entries or deployment binaries are changed. Studio's review wrapper, diagnostic tail/cancellation attribution and archive scanning are external components; browser temporary-path/input adapters and other production candidate patches are not claimed fixed by this PR. The report records the operational follow-ups. These synthetic tests do not establish Golden game acceptance.

@CormickKneey
CormickKneey marked this pull request as ready for review October 1, 2026 16:47
@CormickKneey CormickKneey changed the title perf: preserve Chat prefixes and deduplicate Goal status perf: stabilize Agent cache prefixes and add opt-in Responses continuation Oct 2, 2026
@CormickKneey CormickKneey changed the title perf: stabilize Agent cache prefixes and add opt-in Responses continuation Preserve cache prefixes and task revisions across long-running sessions Oct 3, 2026
@CormickKneey
CormickKneey merged commit d865fe7 into main Oct 3, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants