feat(core): add PDF page citations to document observations - #1490
Conversation
Signed-off-by: phernandez <paul@basicmachines.co>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 11c7334898
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Signed-off-by: phernandez <paul@basicmachines.co>
|
@codex review |
|
Codex Review: Didn't find any major issues. More of your lovely PRs please. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
Why
Refs #1366. Document observations need to point back to the source PDF page so readers can inspect the evidence. This is the page-level slice agreed for v0.24.0, following the amended SPEC-89: the observation is the claim, and citations use OKF footnotes joined to stable
sources[].idvalues.What Changed
page, with an optional separate printedpage_label.sourcesentries and Markdown footnotes. Agents provide page numbers, not destination URLs.docs/DOCUMENT_CITATIONS.md.Implementation Details
The page number is validated against the actual extraction page count. Citation resources use project/bundle-root-relative, percent-encoded file paths with standard
#page=Nfragments. IDs such asdocument-page-2are scoped to the note's single trusted source PDF, not an array index; repeated citations share an entry and reordering does not retarget them. Printed labels never change the navigation page.The existing singular
sourceretains the original PDF checksum and storage-version provenance. The new pluralsourcessupplies citation targets. Trusted-envelope validation checks those entries against the source PDF; generated footnote definitions are not agent-supplied. A separating space prevents trailing Markdown escapes in observation text from escaping the generated marker.Specifications and prior art
page=Naddressing convention.Testing
uv run pytest test-int/test_document_page_citations.py --no-cov -q: 19 passed on SQLite.BASIC_MEMORY_TEST_POSTGRES=1 uv run pytest test-int/test_document_page_citations.py --no-cov -q: 19 passed against real PostgreSQL via testcontainers.uv run pytest tests/schemas/test_document.py tests/schemas/test_document_agent_temporal.py tests/document_ingestion --no-cov -q: 118 passed.just fast-check: passed (lint, formatting, typecheck).just doctor: passed the isolated file/API/index/search/status loop.codex review --uncommitted: run before pushing; its trailing-backslash finding was fixed with a real Markdown-rendering regression test. The second pass reported no actionable defects (76 focused tests passed in the review).Risks / Follow-ups
Review follow-up: generated citation references as well as definitions are reserved across agent-controlled Markdown fields when locators are present. Omitted page labels are compatible with an explicit label in either observation order; only distinct explicit labels conflict. Explicit tests also cover duplicate source IDs and non-PDF sources. The follow-up pre-push review found no actionable regressions (19 citation tests passed).