feat: prepare Entire Graph for Graphify advertised benchmark parity - #77
Closed
suhaanthayyil wants to merge 18 commits into
Closed
feat: prepare Entire Graph for Graphify advertised benchmark parity#77suhaanthayyil wants to merge 18 commits into
suhaanthayyil wants to merge 18 commits into
Conversation
suhaanthayyil
changed the base branch from
main
to
codex/eg-graphify-stable-base
August 1, 2026 21:28
Collaborator
Author
|
Superseded by #84, the single Entire Graph release PR now targeting main. The older PR is preserved for history; its branch is not deleted. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fairness scope
This branch is the candidate arm. The frozen stable arm remains commit 90a3346. The benchmark harness uses generic tools, identical prompts and limits, blinded grading, and disjoint tune/holdout selectors. No benchmark-specific query strings, prompts, candidate caps, or approximate retrieval paths are used. Results will be added only after the sealed tune and disjoint holdout complete.
Verification
Current benchmark status
The 10 percent tune raw run completed 840 memory cells, 24 code-QA cells, and 72 temporal pairs with zero failed runs. Blinded scoring is still running. This draft does not claim that Entire Graph wins; promotion requires exact-equivalence diagnostics and a disjoint sealed holdout.
Known limitation
The self-repository snapshot was degraded (332/342 files parsed); the benchmark report will retain this rather than hiding it.