Skip to content

feat: cluster split validation — Phase 2 benchmark honesty - #6

Closed
cschanhniem wants to merge 1 commit into
mainfrom
feat/cluster-split-validation
Closed

cschanhniem wants to merge 1 commit into
mainfrom
feat/cluster-split-validation

Conversation

@cschanhniem

Copy link
Copy Markdown
Collaborator

Summary

Implements Phase 2: Cluster split requirement from AGENTS.md:

"Pipeline still performs when near-duplicates are removed"

  • cluster_by_similarity(sequences, threshold) — greedy single-linkage clustering; sequences within threshold normalized Levenshtein similarity are co-clustered
  • cluster_split(sequences, threshold) — splits a sequence set into (reference_indices, test_indices) so no cluster spans both sides; guarantees no test sequence has a near-duplicate in the reference set
  • find_contaminated_references() — identifies reference sequences that are near-duplicates of test positives; removes benchmark inflation from reference-set memorization
  • Full benchmark evaluation suite: recall_at_k, random_recall_at_k, enrichment_factor, benchmark_summary added to evaluate.py

What the tests prove

Test Result
Near-duplicates co-cluster at threshold ≥ similarity ✅
Cluster split puts one rep per cluster in reference ✅
Contaminated references correctly identified ✅
Positives score above negatives without any references ✅ (feature-based)
EF > 1.0 after cluster split (no reference) ✅
recall@3 ≥ random baseline after split ✅
≥ 2 of top-3 ranked are AMP-like positives ✅

The key finding: AMP-like sequences score above non-AMP negatives on activity/safety/synthesis features alone, without needing reference proximity — confirming the pipeline is not inflated by memorization.

Test plan

  • make test — 58 tests pass
  • ruff check src tests — clean
  • Example benchmark data at examples/benchmark/cluster_split_{pool,refs}.csv

- Add cluster_by_similarity() and cluster_split() to splits.py: greedy
  single-linkage clustering groups near-duplicate sequences so benchmark
  reference and test sets are never contaminated by each other
- Add find_contaminated_references() to evaluate.py: identifies reference
  sequences that are near-duplicates of test positives (Phase 2 leakage check)
- Add recall_at_k(), random_recall_at_k(), enrichment_factor(),
  benchmark_summary() to evaluate.py (full benchmark evaluation suite)
- Add test_cluster_split.py: 21 tests covering clustering properties,
  split partitioning, contamination detection, and end-to-end enrichment
  verification — all 3 positives rank above negatives without references,
  confirming scoring is feature-based not reference-proximity-based
- Add examples/benchmark/cluster_split_{pool,refs}.csv test data
- Fix ruff F401 unused imports in pipeline.py, test_cli.py, test_pipeline_filters.py
- 58 tests passing, ruff clean
@cschanhniem

Copy link
Copy Markdown
Collaborator Author

Superseded by PR #11 (feat/integrate-all-phases), which merges all Phase 2 + Phase 3 work into a single consolidation PR with 251 tests passing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant