English | δΈζ
OpenNovelty is an LLM-powered agentic pipeline for transparent, evidence-grounded, and verifiable scholarly novelty assessment.
- Website (public reports): https://www.opennovelty.org
- Technical report (arXiv): https://arxiv.org/abs/2601.01576
π If you find OpenNovelty useful, please give us a star on GitHub and an upvote on Hugging Face β your support helps us a lot!
Repository status: The codebase will be incrementally refactored over the next 2 months to improve quality; the current version is a demo.
Novelty is a key criterion in peer review, but manual evaluation is often constrained by time, subjectivity, and retrieval coverage. By grounding analysis in retrieval and evidence alignment, OpenNovelty provides traceable novelty assessments and helps reduce subjective bias.
Figure: The full OpenNovelty workflow β a four-phase pipeline from paper input to report output.
| Phase | Functionality | Key Input | Key Output | Time | Dependencies |
|---|---|---|---|---|---|
| I | Information Extraction | Paper PDF URL | phase1_extracted.json |
~1 min | LLM API |
| II | Literature Retrieval | Phase 1 outputs | citation_index.json |
~10 min | Wispaper API |
| III | Deep Analysis | Phase 2 outputs | phase3_complete_report.json |
~10 min | LLM API |
| IV | Report Generation | Phase 3 outputs | novelty_report.md/pdf |
~30 sec | weasyprint |
Phase I β Information Extraction
- Download the PDF and extract full text + metadata
- Use an LLM to extract one core task and 1β3 contribution claims
- Generate 3 retrieval query variants for each task/contribution
Phase II β Literature Retrieval
- Semantic retrieval of related papers (via WisPaper API [paper])
- Quality filtering (perfect-match, time filtering) and deduplication (
canonical_id+ title normalization) - Build a citation index (Core Task Top-50, Contribution Top-10)
β οΈ Note: Phase 2 depends on Wispaper API which is not yet publicly available. The API will be opened soon β please stay tuned for updates!
Phase III β Deep Analysis
- Build a related-work taxonomy and synthesize a survey
- Textual similarity detection (token-level fuzzy matching;
β οΈ experimental, recommended to skip) - Full-text comparative verification and novelty judgments (
can_refute/cannot_refute/unclear) - π‘ Recommended setting: set
SKIP_TEXTUAL_SIMILARITY=truein.envto skip similarity detection (this experimental module is being updated)
Phase IV β Report Generation
- Template-based rendering for Markdown/PDF reports (no LLM calls)
- Unified citation formatting, evidence snippets, and hierarchical structure
| Type | Requirement |
|---|---|
| OS | Linux (Ubuntu 20.04+) / macOS |
| Python | 3.8+ (recommended: 3.10+) |
| Memory | 8GB+ |
| Network | Access to OpenReview, Wispaper API, and an LLM API |
# Ubuntu/Debian system dependencies
sudo apt-get update && sudo apt-get install -y \
git curl wget \
libpango-1.0-0 libpangocairo-1.0-0 libgdk-pixbuf2.0-0 \
libffi-dev libcairo2 libcairo2-dev libgirepository1.0-dev
# Python dependencies
cd /path/to/OpenNovelty
python3 -m venv venv && source venv/bin/activate
pip install -r requirements.txtMain dependencies: requests, openai, openreview-py, pypdf, weasyprint, python-dotenv, tqdm
Create a .env file in the project root:
# ============ LLM API Configuration (Required) ============
export LLM_API_ENDPOINT="https://openrouter.ai/api/v1" # Example
export LLM_API_KEY="sk-xxxxxxxx"
export LLM_MODEL_NAME="anthropic/claude-sonnet-4.5" # Example
# ============ Wispaper API (Required for Phase 2) ============
# Token is saved to ~/.wispaper_tokens.json by default (see first-time setup below)
# ============ Phase 3 Configuration (Recommended) ============
export SKIP_TEXTUAL_SIMILARITY="true" # Skip similarity detection (under development)
# ============ Optional Configuration ============
export HTTP_PROXY="http://127.0.0.1:7893"
export HTTPS_PROXY="http://127.0.0.1:7893"
β οΈ Coming Soon: Wispaper API is not yet publicly available. The following configuration will be enabled once the API is opened.
Using https://openreview.net/pdf?id=ZgCCDwcGwn as an example:
# Phase 1 - Extraction (~1 min)
python scripts/run_phase1_batch.py \
--papers "https://openreview.net/pdf?id=ZgCCDwcGwn" \
--out-root output/demo \
--force-year 2026 \
2>&1 | tee logs/phase1.log
# Phase 2 - Retrieval (~10 min) β οΈ Requires Wispaper API (coming soon)
bash scripts/run_phase2_concurrent.sh \
openreview_ZgCCDwcGwn_20260118 \
--base-dir output/demo \
2>&1 | tee logs/phase2.log
# Phase 3 - Deep analysis (~10 min)
bash scripts/run_phase3_all.sh \
output/demo/openreview_ZgCCDwcGwn_20260118 \
2>&1 | tee logs/phase3.log
# Phase 4 - Report generation (~30 sec)
bash scripts/run_phase4.sh \
output/demo/openreview_ZgCCDwcGwn_20260118 \
2>&1 | tee logs/phase4.log
# View the result
cat output/demo/openreview_ZgCCDwcGwn_20260118/phase4/novelty_report.mdπ‘ Argument notes:
--paperspaper URL |--out-rootoutput directory |--force-yearforce year |--base-dirsearch directory || teesave logs
# Create a paper list
cat > papers.txt << EOF
https://openreview.net/pdf?id=PAPER_ID_1
https://openreview.net/pdf?id=PAPER_ID_2
https://openreview.net/pdf?id=PAPER_ID_3
EOF
# Phase 1: Batch extraction
python scripts/run_phase1_batch.py \
--paper-file papers.txt \
--out-root output/batch \
--force-year 2026
# Phase 2: Batch retrieval (auto-discover all papers) β οΈ Requires Wispaper API (coming soon)
bash scripts/run_phase2_concurrent.sh \
--base-dir output/batch \
--auto-discover \ # Auto-discover all papers with Phase 1 completed under base-dir
--max-workers 10 # Concurrency (default 10; adjust based on machine capacity)
# Phase 3+4: Batch analysis and report generation
bash scripts/run_phase3_phase4_serial_pending.sh output/batchπ‘ Argument notes:
--paper-file: a list file (one URL per line)--auto-discover: auto-scan all papers that need processing--max-workers: number of parallel workers (Phase 2 makes concurrent API calls)
| Command | Purpose |
|---|---|
python scripts/refresh_wispaper_token.py |
Refresh Wispaper token |
python scripts/run_phase1_batch.py --help |
Show Phase 1 help |
bash scripts/run_phase2_concurrent.sh --help |
Show Phase 2 help |
cat logs/phase2.log | grep ERROR |
Locate error logs |
| Script | Function | Use Case | Time |
|---|---|---|---|
run_phase1_batch.py |
Information extraction | Single / batch | ~1 min / paper |
run_phase2_concurrent.sh |
Literature retrieval |
Single / batch (coming soon) | ~10 min / paper |
run_phase2_only.py |
Retrieval (single paper) |
Coming soon | ~10 min / paper |
run_phase3_all.sh |
Deep analysis (7 sub-steps) | Single / batch | ~10 min / paper |
run_phase4.sh |
Report generation | Single paper | ~30 sec / paper |
run_phase3_phase4_serial_pending.sh |
Batch completion | Auto-discover papers with Phase 2 done | - |
refresh_wispaper_token.py |
Token refresh |
Coming soon | ~10 sec |
output/<run>/<paper_id>/
βββ phase1/ # Phase 1 outputs
β βββ phase1_extracted.json # β Core task and contributions (with query variants)
β βββ paper.json # Paper metadata
β βββ pub_date.json # Publication date
β βββ fulltext_raw.txt # Raw full text from PDF
β βββ fulltext_cleaned.txt # Cleaned full text
β βββ body_for_core_task.txt # Text relevant to the core task
β βββ body_for_claims.txt # Text relevant to contribution claims
β βββ raw_llm_responses/ # Raw LLM responses (reproducibility)
β
βββ phase2/final/ # Phase 2 outputs
β βββ citation_index.json # β Citation index (required for Phase 3)
β βββ core_task_perfect_top50.json # Core task candidates Top-50
β βββ contribution_*_perfect_top10.json # Contribution candidates Top-10
β βββ stats.json # Search statistics
β βββ raw_responses/ # Raw API responses
β
βββ phase3/ # Phase 3 outputs
β βββ phase3_complete_report.json # β Full analysis report (required for Phase 4)
β βββ core_task_survey/ # Taxonomy and survey synthesis
β βββ core_task_comparisons/ # Core-task comparisons
β βββ contribution_analysis/ # Contribution-level analysis
β βββ textual_similarity_detection/ # Textual similarity detection
β βββ cached_paper_texts/ # Cached full texts
β βββ raw_llm_responses/ # Raw LLM responses
β
βββ phase4/ # Phase 4 outputs
βββ novelty_report.md # β Markdown report
βββ novelty_report.pdf # PDF report (requires weasyprint)β Files marked are required inputs for the next phase.
| Issue | Symptom | Solution |
|---|---|---|
| Wispaper token expired | 401 Unauthorized / Token expired |
|
| PDF download failure | ConnectionError / Timeout |
Check network, URL validity, and proxy settings |
| LLM API call failure | API Error / Invalid key |
Check LLM_API_KEY, LLM_MODEL_NAME, and quota |
| PDF generation failure | weasyprint error |
sudo apt-get install libpango-1.0-0 libpangocairo-1.0-0 libgdk-pixbuf2.0-0 |
| Phase 3 appears stuck | No output for > 5 minutes | Usually normal; a single sub-step can take 2β5 minutes |
| Phase 1 year detection fails | Year extraction failed |
Use --force-year YYYY |
View detailed logs:
# Locate errors
cat logs/phase2.log | grep -i error
# Inspect LLM calls
cat logs/phase3.log | grep "LLM call"
# Check generated files
ls -lh output/demo/openreview_XXX/phase2/final/Re-run failed sub-steps (Phase 3):
# Re-run taxonomy generation only
python -m scripts.run_phase3_taxonomy \
--phase2-dir output/demo/openreview_XXX/phase2 \
--out-dir output/demo/openreview_XXX \
--log-level INFO
# Re-run comparison analysis only
python -m scripts.run_phase3_core_task_comparisons \
--phase1-dir output/demo/openreview_XXX/phase1 \
--phase2-dir output/demo/openreview_XXX/phase2 \
--out-dir output/demo/openreview_XXX \
--resume \
--log-level INFOVerify configuration:
# Check environment variables
echo "LLM_API_KEY: ${LLM_API_KEY:0:10}..."
echo "LLM_MODEL_NAME: $LLM_MODEL_NAME"
# Validate Wispaper token (β οΈ coming soon)
# python -c "from paper_novelty_pipeline.services.wispaper_client import WispaperClient; client = WispaperClient(); print('Token valid!')"
# Check Phase 1 output
cat output/demo/openreview_XXX/phase1/phase1_extracted.json | jq '.core_task'@article{zhang2026opennovelty,
title={OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment},
author={Zhang, Ming and Tan, Kexin and Huang, Yueyuan and Shen, Yujiong and Ma, Chunchun and Ju, Li and Zhang, Xinran and Wang, Yuhui and Jing, Wenqing and Deng, Jingyi and others},
journal={arXiv preprint arXiv:2601.01576},
year={2026}
}This project is released under the Apache License 2.0. See LICENSE.
All reports are generated based on retrieved literature and LLM outputs. Coverage is constrained by retrieval recall, and conclusions are provided for reference only and should not be treated as final novelty judgments. OpenNovelty aims to provide transparent, evidence-grounded assistance rather than replacing human review.
- Ming Zhang: mingzhang23@m.fudan.edu.cn
- Kexin Tan: kxtan18@fudan.edu.cn or bluellatan@gmail.com
