Skip to content

feat(distillation): implement OPSA in AReaL - #1768

Open
zahrayousefijamarani wants to merge 9 commits into
areal-project:mainfrom
zahrayousefijamarani:opsa
Open

zahrayousefijamarani wants to merge 9 commits into
areal-project:mainfrom
zahrayousefijamarani:opsa

Conversation

@zahrayousefijamarani

@zahrayousefijamarani zahrayousefijamarani commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Description

Implements OPSA (On-Policy Self-Adaptation) in AReaL for RL post-training of language models.

This PR integrates the OPSA training workflow with AReaL's existing rollout and training infrastructure. The implementation uses token-level policy information from on-policy generations to construct the OPSA training signal without requiring a separate teacher model.

Key components include:

OPSA-based RL training workflow
Integration with AReaL's rollout and policy optimization pipeline
Token-level policy information and entropy computation for OPSA
Self-adaptive advantage construction following the OPSA method
DAPO-Math-17k dataset preparation for training
AIME 2024 dataset preparation for evaluation
OPSA training configuration and example entry point
README documentation covering the method, dataset preparation, configuration, and usage

Type of Change

  • 🐛 Bug fix
  • ✨ New feature
  • 💥 Breaking change
  • 📝 Documentation update
  • ♻️ Refactoring
  • ⚡ Performance improvement
  • ✅ Test coverage improvement

Checklist

  • I have read the
    Contributing Guide
  • Pre-commit hooks pass (pre-commit run --all-files)
  • Relevant tests pass; new tests added for new functionality
  • Documentation updated (if applicable; built with ./docs/build_all.sh)
  • Branch is up to date with main
  • Self-reviewed via /review-pr command
  • This PR was created by a coding agent via /create-pr
  • This PR is a breaking change

Breaking Change Details (if applicable):

Additional Context


Need help? Check the
Contributing Guide
or ask in GitHub Discussions!

@sitabulaixizawaluduo

Copy link
Copy Markdown
Collaborator

Please run pre-commit before submitting

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants