Skip to content

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MACPO: Weak-to-Strong Alignment via Multi-Agent Contrastive Preference Optimization

framework

Abstract

As large language models (LLMs) are rapidly advancing and achieving near-human capabilities, aligning them with human values is becoming more urgent. In scenarios where LLMs outperform humans, we face a weak-to-strong alignment problem where we need to effectively align strong student LLMs through weak supervision generated by weak teachers. Existing alignment methods mainly focus on strong-to-weak alignment and self-alignment settings, and it is impractical to adapt them to the much harder weak-to-strong alignment setting. To fill this gap, we propose a multi-agent contrastive preference optimization (MACPO) framework. MACPO facilitates weak teachers and strong students to learn from each other by iteratively reinforcing unfamiliar positive behaviors while penalizing familiar negative ones. To get this, we devise a mutual positive behavior augmentation strategy to encourage weak teachers and strong students to learn from each other’s positive behavior and further provide higher quality positive behavior for the next iteration. Additionally, we propose a hard negative behavior construction strategy to induce weak teachers and strong students to generate familiar negative behavior by fine-tuning on negative behavioral data. Experimental results on the HH-RLHF and PKU-SafeRLHF datasets, evaluated using both automatic metrics and human judgments, demonstrate that MACPO simultaneously improves alignment performance of strong students and weak teachers. Moreover, as the number of weak teachers increases, MACPO achieves better weak-to-strong alignment performance through more iteration optimization rounds.

Requirements

  1. pip install -e .
  2. if you encounter errors about the environment, you can use pip install -r requirements.txt to fix it.

Please note training code is from open-source framework LLaMA-Factory-files.

Datasets

The data for weak teachers and strong students initialization and iterative optimization are placed in

data/...

Models

Download Llama2-7b-base, Mistral-7b-v0.1-base, Llama3-8b-base, Llama2-70b-base in the model folder

model/...

Training

Based on open-source training framework LLaMA-Factory, following below instructions for training.

Initialization

sh scripts/***/***/pos_initial.sh
sh scripts/***/***/neg_initial

Iterative training

sh scripts/***/***/pos_stage1.sh
sh scripts/***/***/pos_stage2.sh
sh scripts/***/***/pos_stage3.sh

Test_generation

sh scripts/***/***/test.sh

Citation

If this work is helpful to you, welcome to cite our paper as:

@inproceedings{lyu2025macpo,
	title        = {{MACPO}: Weak-to-Strong Alignment via Multi-Agent Contrastive Preference Optimization},
	author       = {Yougang Lyu and Lingyong Yan and Zihan Wang and Dawei Yin and Pengjie Ren and Maarten de Rijke and Zhaochun Ren},
	year         = 2025,
	booktitle    = {Proceedings of ICLR}
}

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages