Skip to content
 
 

Repository files navigation

UPD: This repository is no longer maintained. We moved here: https://github.com/ai-alignment-liaison/ValueShiftProject under AI Alignment Liaison organisation. Please find the latest updates at the new address

This repository contains work done by the ValueShift Team for Moonshot Alignment Program hosted by AI Plans.

General Description

The main research track we worked on is Influence of Value Representations on LLM Outputs. The repository contains our results. It is work in progress, and we plan to continue our research further on.

What we did

We persued multiple distinct research directions:

  • Behavior Editing
  • Domain Dependence of the Refusal Vector Direction
  • Certainty Vector Direction
  • Low-Rank optimization methods for steering vector representations.

In short, we figured ot that

  • behavior editing with ICE, ROME, and other techniques is promosing, but requires a lot of computatiuonal resources
  • refusal vector direction depends on the field (checked for safety and policy)
  • it is possible to extract a single certainty vector, that predicts an LLM's confidence before giving an answer.

As the future work we plan to

  • test our findings on larger models and datasets
  • merge refusal and certainty direction to make an LLM-based AI agent "understand" its limits and avoid producing hallucinations
  • apply our approaches to AI alignment evaluation

Literature we used

About

Results of all our research directions forked from main repo

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages