Skip to content

在GRPO训练中后期,模型有时会生成无关内容至最大token数 #5

Description

@ZhuoranZhao

我发现在GRPO训练的中后期,在采样生成过程中,有时会生成许多无关的内容直到达到最大生成token的数量,这就是所谓的reward hacking吗?这种现象应该在奖励函数上继续进行调整吗?有没有什么办法可以解决呢?强化小白,期待大家的回复!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions