Skip to content

How are QeRL-trained models evaluated on other benchmarks? #24

Description

@LJlimo

Hello, and thanks for your great work.

The repo only released evaluation scripts for the base model without QeRL training on MATH500, AIME24, and AIME25. I'm just wondering how to evaluate QeRL-trained models with LoRA checkpoint on these benchmarks.

Any advice would be greatly appreciated.
Thank you for your help.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions