Skip to content

Clarification on models and prompts used for dataset construction #2

Description

@Wonderdch

Hi authors, thank you for releasing WorldMemArena.

Could you clarify the automated dataset-construction pipeline, particularly for the synthetic Lifelong Evolution split?

The paper states that sessions are generated from hidden world states, gold memory points are extracted and refined, and QA checkpoints are constructed from those memory points. However, we could not find the following information in the paper, dataset card, or repository:

  1. Which model(s) were used to generate the Lifelong sessions?
  2. Which model(s) and prompts were used to extract/merge/revise gold memory points?
  3. Which model(s) and prompts were used to generate questions and gold answers?
  4. Were the automatic validators model-based or rule-based?
  5. What did the 2–3 human annotators review or modify? Are annotation guidelines, edit rates, or agreement statistics available?
  6. Are there plans to release the construction scripts, prompts, model versions, and generation parameters?

This provenance seems important for reproducing the benchmark and assessing possible generator–annotator–judge bias, especially because the Lifelong dialogues, reference memory points, and QA pairs may be produced within the same automated pipeline.

Thanks!

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions