Skip to content

Labels

Labels

  • agent

    Agent-in-the-loop evaluation
  • architecture

    Structure of the benchmark
  • bug

    Something isn't working
  • dataset

    Task set, fixtures and subsets
  • docs

    Documentation
  • documentation

    Improvements or additions to documentation
  • duplicate

    This issue or pull request already exists
  • enhancement

    New feature or request
  • good first issue

    Good for newcomers
  • help wanted

    Extra attention is needed
  • invalid

    This doesn't seem right
  • L1

    Protocol and driver compatibility layer
  • L2

    Web platform semantics layer
  • L3

    Frozen interactive environments layer
  • L4

    End-to-end agent evaluation layer
  • legal

    Licensing and redistribution
  • question

    Further information is requested
  • research

    Open question, not yet scheduled
  • RL

    Reinforcement learning environments
  • scoring

    Scoring rules and denominators
  • wontfix

    This will not be worked on