Skip to content

Things to do before neurips #46

Description

@galv
  • Create two separate datasets to distribute, one CC-BY, one CC-BY-SA.
  • Rerun yamnet on the entire dataset. This means we need to make it more performant See yamnet WIP #40
  • Send data to be hand-transcribed.
    • Optionally, do audio-based deduplication first.
  • Add text deata deduplication to the data creation pipeline.
  • Train kaldi and/or nemo models on the dataset. Provide fixes to the dataset, based on this work.
    Adding more as time goes on...

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions