Description
Hello! Thanks for the great work on Internalize_CoT. I am trying to reproduce your results but I'm a bit confused about the gsm8k-aug dataset.
The Problem
There are currently multiple versions on Hugging Face (zen-E/GSM8k-Aug and whynlp/gsm8k-aug). Since your paper mentions extending the dataset to 385k samples with specific formatting, I want to ensure I'm using the exact same version you used for training.
My Questions
Could you please provide the specific link to the gsm8k-aug dataset you used?Is there any specific preprocessing (like removing NL explanations) that I should be aware of if downloading from a public source?
Thank you for your time!
Description
Hello! Thanks for the great work on Internalize_CoT. I am trying to reproduce your results but I'm a bit confused about the gsm8k-aug dataset.
The Problem
There are currently multiple versions on Hugging Face (zen-E/GSM8k-Aug and whynlp/gsm8k-aug). Since your paper mentions extending the dataset to 385k samples with specific formatting, I want to ensure I'm using the exact same version you used for training.
My Questions
Could you please provide the specific link to the gsm8k-aug dataset you used?Is there any specific preprocessing (like removing NL explanations) that I should be aware of if downloading from a public source?
Thank you for your time!