The dataset isn't very big, but at least for Russian language it can be first trained on the regular dataset of comments from Pikabu(100 mil) and then fine-tuned with some coefficient on insults dataset. Theoretically, it can work very good. Depending on the results, we can consider hardcoding the first words for generated sentences, like "you are ...", "ты ...".
The dataset isn't very big, but at least for Russian language it can be first trained on the regular dataset of comments from Pikabu(100 mil) and then fine-tuned with some coefficient on insults dataset. Theoretically, it can work very good. Depending on the results, we can consider hardcoding the first words for generated sentences, like "you are ...", "ты ...".