I read your excellent paper! The paper mentions that phonemes are modeled to have three states, but I'm not exactly sure what that entails. Did you provide separate position embeddings for the beginning, middle, and end parts of the phoneme?
I read your excellent paper!
The paper mentions that phonemes are modeled to have three states, but I'm not exactly sure what that entails.
Did you provide separate position embeddings for the beginning, middle, and end parts of the phoneme?