Conversation
fix(telex): tone placement on ao/eo clusters, d-adjacency, closed-syllable vowel reach-through
Refactor VietnameseTelexProcessor to handle non-Vietnamese syllables and improve tone application logic.
Added a method to pre-populate Vietnamese bigrams for next-word suggestions, ensuring new installs have initial predictions. This method seeds common word pairs with a low baseline count to encourage learning from user input.
Collaborator
|
Why closed again? |
Author
|
I don't know why the pull request was showing blank for me. I made some updates and just created another one. I am new to this, sorry for the trouble. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR includes several fixes found and verified while using the Vietnamese Telex layout day-to-day, plus one improvement to the (locale-agnostic) suggestion engine.
Telex fixes (VietnameseTelexProcessor.kt):
Tone marks were landing on the wrong vowel in ao/eo diphthongs (e.g. "nao"+f → "naò" instead of "nào"), and in closed syllables like oan/oat (e.g. "toan"+s → "tóan" instead of "toán").
Non-adjacent letters were incorrectly merging (e.g. typing "d","e","s","t","r" would corrupt into đ duplicating/dropping letters).
Shape keys (a/e/o doubling) could incorrectly reach back through an already-closed syllable coda, corrupting an earlier vowel (e.g. "mam"+"a" → "maam" instead of staying literal).
Added a lightweight "is this even a plausible Vietnamese syllable" check (valid onset/coda validation) with a rollback mechanism, so common English words typed while Telex is active (e.g. "destroyed", "wolf", "straight") no longer get corrupted into nonsense.
Deliberate behavior change: a third press of d on đ now reverts fully to literal dd, matching how the other shape keys already escape (e.g. ô + o → oo). Previously it appended a literal d instead (đd). Confirmed against real device behavior (Gboard/UniKey) before changing. This changes what one existing test (shape keys convert base vowels) expected — updated accordingly.
All changes verified against the full existing test suite (zero regressions) plus a broad set of real Vietnamese words with varied onsets/codas.
Suggestion engine (SuggestionEngine.kt):
The "does this candidate have an accent" check was hardcoded to a small set of Italian-only accented characters, so it silently never fired for the vast majority of Vietnamese diacritics (missed ô, ơ, ư, â, ă, đ, and any combined tone+shape mark). Replaced with the general NFD-based accent-stripping check already used elsewhere in the same file, so it works correctly for any language, not just Italian.
Next-word prediction (UserNGramStore.kt):
Previously started with zero data for every new install — no shipped bigram model at all. Added a small seed set of ~130 common Vietnamese word-pairs, inserted at a low baseline count so real usage naturally overtakes it over time.