Skip to content

Fix Vietnamese Telex tone/shape-key bugs, correct accent detection for non-Italian locales, seed next-word predictions - #297

Closed
Oxydox wants to merge 5 commits into
palsoftware:mainfrom
Oxydox:main
Closed

Oxydox wants to merge 5 commits into
palsoftware:mainfrom
Oxydox:main

Conversation

@Oxydox

@Oxydox Oxydox commented Sep 5, 2026

Copy link
Copy Markdown

This PR includes several fixes found and verified while using the Vietnamese Telex layout day-to-day, plus one improvement to the (locale-agnostic) suggestion engine.

Telex fixes (VietnameseTelexProcessor.kt):

Tone marks were landing on the wrong vowel in ao/eo diphthongs (e.g. "nao"+f → "naò" instead of "nào"), and in closed syllables like oan/oat (e.g. "toan"+s → "tóan" instead of "toán").
Non-adjacent letters were incorrectly merging (e.g. typing "d","e","s","t","r" would corrupt into đ duplicating/dropping letters).
Shape keys (a/e/o doubling) could incorrectly reach back through an already-closed syllable coda, corrupting an earlier vowel (e.g. "mam"+"a" → "maam" instead of staying literal).
Added a lightweight "is this even a plausible Vietnamese syllable" check (valid onset/coda validation) with a rollback mechanism, so common English words typed while Telex is active (e.g. "destroyed", "wolf", "straight") no longer get corrupted into nonsense.
Deliberate behavior change: a third press of d on đ now reverts fully to literal dd, matching how the other shape keys already escape (e.g. ô + o → oo). Previously it appended a literal d instead (đd). Confirmed against real device behavior (Gboard/UniKey) before changing. This changes what one existing test (shape keys convert base vowels) expected — updated accordingly.

All changes verified against the full existing test suite (zero regressions) plus a broad set of real Vietnamese words with varied onsets/codas.

Suggestion engine (SuggestionEngine.kt):

The "does this candidate have an accent" check was hardcoded to a small set of Italian-only accented characters, so it silently never fired for the vast majority of Vietnamese diacritics (missed ô, ơ, ư, â, ă, đ, and any combined tone+shape mark). Replaced with the general NFD-based accent-stripping check already used elsewhere in the same file, so it works correctly for any language, not just Italian.

Next-word prediction (UserNGramStore.kt):

Previously started with zero data for every new install — no shipped bigram model at all. Added a small seed set of ~130 common Vietnamese word-pairs, inserted at a low baseline count so real usage naturally overtakes it over time.

fix(telex): tone placement on ao/eo clusters, d-adjacency, closed-syllable vowel reach-through
Refactor VietnameseTelexProcessor to handle non-Vietnamese syllables and improve tone application logic.
Added a method to pre-populate Vietnamese bigrams for next-word suggestions, ensuring new installs have initial predictions. This method seeds common word pairs with a low baseline count to encourage learning from user input.
@Oxydox Oxydox closed this Sep 6, 2026
@pzauner

pzauner commented Sep 6, 2026

Copy link
Copy Markdown
Collaborator

Why closed again?

@Oxydox

Oxydox commented Sep 6, 2026 •

Copy link
Copy Markdown
Author

I don't know why the pull request was showing blank for me. I made some updates and just created another one. I am new to this, sorry for the trouble.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants