Skip to content

Restore the remaining missing diacritics in the Portuguese locale - #327

Merged
nank1ro merged 1 commit into
mainfrom
fix/portuguese-diacritics-rest
Sep 18, 2026
Merged

nank1ro merged 1 commit into
mainfrom
fix/portuguese-diacritics-rest

Conversation

@nank1ro

@nank1ro nank1ro commented Sep 17, 2026

Copy link
Copy Markdown
Owner

Follow-up to #326, which only caught part of this.

That PR's detector looked for description prose of 120+ characters containing no diacritic at all. It therefore missed every section that had some accents alongside a few misspelled words, and it never looked at --instructions--, --answers-- or --solutions--.

A word-level detector finds the rest — 811 sections in 540 files:

word occurrences
codigo 456
variavel 135
saida 131
nao 124
condicao 97
instrucao 66

Repo-wide the accented spellings dominate — "código" 1459 to "codigo" 540, "função" 2037 to "funcao" 26 — so the unaccented text is a straggler, not a house style.

What changed

The missing accents, and nothing else. No rewording, no reordering, no punctuation.

The judgement calls are the ones that change meaning rather than spelling: e (and) vs é (is), a vs à (crase), esta vs está, tem vs têm. Each was resolved by reading the clause — several lines contain both readings in sequence, e.g. "time é menor que 12, e no bloco de código imprima".

Diff shape

  • 537 exercise files, 1064 insertions / 1064 deletions — line-for-line.
  • 19 _theory.md files, +49 / −49.
  • curriculum.json unchanged.

Checks

  • Round-trip proof, 811 of 811. Stripping every combining mark from the new text (and mapping ç→c) reproduces the old text character-for-character. That is what proves only accents moved and no prose was quietly rewritten.
  • Every changed exerciseType: 3 file still has its --solutions-- bullet character-identical to one of its --answers-- bullets — the validator matches those as exact strings.
  • --output--, --asserts--, --seed-- and the other executed sections: byte-identical. This mattered: it's equivalent defect hides inside print("E' ora del caffe'!"), which is byte-compared against --output--, so accenting blindly would have broken exercises.
  • Fenced code byte-identical everywhere except pt/python/dictionaries/2.md, where the change is a Portuguese prose comment (# obtem → # obtém). The code itself is untouched.
  • Detector re-run: three apparent hits remain and all are correct as written — "Codigo" is a string literal, indices/lastIndex are Kotlin API names, and especifica there is the verb (no accent), not the adjective.
  • validator: +275436: All tests passed!

@nank1ro
nank1ro merged commit 34b0ffd into main Sep 18, 2026
4 checks passed
@nank1ro
nank1ro deleted the fix/portuguese-diacritics-rest branch September 18, 2026 09:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant