Repository navigation
Structuré: rewrite the prompt short, one example set per Apple FM language, and refuse an invented incompleteness (refs #587) - #588
Conversation
…all (refs #587) Decision 11's four bars, how each is measured, the fixture sets, the stop rule for the language ladder, and the fabrication check's specification, committed before the first model call of the round. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…refs #587) Decision 6: rule 7 stays, and an output reporting the speaker's recall failing is refused when the transcript carries no such sentence. Narrow on purpose: recall-failure phrasings only, in six languages, matched strictly on the output and leniently on the input so a spoken form rewritten into a written one is never refused. Its own guardrail slug, a contract field off by default, turned on for Structuré alone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ranslated languages (refs #587) device-structured-0918-fr.json: nine French Structuré dictations of 2026-09-18 to 09-20, raw verbatim, including the three accepted fabrications and the bullet-per-sentence output. device-structured-en.json: the five English dictations #585 measured coming back in French. translated-structured.json: six French dictations translated by the agent into English and the 14 other Apple FM languages, for output language only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…delity rounds (refs #587) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…(refs #587) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…n Danish (refs #587) Captures of every run, the console output, and the script that reads them against bars.md. C1 holds the output language in 15 of 16 Apple FM languages and translates the Danish T3 fixture into English 3 of 3, which is over decision 5's 10 % bar: the round stops, C1 is not landed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: getdictus/dictus-ios/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (97)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThis PR updates structured-mode prompts and guardrails, adds short-input Smart Mode handling, records skip telemetry, expands the fidelity harness, and adds structured-rewrite fixtures, benchmark captures, and research reports. ChangesStructured rewrite and polish pipeline
Estimated code review effort: 5 (Critical) | ~90 minutes Merge Risk: ⚪ Minimal · up to This change shortens the Structuré prompt, adds localized examples, and adds a guard against invented incompleteness. It also accepts close spellings of names, cleans trailing layout marks, and skips short inputs. No concrete defect remains after review. The identical before and after grounding captures are the expected result, because the new spelling rule changes no verdict on that corpus. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 75.24% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 105 functions across 23 files. (65 skipped: 65 unsupported.)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Corroborating evidence from #571 (
|
…e (refs #587) Decision 5 step 2's seam, mode-neutral so #571 can carry the same patch. SmartModePrompt gains an optional table of instructions keyed by NLLanguage code, looked up by exact code then base subtag, falling back to the one prompt. PolishJob carries the transcript language: the forced transcription language if set, else the detected one. The pipeline resolves the task once before the engine, and the engine keys its session cache on the prompt text, so a session warmed with the fallback is never reused for another language's prompt. Prewarm resolves with the forced language when there is one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…587) A fixture's lang stands for the forced transcription language on the per-language path; the auto path passes the detected one, as PolishService does, so a Smart Mode's localized examples are exercised by the harness. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
C1 closed 18 of 354 Mac outputs on a lone ``` or --- line, all accepted. The post-pass now removes trailing lines made only of those tokens, before the guardrails; a line carrying any other character is left alone, so no word can be removed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…he transcript's language (refs #587) C2 keeps C1's rules byte for byte and translates its two worked examples into each of the 16 Apple FM language codes; its fallback is C1 exactly. The sets live in SmartModeStructuredExamples and the arm files are dumped from the Swift composition, so a landed C2 is the benched C2. Not wired into the catalogue yet. The fidelity round accepts a directory arm and resolves it per fixture with the app's own lookup rule. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The tables store Norwegian under nb, the code NLLanguageRecognizer answers; a transcript labelled with the macrolanguage code no fell back silently. Found by CodeRabbit on PR #589. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…0 outputs: fails, not landed (refs #587) Round 2, bars.md §10. C2 returns no wrong-language output in any of the 16 Apple FM languages and holds B2 and B4, but fails B1b as the bar reads it in Danish (2 of 18 refused on language, both a correct Danish sentence the recogniser reads as Norwegian at 0.99) and fails B3a (11 of 290 prose outputs carry a list line, against C1's 1 of 288). One accepted Norwegian output closed on the Norwegian example's own line. The round stops; step 3 is not attempted and C2 is not wired. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ature/587-smart-mode-prompt-campaign
…s example order, more trailing artefacts (refs #587) The language check now treats da/nb/no/sv as answering for each other, which is what C2's Danish failure was: a correct Danish sentence read as Bokmål at 0.993. The candidate C3 is C2 with the enumeration example first and the prose example last, on #571's measurement that the last example is the one the model copies. A lone dash and an ellipsis join the trailing-artefact list. Bars, fixtures and the 10 % threshold are unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ge rule, 710 outputs (refs #587) Round 3, bars.md §11. C3 holds B1 in every one of the 16 Apple FM languages (0 refused on language anywhere, 0 wrong-language accepted after the hand read), B2 (0 fabricated incompleteness, rule 7 kept 3/3) and B4. It still fails B3a: 5 of 290 prose outputs carry a list line, against 11 for C2 and 1 for C1, and all 5 append a list that restates the paragraph above it. The round stops; C3 is not landed and step 3 is not attempted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every set carries the same English rules, the enumeration example first and the prose example last, no sentence about the speaker's memory, no person named, and stays under the declared size. The last assertion records that the sets are benched and not wired. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2 966 characters against 5 556, seven one-line rules, two worked examples per Apple FM language, no counter-example block. Benched over three rounds and 2 130 Mac outputs: 0 % refused on check=language in all 16 languages against up to 29 %, English 33/33, bullets on prose 5/290 against 72/289, fabricated incompleteness 0/356 against 15/354, fidelity not worse on #583's axes. The largest dictation that fits rises from 3 972 to about 4 900 characters of speech. The one bar it does not hold is recorded rather than chased: 5 of 290 prose outputs append a list restating the paragraph above, on agent-translated Danish, Dutch and Swedish fixtures. No guardrail is added for it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
) Device round of 2026-09-23: the speaker said Hermes, Parakeet wrote airmes, the model wrote Airmesh, and the grounding check refused a faithful four-item English list. A name is now also grounded by an input word, or two adjacent input words joined, within an insertion-deletion distance of 1 for six letters and 2 from seven; five letters or fewer stay exact. A substituted letter is never accepted: a Damerau-Levenshtein draft grounded Sophie against Sophia, #414's W2 fabrication. Replayed on the #414 corpus, the final rule is identical to the old one anchor by anchor (grounding caught 7/11, 0/328 false rejections, before and after). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two of the three meaning deviations of the 2026-09-23 device round were on short input: a 28-character question moved into the first person, and a 90-character line gained a leading Voici. Below its floor the mode is now skipped, not refused: the dictation takes the path it takes with nothing armed (Normal polish with the toggle on, raw text otherwise), and nothing is announced. The floor is a SmartMode field, set on Structuré only, read from the 51 device dictations of 2026-09-13 to 09-23: every one under 120 characters is a single sentence, and none under 200 has more than two. The skip is logged as smartModeSkipped reason=shortInput and carried on the Normal polish's metrics event, with a per-mode count in the debug export. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…587) The previous commit said no device dictation under 200 characters has more than two sentences. One does: a 149-character dictation counting off three tasks, which is also the floor's known cost. 15 of the 16 are one or two. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t (refs #587) Two gaps the 2026-09-23 device round left. The user is told. A skip travels on the same SmartModeFailure channel as every other did-not-run sentence, with the text inserted, so the keyboard raises a notice in the #580 slot: "Structuré: dictation too short, text inserted as dictated." / "Structuré : dictée trop courte, texte inséré tel quel." It names its cause, unlike the refusal sentences, because this one is knowable and dictating more is what changes it. The skip is its own outcome, smartModeSkippedShortInput, carrying the mode, the character count and the floor. It is emitted before the delegated call, which is the fix: with the polish toggle OFF that call returns the raw text without writing anything, so the export showed the dictation as no event at all. It is now recorded with polish on and off. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The polish-event detail screen filled its Engine output section with "(engine did not run successfully)" on a smartModeSkippedShortInput event. Nothing failed there: the engine was never called, by design. The skip now gets its own line, with the character count and the floor, because the reader of a debug export is usually an agent triaging a report (#255). Debug UI only, off the dictation path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ing tolerance covers the case Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
refs #587 (PR 1 of decision 8). Also carries evidence for #570, #581, #585 and #523's 2026-09-20 verdict.
What this was for
Structuréwas failing Pierre in three ways at once, and #587 traced all three to the prompt's worked examples rather than its rules: an English dictation came back in French every time (#585, 4 of 4 on device); outputs closed on a sentence the speaker never said,Il y avait un autre truc, mais ça m'échappe., copied from the prompt's own example (#581, 7 device outputs, 3 inserted); and six sentences of prose came back as six bullets (#523, 2026-09-20).This PR rewrites the prompt short, gives it one example set per Apple FM language, adds a narrow check against the invented incompleteness sentence, and lands it on three benched rounds — 2 130 scored Mac outputs, bars committed before each round's first call.
To test by hand — after round 5
Structuréarmed (e.g. "Comment tu vas aujourd'hui ?"). Expected: the toolbar saysStructuré : dictée trop courte, texte inséré tel quel., the text goes in unchanged by Structuré (Normal polish, or raw with the toggle off), and the polish export carries an event withoutcome: smartModeSkippedShortInputandsmartModeLengthSkip: {mode: structured, characters: 28, floor: 200}. Worth doing twice, once with the polish toggle off: that is the case the previous build recorded nowhere.ac1c689.To test by hand — after round 4
These three checks come first, on the head commit of this PR:
check=groundingin the export.Structuréarmed. Expected: noStructurérewrite, so novais-jeand noVoici. You get Normal polish (or your raw words if polish is off), and no notice. The debug log showssmartModeSkipped mode=structured reason=shortInput chars=28 floor=200, and the export event carriessmartModeSkippedForLength: structured.08f7ba0, meaning French, paragraphs, first person kept, nothing added, no bullets on prose.Decision 12's device round (done on
08f7ba0, reference)iPhone on iOS 27, this build installed,
Structurépinned, keyboardfr, transcription on auto-detect. 8 dictations:```,---,-or...).Bars: 0 language error, 0 fabricated sentence, 0 bullet on prose. Flag every meaning distortion you see — a dropped negation, a changed name, a lost clause. Per decision 10, anything you flag opens #570's guardrail grilling, and zero takes #570 off the 2.0.0 gate.
Three extra checks, cheap while you are in there:
Structuré : n'a pas pu s'appliquer, texte inséré tel quel.and the export showsguardrailCheck=incompleteness.ListeandMessageonce each. Expected: unchanged — this PR moves nothing in them (that is PR 2, Smart Mode prompt campaign: rewrite Structuré short, make the output language hold, stop examples leaking #587 decision 9).Round 5 — after Pierre's device round on
ac1c689(2026-09-23)Both round-4 changes held on device: the English enumeration is accepted with its list (
air messkept as dictated, so the anchor is grounded), and a long French dictation is clean. Two gaps closed:1. The skip now shows a notice. A user who armed
Structuréand receives plain text cannot tell a deliberate skip from a broken mode. The skip travels on the sameSmartModeFailurechannel as every other did-not-run sentence, with the text inserted — which is whatisDegradedalready means — so the keyboard raises a notice in #580's slot and register (notice, not error; no red; #313's decisions):Structuré: dictation too short, text inserted as dictated./Structuré : dictée trop courte, texte inséré tel quel.Both locales are inDictusKeyboard/Localizable.xcstrings. It is the one did-not-run sentence that names its cause, because this one is knowable and repeatable and dictating more is what changes it.2. The skip was missing from the export, and
polishEnabled: falsewas indeed the reason. The skip delegated to the same method with no mode armed; with the toggle off, that call returns the raw text before writing any event, so nothing was recorded. (skippedShort: 0in Pierre's export is a different gate: that one is the recording-duration gate on the free polish.) The skip is now its own outcome,smartModeSkippedShortInput, emitted before the delegated call, carryingsmartModeLengthSkip: {mode, characters, floor}— so it is recorded with polish on and off, seeded in the export'soutcomes, split per mode insmartModesSkippedForLength, and visible in the in-app debug view. With the toggle on, a skipped dictation writes two events: the skip, then the Normal polish that ran in its place.After the device validation of
b30a954, one cosmetic fix: the polish-event detail screen filled itsEngine outputsection with(engine did not run successfully)on a skip. Nothing failed there, so it now reads(mode skipped: dictation too short, 46 chars, floor 200). Debug UI only, off the dictation path.One thing from that round that is not mine to fix here: the long French output opened with
Voici, a word the speaker never said (the 0.81-ratio dictation). It is not a length problem; it is being recorded on #570.Round 4 — after Pierre's device round on
08f7ba0(2026-09-23)Nine
Structurédictations on iPhone16,2, iOS 27.0: 0 language errors (English 3 of 3 in English), 0 fabricated incompleteness, 0 bullets on prose. A genuine four-item English enumeration correctly became a list, and was then refused by our own grounding check. Three meaning deviations were flagged, two of them on short dictations. Two changes answer that, both approved by Pierre:1. A name the model respelled from the speaker's word is grounded. The speaker said Hermes, Parakeet wrote
airmes, the model wroteAirmesh, andNLTaggertagged it as a personal name. The exact match then refused the whole list (check=grounding). The rule now: a name is grounded by an input word, or two adjacent input words joined, within an insertion-deletion distance of 1 for six-letter names and 2 from seven letters up. Five letters or fewer stay exact, and a changed letter is never accepted.SophieagainstSophia, which is A Smart Mode prompt's worked example was copied verbatim into the user's output #414'sW2-nom-prefixe, a labelled fabrication. It was still refused, but only because the same output also inventedMarion.PolishGroundingTests(Marc≠Marco,Mülle≠Müller). The floor moved to six rather than the pins.docs/research/587-structured-rewrite/round4/.TypeLessagainsttype less) is now grounded too, through the joined pair.2.
Structuréskips a dictation shorter than 200 characters. Below that, the mode does not call the model: the dictation takes the path it takes with nothing armed (Normal polish with the toggle on, raw text otherwise), with no notice. It is logged assmartModeSkipped mode=structured reason=shortInput chars=<n> floor=200 disarmed=false, and carried on the Normal event assmartModeSkippedForLength, with a per-mode count in the debug export.Comment tu vas→Comment vais-je),Voiciadded at 90, clauses swapped at 65,s'il te plaîtdropped at 192, tense and register lifted at 149.Structuréonly. It is aSmartModefield (minimumInputCharacters),nilon every other mode, includingMessageandRésumé.What changed
The prompt, landed. 2 966 characters against 5 556: seven one-line rules, two worked examples, no counter-example block, rule 1 the language, rule 6 the list, rule 7 conditioned on the transcript and quoting no phrasing. The rules are one English text; only the two examples are translated, one set per Apple FM language (
SmartModeStructuredExamples, 16 sets, agent-translated). The enumeration example comes first and the prose example last, which is a measurement rather than a preference (below). The largest dictation that fits the context rises from 3 972 to about 4 900 characters of speech.Three checks around it:
PolishIncompleteness(decision 6): refuses an output that reports the speaker's recall failing when the transcript carries none. Recall-failure phrasings only, in FR, EN, ES, DE, IT, PT; strict on the output, lenient on the transcript, so a spoken form rewritten into a written one is never refused. Its ownincompletenessslug; a contract field off by default and on forStructuréalone.PolishPostpass.stripTrailingFenceLines: drops trailing lines made only of```,---,***,-,...or…, before the guardrails. It cannot remove a word.da,nb,noandsvas one family. Measured cause:Jeg tror, jeg kan eksportere loggene, som de er., copied verbatim from a Danish transcript, reads asnbat 0.993, and the per-segment check (PolishGuardrail's language check is blind to bilingual output, which is the one shape Notes produces #413) refused the whole output twice. It is a whitelist of four codes, not a notion of "similar languages": Spanish/Portuguese and Simplified/Traditional stay refused, and an English chat reply on a Danish dictation stays refused.Also in the PR: the per-language seam is in
developvia #589 and only consumed here; the harness passes the transcript language, accepts a directory arm and records output language, list lines and the check's verdict; new fixtures (9 French and 5 English device dictations, plus 6 French dictations translated into 15 languages, valid for output language only, not fidelity); 40 new unit tests; anddocs/research/587-structured-rewrite/withbars.md, the three arms, every capture andsummarise.py, which reproduces every number below.The three rounds
check=languageWhy the example order moved: with the list example last, bullets on prose went from 1 of 288 (untranslated examples) to 11 of 290; #571 measured the same position effect on the output language. Moving it off the end took it to 5 of 290.
Vas-y let's go.×3 (your own 15 characters) and three correct Danish outputs the new family rule now accepts5-ramblingThe example's own content reached no output in round 3, in either arm: the Norwegian leak of round 2 did not recur after the reorder.
PolishIncompletenesswould not have caught it — that sentence was about a hotel booking, not about the speaker's memory — and the check is blind to that by design.The one bar that fails, unclosed and recorded here
5 of 290 prose outputs append a list that restates the paragraph above it, two of them accepted. All five are on agent-translated fixtures, in Danish (3), Dutch (1) and Swedish (1) — synthetic text in languages nobody on this project reads, measured on a Mac that already understates the phone. It is added content, not a formatting preference, and no guardrail sees it: the lines reuse the speaker's own words, so
segmentOverlappasses them.It is not fixed and not guarded in this PR, deliberately. The numbers are in
docs/research/587-structured-rewrite/round3/(summarise.py round3, section B3), and #587 carries a comment pointing at them. In French and English — the languages of the device round, and Pierre's — C3 produces no bullet on prose at all, which is why the decision was to land rather than spend a fourth round against a proxy.What contradicted the issue, and what I am not comfortable with
Round 4:
MartineforMartin). That is meaning damage (Structured rewrites past the guardrails: a product name replaced and a negation dropped #570's class), not an invention, and it is the trade the device evidence asked for. Changed letters and names of five letters or fewer are untouched.PolishServicehas no unit test. The service builds the real Apple FM engine and reads the App Group, so a test would call the model. The predicate, the floor, its decoding, the outcome, its detail, the degraded-outcome shape the notice keys on and the log line are all tested; the branch that assembles them is what check 1 above verifies on device.outcomescounts one dictation twice across two keys.Rounds 1-3:
Verification at the head commit:
swift test2142 tests, 0 failures (1 skipped).swiftlint lint --strict0 violations in 293 files.xcodebuild buildDictusApp with the keyboard on the iOS Simulator: BUILD SUCCEEDED. Round 4's changes are device-validated (ac1c689). Round 5's notice and export event are not; that is check 1 at the top.🤖 Generated with Claude Code