Skip to content

Structuré: rewrite the prompt short, one example set per Apple FM language, and refuse an invented incompleteness (refs #587) - #588

Merged
Pivii merged 23 commits into
developfrom
feature/587-smart-mode-prompt-campaign
Sep 23, 2026
Merged

Pivii merged 23 commits into
developfrom
feature/587-smart-mode-prompt-campaign

Conversation

@Pivii

@Pivii Pivii commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

refs #587 (PR 1 of decision 8). Also carries evidence for #570, #581, #585 and #523's 2026-09-20 verdict.

What this was for

Structuré was failing Pierre in three ways at once, and #587 traced all three to the prompt's worked examples rather than its rules: an English dictation came back in French every time (#585, 4 of 4 on device); outputs closed on a sentence the speaker never said, Il y avait un autre truc, mais ça m'échappe., copied from the prompt's own example (#581, 7 device outputs, 3 inserted); and six sentences of prose came back as six bullets (#523, 2026-09-20).

This PR rewrites the prompt short, gives it one example set per Apple FM language, adds a narrow check against the invented incompleteness sentence, and lands it on three benched rounds — 2 130 scored Mac outputs, bars committed before each round's first call.

To test by hand — after round 5

  1. Dictate a short sentence with Structuré armed (e.g. "Comment tu vas aujourd'hui ?"). Expected: the toolbar says Structuré : dictée trop courte, texte inséré tel quel., the text goes in unchanged by Structuré (Normal polish, or raw with the toggle off), and the polish export carries an event with outcome: smartModeSkippedShortInput and smartModeLengthSkip: {mode: structured, characters: 28, floor: 200}. Worth doing twice, once with the polish toggle off: that is the case the previous build recorded nowhere.
  2. Redo the English enumeration dictation. Expected: accepted with its list, as on ac1c689.
  3. One long French dictation, as a regression check.

To test by hand — after round 4

These three checks come first, on the head commit of this PR:

  1. Redo the English enumeration dictation (the one naming the masterclass, T3 Code, the Hermes agent on the VPS, and the HTML presentation). Expected: accepted, with its list. No toolbar notice, and no check=grounding in the export.
  2. One single-line dictation (e.g. "Comment tu vas aujourd'hui ?") with Structuré armed. Expected: no Structuré rewrite, so no vais-je and no Voici. You get Normal polish (or your raw words if polish is off), and no notice. The debug log shows smartModeSkipped mode=structured reason=shortInput chars=28 floor=200, and the export event carries smartModeSkippedForLength: structured.
  3. One long French dictation (1 minute or more), as a regression check. Expected: as on 08f7ba0, meaning French, paragraphs, first person kept, nothing added, no bullets on prose.

Decision 12's device round (done on 08f7ba0, reference)

iPhone on iOS 27, this build installed, Structuré pinned, keyboard fr, transcription on auto-detect. 8 dictations:

  1. Two long French dictations (1 minute or more, rambling, one change of topic). Expected: French, real paragraphs, your first person kept, nothing you did not say, and no bullets unless you counted items off out loud.
  2. Two English dictations (~20 s each). Expected: both come back in English and are accepted — not the raw text with the toolbar notice. This is Smart Modes rewrite an English dictation in French: the worked examples, not the rules, set the output language #585's bar.
  3. Two short dictations (1 or 2 sentences). Expected: returned nearly as said, nothing appended, and no stray last line (```, ---, - or ...).
  4. Two prose dictations that enumerate nothing (6 sentences on one topic, like your 2026-09-20 recording-conditions test). Expected: 0 bullets, paragraphs only.

Bars: 0 language error, 0 fabricated sentence, 0 bullet on prose. Flag every meaning distortion you see — a dropped negation, a changed name, a lost clause. Per decision 10, anything you flag opens #570's guardrail grilling, and zero takes #570 off the 2.0.0 gate.

Three extra checks, cheap while you are in there:

  1. Dictate a French message that really ends on "ah et il y avait un dernier truc, ça m'échappe". Expected: accepted, sentence kept (rule 7).
  2. Dictate an ordinary French message with nothing about forgetting, several times. Expected: no output ever ends on a sentence about your memory. If one does, the field gets your raw words with Structuré : n'a pas pu s'appliquer, texte inséré tel quel. and the export shows guardrailCheck=incompleteness.
  3. Arm Liste and Message once each. Expected: unchanged — this PR moves nothing in them (that is PR 2, Smart Mode prompt campaign: rewrite Structuré short, make the output language hold, stop examples leaking #587 decision 9).

Round 5 — after Pierre's device round on ac1c689 (2026-09-23)

Both round-4 changes held on device: the English enumeration is accepted with its list (air mess kept as dictated, so the anchor is grounded), and a long French dictation is clean. Two gaps closed:

1. The skip now shows a notice. A user who armed Structuré and receives plain text cannot tell a deliberate skip from a broken mode. The skip travels on the same SmartModeFailure channel as every other did-not-run sentence, with the text inserted — which is what isDegraded already means — so the keyboard raises a notice in #580's slot and register (notice, not error; no red; #313's decisions): Structuré: dictation too short, text inserted as dictated. / Structuré : dictée trop courte, texte inséré tel quel. Both locales are in DictusKeyboard/Localizable.xcstrings. It is the one did-not-run sentence that names its cause, because this one is knowable and repeatable and dictating more is what changes it.

2. The skip was missing from the export, and polishEnabled: false was indeed the reason. The skip delegated to the same method with no mode armed; with the toggle off, that call returns the raw text before writing any event, so nothing was recorded. (skippedShort: 0 in Pierre's export is a different gate: that one is the recording-duration gate on the free polish.) The skip is now its own outcome, smartModeSkippedShortInput, emitted before the delegated call, carrying smartModeLengthSkip: {mode, characters, floor} — so it is recorded with polish on and off, seeded in the export's outcomes, split per mode in smartModesSkippedForLength, and visible in the in-app debug view. With the toggle on, a skipped dictation writes two events: the skip, then the Normal polish that ran in its place.

After the device validation of b30a954, one cosmetic fix: the polish-event detail screen filled its Engine output section with (engine did not run successfully) on a skip. Nothing failed there, so it now reads (mode skipped: dictation too short, 46 chars, floor 200). Debug UI only, off the dictation path.

One thing from that round that is not mine to fix here: the long French output opened with Voici, a word the speaker never said (the 0.81-ratio dictation). It is not a length problem; it is being recorded on #570.

Round 4 — after Pierre's device round on 08f7ba0 (2026-09-23)

Nine Structuré dictations on iPhone16,2, iOS 27.0: 0 language errors (English 3 of 3 in English), 0 fabricated incompleteness, 0 bullets on prose. A genuine four-item English enumeration correctly became a list, and was then refused by our own grounding check. Three meaning deviations were flagged, two of them on short dictations. Two changes answer that, both approved by Pierre:

1. A name the model respelled from the speaker's word is grounded. The speaker said Hermes, Parakeet wrote airmes, the model wrote Airmesh, and NLTagger tagged it as a personal name. The exact match then refused the whole list (check=grounding). The rule now: a name is grounded by an input word, or two adjacent input words joined, within an insertion-deletion distance of 1 for six-letter names and 2 from seven letters up. Five letters or fewer stay exact, and a changed letter is never accepted.

2. Structuré skips a dictation shorter than 200 characters. Below that, the mode does not call the model: the dictation takes the path it takes with nothing armed (Normal polish with the toggle on, raw text otherwise), with no notice. It is logged as smartModeSkipped mode=structured reason=shortInput chars=<n> floor=200 disarmed=false, and carried on the Normal event as smartModeSkippedForLength, with a per-mode count in the debug export.

  • The floor is read from the 51 device dictations of 2026-09-13 to 09-23: all 8 under 120 characters are one sentence, and 15 of the 16 under 200 are one or two sentences.
  • Deviations under 200 characters: the person swapped at 28 (Comment tu vas → Comment vais-je), Voici added at 90, clauses swapped at 65, s'il te plaît dropped at 192, tense and register lifted at 149.
  • The known cost: a 149-character dictation counting off three tasks now gets Normal polish instead of a list.
  • It applies to Structuré only. It is a SmartMode field (minimumInputCharacters), nil on every other mode, including Message and Résumé.

What changed

The prompt, landed. 2 966 characters against 5 556: seven one-line rules, two worked examples, no counter-example block, rule 1 the language, rule 6 the list, rule 7 conditioned on the transcript and quoting no phrasing. The rules are one English text; only the two examples are translated, one set per Apple FM language (SmartModeStructuredExamples, 16 sets, agent-translated). The enumeration example comes first and the prose example last, which is a measurement rather than a preference (below). The largest dictation that fits the context rises from 3 972 to about 4 900 characters of speech.

Three checks around it:

  • PolishIncompleteness (decision 6): refuses an output that reports the speaker's recall failing when the transcript carries none. Recall-failure phrasings only, in FR, EN, ES, DE, IT, PT; strict on the output, lenient on the transcript, so a spoken form rewritten into a written one is never refused. Its own incompleteness slug; a contract field off by default and on for Structuré alone.
  • PolishPostpass.stripTrailingFenceLines: drops trailing lines made only of ```, ---, ***, -, ... or …, before the guardrails. It cannot remove a word.
  • The language check treats da, nb, no and sv as one family. Measured cause: Jeg tror, jeg kan eksportere loggene, som de er., copied verbatim from a Danish transcript, reads as nb at 0.993, and the per-segment check (PolishGuardrail's language check is blind to bilingual output, which is the one shape Notes produces #413) refused the whole output twice. It is a whitelist of four codes, not a notion of "similar languages": Spanish/Portuguese and Simplified/Traditional stay refused, and an English chat reply on a Danish dictation stays refused.

Also in the PR: the per-language seam is in develop via #589 and only consumed here; the harness passes the transcript language, accepts a directory arm and records output language, list lines and the check's verdict; new fixtures (9 French and 5 English device dictations, plus 6 French dictations translated into 15 languages, valid for output language only, not fidelity); 40 new unit tests; and docs/research/587-structured-rewrite/ with bars.md, the three arms, every capture and summarise.py, which reproduces every number below.

The three rounds

Arm Worst language, refused on check=language English Bullets on prose Fabricated incompleteness
The prompt in production 29 % (nl), 24 % (en), 17 % (da, de) 22/33 English 72 / 289 15 / 354
C1 — one FR + one EN example 28 % (da), 3 genuine translations 33/33 1 / 288 0
C2 — examples translated, list example last 11 % (da), 0 genuine 33/33 11 / 290 0
C3 — landed 0 % in all 16 languages 33/33 5 / 290 0 / 356

Why the example order moved: with the list example last, bullets on prose went from 1 of 288 (untranslated examples) to 11 of 290; #571 measured the same position effect on the output language. Moving it off the end took it to 5 of 290.

Bar (round 3) Status Evidence
B1a, 0 wrong-language accepted ✅ the only flags are Vas-y let's go. ×3 (your own 15 characters) and three correct Danish outputs the new family rule now accepts
B1b, ≤ 10 % refused on language per language ✅ 0 % in all 16
B1c, English 100 % English ✅ 33/33, including your five #585 dictations
B2a, 0 fabricated incompleteness accepted ✅ 0/356. The check flags 15/354 on the old prompt, all genuine, and 0 false flags in 2 130 outputs
B2b, rule 7 kept 3/3 on 5-rambling ✅ both arms
B3a, 0 bullets on prose ❌ 5 / 290 see below
B3b, a list on the enumerating fixture ≥ 2/3 ❌ 0/3, both arms declared in advance as a Mac limit; #523 measured the same
B4, not worse on #583's fidelity axes ✅ unrecalled 6→8 (inside the declared ±2), personLost 7→4, hedgeLost 5→5, hardened 1→0, fabricated 2→0

The example's own content reached no output in round 3, in either arm: the Norwegian leak of round 2 did not recur after the reorder. PolishIncompleteness would not have caught it — that sentence was about a hotel booking, not about the speaker's memory — and the check is blind to that by design.

The one bar that fails, unclosed and recorded here

5 of 290 prose outputs append a list that restates the paragraph above it, two of them accepted. All five are on agent-translated fixtures, in Danish (3), Dutch (1) and Swedish (1) — synthetic text in languages nobody on this project reads, measured on a Mac that already understates the phone. It is added content, not a formatting preference, and no guardrail sees it: the lines reuse the speaker's own words, so segmentOverlap passes them.

It is not fixed and not guarded in this PR, deliberately. The numbers are in docs/research/587-structured-rewrite/round3/ (summarise.py round3, section B3), and #587 carries a comment pointing at them. In French and English — the languages of the device round, and Pierre's — C3 produces no bullet on prose at all, which is why the decision was to land rather than spend a fourth round against a proxy.

What contradicted the issue, and what I am not comfortable with

Round 4:

  • The near-spelling rule admits an alteration of a named person by one or two added or dropped letters (Martine for Martin). That is meaning damage (Structured rewrites past the guardrails: a product name replaced and a negation dropped #570's class), not an invention, and it is the trade the device evidence asked for. Changed letters and names of five letters or fewer are untouched.
  • The skip's branch in PolishService has no unit test. The service builds the real Apple FM engine and reads the App Group, so a test would call the model. The predicate, the floor, its decoding, the outcome, its detail, the degraded-outcome shape the notice keys on and the log line are all tested; the branch that assembles them is what check 1 above verifies on device.
  • A skipped dictation writes two export events when the polish toggle is on (the skip, then the Normal polish). That is deliberate — both happened — but it means outcomes counts one dictation twice across two keys.
  • The third meaning deviation of the device round (not on a short dictation) is not addressed here. Per decision 10 it belongs to Structured rewrites past the guardrails: a product name replaced and a negation dropped #570's guardrail grilling.

Rounds 1-3:

  1. Step 1 of decision 5's ladder does not hold for this mode. English rules with a language rule were measured translating a Danish dictation into English 3 of 3. Step 2 is what works, and it is what ships.
  2. The Danish failure of round 2 was ours, not the model's. Two refusals of correct Danish, by our own per-segment language check. That is now fixed, and its cost is unmeasured in the other direction: nothing here proves a genuinely Swedish output on a Danish dictation would be caught, because nothing produced one.
  3. 14 of the 16 example sets are agent-translated and have never been read by a native speaker. Two of the three languages that still produce the list defect are among them.
  4. The Mac barely rewrites (median length ratio 1.00 on the French fixtures), so the fidelity numbers are partly that. The device round is the gate.
  5. B3b is out of reach on the Mac for every prompt tried, which matches Structuré — a Smart Mode that rewrites a long dictation into clear paragraphs #523's own measurement. Only the device can say whether a genuine spoken enumeration becomes a list.

Verification at the head commit: swift test 2142 tests, 0 failures (1 skipped). swiftlint lint --strict 0 violations in 293 files. xcodebuild build DictusApp with the keyboard on the iOS Simulator: BUILD SUCCEEDED. Round 4's changes are device-validated (ac1c689). Round 5's notice and export event are not; that is check 1 at the top.

🤖 Generated with Claude Code

Pivii and others added 6 commits September 21, 2026 21:43
…all (refs #587)

Decision 11's four bars, how each is measured, the fixture sets, the stop
rule for the language ladder, and the fabrication check's specification,
committed before the first model call of the round.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…refs #587)

Decision 6: rule 7 stays, and an output reporting the speaker's recall
failing is refused when the transcript carries no such sentence. Narrow on
purpose: recall-failure phrasings only, in six languages, matched strictly
on the output and leniently on the input so a spoken form rewritten into a
written one is never refused. Its own guardrail slug, a contract field off
by default, turned on for Structuré alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ranslated languages (refs #587)

device-structured-0918-fr.json: nine French Structuré dictations of
2026-09-18 to 09-20, raw verbatim, including the three accepted
fabrications and the bullet-per-sentence output. device-structured-en.json:
the five English dictations #585 measured coming back in French.
translated-structured.json: six French dictations translated by the agent
into English and the 14 other Apple FM languages, for output language only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…delity rounds (refs #587)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…(refs #587)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…n Danish (refs #587)

Captures of every run, the console output, and the script that reads them
against bars.md. C1 holds the output language in 15 of 16 Apple FM
languages and translates the Danish T3 fixture into English 3 of 3, which
is over decision 5's 10 % bar: the round stops, C1 is not landed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: getdictus/dictus-ios/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 382aa6e2-f02c-43b5-a520-37cdcfbde03e

📥 Commits

Reviewing files that changed from the base of the PR and between ec5d4c5 and ac1c689.

📒 Files selected for processing (97)
  • DictusApp/Polish/PolishDebugExporter.swift
  • DictusCore/Sources/DictusCore/LogEvent.swift
  • DictusCore/Sources/DictusCore/Polish/PolishAcceptanceContract.swift
  • DictusCore/Sources/DictusCore/Polish/PolishGrounding.swift
  • DictusCore/Sources/DictusCore/Polish/PolishGuardrail.swift
  • DictusCore/Sources/DictusCore/Polish/PolishIncompleteness.swift
  • DictusCore/Sources/DictusCore/Polish/PolishMetrics.swift
  • DictusCore/Sources/DictusCore/Polish/PolishPipeline.swift
  • DictusCore/Sources/DictusCore/Polish/PolishPostpass.swift
  • DictusCore/Sources/DictusCore/Polish/PolishService.swift
  • DictusCore/Sources/DictusCore/Polish/Prompts/SmartModeStructuredExamples.swift
  • DictusCore/Sources/DictusCore/Polish/Prompts/SmartModeStructuredPrompt.swift
  • DictusCore/Sources/DictusCore/Polish/SmartMode.swift
  • DictusCore/Sources/DictusCore/Polish/SmartModeCatalogue.swift
  • DictusCore/Sources/polish-harness/FidelityRound.swift
  • DictusCore/Sources/polish-harness/fixtures/device-structured-0918-fr.json
  • DictusCore/Sources/polish-harness/fixtures/device-structured-en.json
  • DictusCore/Sources/polish-harness/fixtures/translated-structured.json
  • DictusCore/Tests/DictusCoreTests/Polish/PolishIncompletenessTests.swift
  • DictusCore/Tests/DictusCoreTests/Polish/PolishNearSpellingGroundingTests.swift
  • DictusCore/Tests/DictusCoreTests/Polish/PolishScandinavianLanguageTests.swift
  • DictusCore/Tests/DictusCoreTests/Polish/PolishTrailingFenceTests.swift
  • DictusCore/Tests/DictusCoreTests/Polish/SmartModeCatalogueTests.swift
  • DictusCore/Tests/DictusCoreTests/Polish/SmartModeShortInputSkipTests.swift
  • DictusCore/Tests/DictusCoreTests/Polish/SmartModeStructuredPromptTests.swift
  • docs/research/587-structured-rewrite/arms/C1.txt
  • docs/research/587-structured-rewrite/arms/C2/da.txt
  • docs/research/587-structured-rewrite/arms/C2/de.txt
  • docs/research/587-structured-rewrite/arms/C2/en.txt
  • docs/research/587-structured-rewrite/arms/C2/es.txt
  • docs/research/587-structured-rewrite/arms/C2/fallback.txt
  • docs/research/587-structured-rewrite/arms/C2/fr.txt
  • docs/research/587-structured-rewrite/arms/C2/it.txt
  • docs/research/587-structured-rewrite/arms/C2/ja.txt
  • docs/research/587-structured-rewrite/arms/C2/ko.txt
  • docs/research/587-structured-rewrite/arms/C2/nb.txt
  • docs/research/587-structured-rewrite/arms/C2/nl.txt
  • docs/research/587-structured-rewrite/arms/C2/pt.txt
  • docs/research/587-structured-rewrite/arms/C2/sv.txt
  • docs/research/587-structured-rewrite/arms/C2/tr.txt
  • docs/research/587-structured-rewrite/arms/C2/vi.txt
  • docs/research/587-structured-rewrite/arms/C2/zh-Hans.txt
  • docs/research/587-structured-rewrite/arms/C2/zh-Hant.txt
  • docs/research/587-structured-rewrite/arms/C3/da.txt
  • docs/research/587-structured-rewrite/arms/C3/de.txt
  • docs/research/587-structured-rewrite/arms/C3/en.txt
  • docs/research/587-structured-rewrite/arms/C3/es.txt
  • docs/research/587-structured-rewrite/arms/C3/fallback.txt
  • docs/research/587-structured-rewrite/arms/C3/fr.txt
  • docs/research/587-structured-rewrite/arms/C3/it.txt
  • docs/research/587-structured-rewrite/arms/C3/ja.txt
  • docs/research/587-structured-rewrite/arms/C3/ko.txt
  • docs/research/587-structured-rewrite/arms/C3/nb.txt
  • docs/research/587-structured-rewrite/arms/C3/nl.txt
  • docs/research/587-structured-rewrite/arms/C3/pt.txt
  • docs/research/587-structured-rewrite/arms/C3/sv.txt
  • docs/research/587-structured-rewrite/arms/C3/tr.txt
  • docs/research/587-structured-rewrite/arms/C3/vi.txt
  • docs/research/587-structured-rewrite/arms/C3/zh-Hans.txt
  • docs/research/587-structured-rewrite/arms/C3/zh-Hant.txt
  • docs/research/587-structured-rewrite/bars.md
  • docs/research/587-structured-rewrite/capture-r1-device.json
  • docs/research/587-structured-rewrite/capture-r1-longform.json
  • docs/research/587-structured-rewrite/capture-r2.json
  • docs/research/587-structured-rewrite/capture-r3.json
  • docs/research/587-structured-rewrite/capture-r4.json
  • docs/research/587-structured-rewrite/raw/r1-device.txt
  • docs/research/587-structured-rewrite/raw/r1-longform.txt
  • docs/research/587-structured-rewrite/raw/r2.txt
  • docs/research/587-structured-rewrite/raw/r3.txt
  • docs/research/587-structured-rewrite/raw/r4.txt
  • docs/research/587-structured-rewrite/raw/summary.txt
  • docs/research/587-structured-rewrite/round2/capture-r1-device.json
  • docs/research/587-structured-rewrite/round2/capture-r1-longform.json
  • docs/research/587-structured-rewrite/round2/capture-r2.json
  • docs/research/587-structured-rewrite/round2/capture-r3.json
  • docs/research/587-structured-rewrite/round2/capture-r4.json
  • docs/research/587-structured-rewrite/round2/raw/r1-device.txt
  • docs/research/587-structured-rewrite/round2/raw/r1-longform.txt
  • docs/research/587-structured-rewrite/round2/raw/r2.txt
  • docs/research/587-structured-rewrite/round2/raw/r3.txt
  • docs/research/587-structured-rewrite/round2/raw/r4.txt
  • docs/research/587-structured-rewrite/round2/raw/summary.txt
  • docs/research/587-structured-rewrite/round3/capture-r1-device.json
  • docs/research/587-structured-rewrite/round3/capture-r1-longform.json
  • docs/research/587-structured-rewrite/round3/capture-r2.json
  • docs/research/587-structured-rewrite/round3/capture-r3.json
  • docs/research/587-structured-rewrite/round3/capture-r4.json
  • docs/research/587-structured-rewrite/round3/raw/r1-device.txt
  • docs/research/587-structured-rewrite/round3/raw/r1-longform.txt
  • docs/research/587-structured-rewrite/round3/raw/r2.txt
  • docs/research/587-structured-rewrite/round3/raw/r3.txt
  • docs/research/587-structured-rewrite/round3/raw/r4.txt
  • docs/research/587-structured-rewrite/round3/raw/summary.txt
  • docs/research/587-structured-rewrite/round4/guardrail-414-corpus-after.txt
  • docs/research/587-structured-rewrite/round4/guardrail-414-corpus-before.txt
  • docs/research/587-structured-rewrite/summarise.py

Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

This PR updates structured-mode prompts and guardrails, adds short-input Smart Mode handling, records skip telemetry, expands the fidelity harness, and adds structured-rewrite fixtures, benchmark captures, and research reports.

Changes

Structured rewrite and polish pipeline

Layer / File(s) Summary
Guardrails and output cleanup
DictusCore/Sources/DictusCore/Polish/*, DictusCore/Tests/DictusCoreTests/Polish/*
Adds fabricated-incompleteness detection, bounded near-spelling grounding, Scandinavian language matching, and removal of trailing layout-only lines.
Localized structured prompts
DictusCore/Sources/DictusCore/Polish/Prompts/*, DictusCore/Sources/DictusCore/Polish/SmartModeCatalogue.swift, docs/research/587-structured-rewrite/arms/*
Replaces the inline structured prompt with shared rules and language-specific examples for 16 language codes.
Short-input handling and telemetry
DictusCore/Sources/DictusCore/Polish/SmartMode.swift, DictusCore/Sources/DictusCore/Polish/PolishService.swift, DictusApp/Polish/PolishDebugExporter.swift
Adds optional Smart Mode input floors. Structured mode skips inputs below 200 characters, reruns without the mode, and records per-mode skip metrics and export counts.
Fidelity harness and fixtures
DictusCore/Sources/polish-harness/*, DictusCore/Sources/polish-harness/fixtures/*, docs/research/587-structured-rewrite/summarise.py
Adds localized arm resolution and reporting for language, foreign sentences, lists, incompleteness, fidelity, stray lines, and run outcomes.
Benchmark and research records
docs/research/587-structured-rewrite/*
Adds prompt bars, raw benchmark runs, JSON captures, grounding corpora, translated fixtures, and round summaries for the structured-rewrite evaluation.

Estimated code review effort: 5 (Critical) | ~90 minutes

Merge Risk: ⚪ Minimal · up to ac1c6

This change shortens the Structuré prompt, adds localized examples, and adds a guard against invented incompleteness. It also accepts close spellings of names, cleans trailing layout marks, and skips short inputs. No concrete defect remains after review. The identical before and after grounding captures are the expected result, because the new spelling rule changes no verdict on that corpus.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 75.24% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 105 functions across 23 files. (65 skippe… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main changes: shortening the Structuré prompt, adding language-specific examples, and rejecting fabricated incompleteness. It is specific and related to the pull reque…
Full details: Docstring Coverage

Explanation

Docstring coverage is 75.24% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 105 functions across 23 files. (65 skipped: 65 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Pivii

Pivii commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator Author

Corroborating evidence from #571 (Résumé), same Mac, same day — not re-measured here

The #571 agent reports, from its own runs (docs/research/571-summary/ on feature/571-resume-mode):

  • Step 1 fails harder on a mode that genuinely rewrites. Same shape as C1 (English rules, rule 1 = language, one EN + one FR example): 55 of 130 outputs came back in English across the 13 non-FR/EN Apple FM languages (de 10/10, ja 10/10, es/it/ko/pt/vi/zh ~5/10; da/nb/nl/sv ~0).
  • Example order matters. With the English example last, 9/35 French dictations came back in English; with the French example last, 0/35 FR and 0/20 EN. C1 puts the English example last, and every C1 translation in this bench went to English (3× Danish T3, 1× French F5), which fits that effect.
  • Step 2 probe (same rules, only the two examples translated into the transcript's language): de/es/ja/zh wrong-language 0/40 against 28/40 on step 1.
  • Adding "never in English unless the text is English" to the language line did nothing (35/77 vs 31/78).

Read together with this PR's numbers, it supports climbing to step 2 rather than re-wording step 1. That decision remains Pierre's. These are the #571 agent's measurements; this PR did not reproduce them.

🤖 Generated with Claude Code

Pivii and others added 6 commits September 22, 2026 15:07
…e (refs #587)

Decision 5 step 2's seam, mode-neutral so #571 can carry the same patch.
SmartModePrompt gains an optional table of instructions keyed by NLLanguage
code, looked up by exact code then base subtag, falling back to the one
prompt. PolishJob carries the transcript language: the forced transcription
language if set, else the detected one. The pipeline resolves the task once
before the engine, and the engine keys its session cache on the prompt
text, so a session warmed with the fallback is never reused for another
language's prompt. Prewarm resolves with the forced language when there is
one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…587)

A fixture's lang stands for the forced transcription language on the
per-language path; the auto path passes the detected one, as PolishService
does, so a Smart Mode's localized examples are exercised by the harness.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
C1 closed 18 of 354 Mac outputs on a lone ``` or --- line, all accepted.
The post-pass now removes trailing lines made only of those tokens, before
the guardrails; a line carrying any other character is left alone, so no
word can be removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…he transcript's language (refs #587)

C2 keeps C1's rules byte for byte and translates its two worked examples
into each of the 16 Apple FM language codes; its fallback is C1 exactly.
The sets live in SmartModeStructuredExamples and the arm files are dumped
from the Swift composition, so a landed C2 is the benched C2. Not wired
into the catalogue yet. The fidelity round accepts a directory arm and
resolves it per fixture with the app's own lookup rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The tables store Norwegian under nb, the code NLLanguageRecognizer answers;
a transcript labelled with the macrolanguage code no fell back silently.
Found by CodeRabbit on PR #589.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…0 outputs: fails, not landed (refs #587)

Round 2, bars.md §10. C2 returns no wrong-language output in any of the 16
Apple FM languages and holds B2 and B4, but fails B1b as the bar reads it in
Danish (2 of 18 refused on language, both a correct Danish sentence the
recogniser reads as Norwegian at 0.99) and fails B3a (11 of 290 prose
outputs carry a list line, against C1's 1 of 288). One accepted Norwegian
output closed on the Norwegian example's own line. The round stops; step 3
is not attempted and C2 is not wired.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Pivii Pivii changed the title Structuré campaign PR 1: fabrication check lands; the short rewrite fails step 1 in Danish and is not landed (refs #587) Structuré campaign PR 1: fabrication check, fence strip and per-language seam land; C1 and C2 prompts benched, not landed (refs #587) Sep 22, 2026
Pivii and others added 4 commits September 22, 2026 19:51
…s example order, more trailing artefacts (refs #587)

The language check now treats da/nb/no/sv as answering for each other, which
is what C2's Danish failure was: a correct Danish sentence read as Bokmål at
0.993. The candidate C3 is C2 with the enumeration example first and the
prose example last, on #571's measurement that the last example is the one
the model copies. A lone dash and an ellipsis join the trailing-artefact
list. Bars, fixtures and the 10 % threshold are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ge rule, 710 outputs (refs #587)

Round 3, bars.md §11. C3 holds B1 in every one of the 16 Apple FM languages
(0 refused on language anywhere, 0 wrong-language accepted after the hand
read), B2 (0 fabricated incompleteness, rule 7 kept 3/3) and B4. It still
fails B3a: 5 of 290 prose outputs carry a list line, against 11 for C2 and
1 for C1, and all 5 append a list that restates the paragraph above it. The
round stops; C3 is not landed and step 3 is not attempted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every set carries the same English rules, the enumeration example first and
the prose example last, no sentence about the speaker's memory, no person
named, and stays under the declared size. The last assertion records that
the sets are benched and not wired.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Pivii Pivii changed the title Structuré campaign PR 1: fabrication check, fence strip and per-language seam land; C1 and C2 prompts benched, not landed (refs #587) Structuré campaign PR 1: fabrication check, fence strip and Scandinavian language fix land; three prompt candidates benched, none landed (refs #587) Sep 22, 2026
2 966 characters against 5 556, seven one-line rules, two worked examples
per Apple FM language, no counter-example block. Benched over three rounds
and 2 130 Mac outputs: 0 % refused on check=language in all 16 languages
against up to 29 %, English 33/33, bullets on prose 5/290 against 72/289,
fabricated incompleteness 0/356 against 15/354, fidelity not worse on
#583's axes. The largest dictation that fits rises from 3 972 to about
4 900 characters of speech.

The one bar it does not hold is recorded rather than chased: 5 of 290 prose
outputs append a list restating the paragraph above, on agent-translated
Danish, Dutch and Swedish fixtures. No guardrail is added for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Pivii Pivii changed the title Structuré campaign PR 1: fabrication check, fence strip and Scandinavian language fix land; three prompt candidates benched, none landed (refs #587) Structuré: rewrite the prompt short, one example set per Apple FM language, and refuse an invented incompleteness (refs #587) Sep 22, 2026
@Pivii
Pivii marked this pull request as ready for review September 22, 2026 18:42
Pivii and others added 2 commits September 23, 2026 16:52
)

Device round of 2026-09-23: the speaker said Hermes, Parakeet wrote airmes,
the model wrote Airmesh, and the grounding check refused a faithful
four-item English list. A name is now also grounded by an input word, or
two adjacent input words joined, within an insertion-deletion distance of
1 for six letters and 2 from seven; five letters or fewer stay exact.

A substituted letter is never accepted: a Damerau-Levenshtein draft grounded
Sophie against Sophia, #414's W2 fabrication. Replayed on the #414 corpus,
the final rule is identical to the old one anchor by anchor (grounding
caught 7/11, 0/328 false rejections, before and after).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two of the three meaning deviations of the 2026-09-23 device round were on
short input: a 28-character question moved into the first person, and a
90-character line gained a leading Voici. Below its floor the mode is now
skipped, not refused: the dictation takes the path it takes with nothing
armed (Normal polish with the toggle on, raw text otherwise), and nothing
is announced.

The floor is a SmartMode field, set on Structuré only, read from the 51
device dictations of 2026-09-13 to 09-23: every one under 120 characters is
a single sentence, and none under 200 has more than two. The skip is
logged as smartModeSkipped reason=shortInput and carried on the Normal
polish's metrics event, with a per-mode count in the debug export.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pivii and others added 3 commits September 23, 2026 16:55
…587)

The previous commit said no device dictation under 200 characters has more
than two sentences. One does: a 149-character dictation counting off three
tasks, which is also the floor's known cost. 15 of the 16 are one or two.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t (refs #587)

Two gaps the 2026-09-23 device round left.

The user is told. A skip travels on the same SmartModeFailure channel as
every other did-not-run sentence, with the text inserted, so the keyboard
raises a notice in the #580 slot: "Structuré: dictation too short, text
inserted as dictated." / "Structuré : dictée trop courte, texte inséré tel
quel." It names its cause, unlike the refusal sentences, because this one
is knowable and dictating more is what changes it.

The skip is its own outcome, smartModeSkippedShortInput, carrying the mode,
the character count and the floor. It is emitted before the delegated call,
which is the fix: with the polish toggle OFF that call returns the raw text
without writing anything, so the export showed the dictation as no event at
all. It is now recorded with polish on and off.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The polish-event detail screen filled its Engine output section with
"(engine did not run successfully)" on a smartModeSkippedShortInput event.
Nothing failed there: the engine was never called, by design. The skip now
gets its own line, with the character count and the floor, because the
reader of a debug export is usually an agent triaging a report (#255).

Debug UI only, off the dictation path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant