Translate the docs and code-comments of an OSS repo from one human language to another (FR → EN, EN → ES, EN → DE, …) — Markdown prose for any OSS repo, plus code-comments in Python, JS/TS, PHP — with a locked glossary and structural validators that refuse to merge a broken translation.
Battle-tested on a 537-file FR → EN migration ; verified on EN → ES. Built on Claude Code in headless mode for overnight batches.
If you maintain an OSS repo where translatable content is scattered across prose (README.md, MDX docs) and code (Python docstrings, JSDoc, PHPDoc), the existing tooling lands you with one of three trade-offs :
- Product i18n tools (lingo.dev, Intlayer, Crowdin, …) translate JSON i18n keys for app strings, but don't touch your
README,CONTRIBUTING, or in-code docstrings. - Markdown-only translators (richawo/llm-translator, auto-translate-readme, astro-llm-translator, Co-op Translator, …) ignore the rest of your repo and can't enforce that
loopstaysloopin chapter 3 the way it does in chapter 12. - Generic LLM scripts translate Markdown AND code-comments (e.g. OpenAI Continuous Translation Action), but rely on humans to catch every drift : no locked glossary, no structural validators, no automatic refusal to merge if a frontmatter key vanished or
createClientbecamecréerClient.
repoglot is built for the seam between them : a single run, a single glossary lock, profiles per file type (markdown, code-comments, extensible), structural validators that refuse to merge a translation that broke a frontmatter key or renamed createClient to créerClient.
What you get:
- Translate hundreds of doc files in one overnight run — with per-file checkpoint and automatic recovery if Claude's quota resets mid-night. Files are grouped into ordered batches called tiers, so you can review the showcase content (READMEs, main guides) before shipping the rest.
- Keep terminology consistent across every file — a locked glossary prevents the model from translating
loopascyclein chapter 3 butloopin chapter 12. Cross-file drift is rejected at commit time. - Preserve internal links, slash commands, and code identifiers byte-for-byte — validators refuse to merge a translation that broke a frontmatter key or renamed
createClienttocréerClient. - Translate code-comments without touching the code itself — your source files never reach Claude raw ; only the comments do. Imports, function names, and syntax stay byte-identical.
- Editorial scaffolding so you keep control —
repoglot scanemits a per-file confidence-gradedi18n-manifest.yamlyou review before committing budget ;repoglot glossaryseeds a frequency-ranked starting draft for your editorial pass. - Headless overnight batching with Claude Max (
claude setup-token) or API key (ANTHROPIC_API_KEY).
- A 537-file Markdown documentation repo (FR → EN, May 2026): ~285k words, 8 tiers, 4 days. See
examples/case-study/for the actual specs / plans / journal / tasks. - A second OSS repo, EN → ES (internal pre-launch dogfood): ~6.7k words tier-1 vitrine, structural validators caught 11 bugs before v1.0.0 — all fixed in PRs #20–#27 (see CHANGELOG). Kept internal per policy; the chiffres are the receipt.
About this case-study being in French: the artefacts in
examples/case-study/are intentionally untranslated. They're the original specs/plans/journal/tasks that drove the FR → EN migration, written by the maintainer in his working language. We keep them as the historical record of what repoglot was built to solve — not as a vitrine of the translation output (that lives in the migrated source repo itself).
| Source → Target | Status | Notes |
|---|---|---|
| FR → EN | ✅ Battle-tested | Case study, 537 files, ~285k words |
| Any Latin-script pair | 🟢 Code-path shared | Generic source: / target: glossary schema since M1. FR → EN battle-tested ; EN → ES exercised via structural smoke fixture ; EN → DE, EN → IT, EN → PT-BR etc. share the same code path but are untested in CI — please report any pair you try. |
| EN ↔ JA, ZH, KO (CJK) | 🟡 v2.0 (planned) | Token-level glossary + no-space tokenization |
| RTL (AR, HE) | 🟡 v2.1 (planned) | Bidirectional markdown handling |
| Multi-target parallel | 🟡 v3.0 (vision) | One source repo, multiple targets in one run |
The architecture is language-agnostic by design. To start a new pair, edit the source: / target: labels in glossary.yaml and the {{SOURCE_LANG}} / {{TARGET_LANG}} placeholders in profiles/markdown/prompt.md — no code change needed.
| Your situation | Tool that fits |
|---|---|
| You ship a product (Next.js / React / Vue / Flutter) and need i18n JSON keys translated for end users | A full-product i18n SaaS like lingo.dev or Intlayer |
You want a single README.md translated, no glossary, no structural enforcement |
A plain LLM call or a Markdown-only translator like richawo/llm-translator / the auto-translate-readme GH Action |
| You want Markdown + code-comments translated and trust your reviewers to catch every drift manually | OpenAI Continuous Translation Action — closest existing alternative, ships without a locked glossary or structural validators (relies on human review) |
| You translate a JSON i18n bundle and that's the whole job | A JSON-focused tool (ai-i18n, i18n-ally, …) |
| You maintain an OSS repo with docs + code-comments + maybe i18n strings, hundreds of files, polyglot stack, and translation drift across them is not acceptable | repoglot |
What repoglot uniquely offers:
- Markdown + code-comments in the same run with a single locked glossary (no cross-format terminology drift).
- Headless overnight batching with Claude Max quota-cycle recovery — your migration survives a mid-night interruption.
- Structural validators that refuse to merge a translation that broke a frontmatter key, mis-counted a heading, or renamed
createClienttocréerClient. - Per-file
confidencemanifest produced byrepoglot scan, so you review the edge cases instead of trusting a black-box pipeline.
If your needs map cleanly to one of the other tools above, use that. repoglot is for the seam between them.
One-liner via the bootstrap script — clones the repo into ~/.repoglot and symlinks repoglot into ~/.local/bin:
curl -fsSL https://raw.githubusercontent.com/christopherlouet/repoglot/main/install.sh | bashOr, for the trust-but-verify crowd:
git clone https://github.com/christopherlouet/repoglot.git
cd repoglot
make install # symlinks bin/repoglot into ~/.local/bin
# make install PREFIX=/usr/local # system-wide, requires sudoBoth paths produce a single binary, repoglot, with the eight subcommands described below. After install, verify with:
repoglot --version # repoglot 1.0.1
repoglot --help # subcommand referenceRequirements: git, python3 with PyYAML, make, and jq. The claude CLI is needed at translation time (the bootstrap warns if absent but does not fail).
- Linux (Debian/Ubuntu) :
sudo apt install git python3-yaml jq make - macOS :
brew install python3 jq(git + make come with Xcode CLT). The harness'ssha256checksum step auto-falls back toshasum -a 256so GNU coreutils is not required.
To uninstall: cd ~/.repoglot && make uninstall.
repoglot calls claude --print per file. Two modes :
- Claude Max subscription (
claude setup-token) — flat-fee monthly, quota-bounded. The tool includes quota-cycle recovery : a tier interrupted by a quota reset resumes from the per-file checkpoint when the cycle refreshes. Best fit for one-off migrations. - API key (
ANTHROPIC_API_KEY) — pay-per-token, no quota cycles. As a back-of-envelope reference, the 537-file FR → EN case study (~285k source words) sent roughly 1.5M input tokens + 1.5M output tokens to the model over the four nights. Use the Claude pricing page to convert at your contracted rate. For most polyglot OSS repos, the dominant cost is your time on the editorial review of the glossary draft — not the model itself.
End-to-end timing on the case study, useful as an order-of-magnitude reference :
| Tier | Files | Words | Wall time | Notes |
|---|---|---|---|---|
| 1 (vitrine) | 42 | ~50k | ~9h headless + ~2h human review (Friday morning checkpoint) | Glossary locked here |
| 2-4 | ~370 | ~235k | 3 overnight runs (Fri/Sat/Sun) | No human review per night, just sampling next morning |
Rule of thumb : on Opus, expect ~10-30 seconds per file for markdown ; longer for code-comments (extract → claude → reinject is two model calls per file in some configurations).
The casual-user flow — 5 commands, 1 manual editorial pass :
cd /path/to/your/project
# 1. Bootstrap : scan + frequency-ranked glossary draft + blacklist template.
# --target-lang is required ; --source-lang defaults to en.
repoglot init --target-lang es
# 2. Edit specs/i18n/glossary.draft.yaml :
# - fill the `target:` field per term
# - strip the `_`-prefixed metadata (_frequency, _sources, _top)
# - rename to specs/i18n/glossary.yaml
# This is the irreducible human step — only you know your domain.
# 3. Dry-run preview : lock + inventory + batch --dry-run.
# No claude calls, no files modified ; you see the pipeline run.
repoglot run
# 4. Commit real translations (this is the irreversible step).
repoglot run --apply
# 5. Inspect progress / re-run on failures.
repoglot statusThe --dry-run default of repoglot run follows the terraform plan /
terraform apply convention : the safe-by-default preview comes first,
then the explicit --apply opts into the real call to Claude. The
intermediate steps (lock, inventory, batch) are silent on success ; you
see a 3-phase summary block + a next-step hint.
If you want finer control over each step (review the manifest before building the glossary, lock the glossary in a separate ceremony, run specific tiers, etc.), the individual subcommands still work :
# Equivalent to `init`
repoglot scan --root /path/to/your/project --source-lang en --target-lang es
repoglot glossary --root /path/to/your/project
cp "$(dirname "$(readlink -f "$(command -v repoglot)")")/../templates/blacklist.txt" \
/path/to/your/project/specs/i18n/
# Edit glossary as above
# Equivalent to `run` (or `run --apply`)
repoglot lock --root /path/to/your/project
repoglot inventory --from-manifest --root /path/to/your/project
repoglot batch --tier 1 --root /path/to/your/project # add --dry-run for preview
repoglot status --root /path/to/your/project --watchrepoglot ... |
Wraps | Purpose |
|---|---|---|
init |
scripts/init.sh |
One-shot bootstrap (scan + glossary + blacklist) |
scan |
scripts/scan-project.sh |
Emit i18n-manifest.yaml |
glossary |
scripts/build-glossary-draft.sh |
Frequency-ranked glossary draft |
lock |
scripts/lock-glossary.sh |
Lock the glossary after review |
inventory |
scripts/inventory.sh |
Generate inventory.json from the manifest |
batch |
scripts/translate-batch.sh |
Translate a tier (auto-routes per profile) |
run |
scripts/run.sh |
One-shot pipeline (lock + inventory + batch ; dry-run by default) |
status |
scripts/status.sh |
Migration progress dashboard |
verify |
scripts/validate-translation.sh |
Run validators on a translated file |
pr |
scripts/create-tier-pr.sh |
Open a PR for a translated tier |
All flags are forwarded verbatim, so any documentation that refers to ./scripts/<name>.sh --foo works identically as repoglot <subcommand> --foo.
scan-project.sh writes an editable i18n-manifest.yaml at the target root.
Each candidate file gets a confidence: high | medium | low verdict and a
pre-attributed profile (markdown or code-comments). Edit before the
first translate-batch — bump verdicts after manual review or remove entries
you don't want translated. See templates/i18n-manifest.example.yaml for
the full schema. Bypass the refusal with --force-low-confidence if you
know what you're doing; repos that predate this layer (case study FR→EN)
still run without a manifest (legacy warn mode).
A typical excerpt looks like this :
version: 1
generated_at: '2026-05-19T08:14:23+00:00'
generated_by: scan-project.sh v1
root: /home/dev/my-monorepo
source_lang: en
target_lang: es
entries:
- path: README.md
profile: markdown
confidence: high
reason: doc file at repo root
notes: ''
- path: docs/guides/getting-started.md
profile: markdown
confidence: high
reason: .md inside docs/ tree
notes: ''
- path: src/handlers.py
profile: code-comments
confidence: medium
reason: 22% comment lines
notes: ''
- path: docs/reference/api.md
profile: markdown
confidence: low
reason: .md without strong location signal
notes: '' # set to 'skip' or bump confidence after reviewYou read it, you decide. The runner trusts your edits.
Scan is git-agnostic by design. It does NOT consult .gitignore or
.git/info/exclude — the tool must work on directories that aren't
git repos at all (npm tarballs, vendored copies, arbitrary trees).
Hardcoded skips live in EXCLUDE_DIRS (scripts/scan/classify.py):
.git, node_modules, dist, build, target, vendor,
__pycache__, venv, .venv, .next, .nuxt, .cache,
integration-tests, __fixtures__. To exclude something else, either
rename the directory temporarily or delete the resulting manifest
entries by hand before invoking batch.
build-glossary-draft.sh consumes the manifest and writes a frequency-ranked
glossary.draft.yaml at <root>/specs/i18n/glossary.draft.yaml. The top
source terms (default --top 200) are extracted with a profile-aware
tokenizer (markdown strips frontmatter and code fences; code-comments
extracts comments per language), lowercased, and filtered against a
stop-word list (templates/stopwords-<lang>.txt; en and fr ship in v1).
The draft schema is a strict superset of the lockable glossary.yaml.
Each entry carries _-prefixed metadata (_frequency, _sources,
_generated_at, _generated_by, _top) — fill the target: field,
strip the _* keys, rename to glossary.yaml, then proceed to
lock-glossary.sh. The draft is an editorial starting point, not an
oracle: review every term, add domain-specific ones the scan missed,
remove noise.
Repos that mix prose (.md) with source files where only comments need
translating (.py, .js, .ts, .jsx, .tsx, .mjs, .cjs, .php)
get a second profile: code-comments. Selection is automatic — the
multi-profile resolver picks markdown for *.md/*.mdx and
code-comments for the source-file globs. Set PROFILE= in the
environment to override.
The translation pipeline for code-comments is structurally different
from the markdown one. Instead of sending the whole source file to
claude --print, the runner:
- Extracts comments to a JSON Lines record stream using one of the
three language-specific extractors (
extract-comments-py.pyviatokenize+ast,extract-comments-js.pyvia a hand-rolled state machine,extract-comments-php.pywith<?php/ heredoc / nowdoc awareness). - Translates only the
textfield of each record via Claude. The schema preservesraw_open,raw_text,raw_close,marker,line,col,placement,leading_ws,prefix_code,translatablefor lossless reinjection. - Reinjects the translated records back into the source file via
reinject-comments.py— language-agnostic, splices in reverse order so byte offsets stay valid. JSDoc/PHPDoc multi-line*continuations are re-applied; trailing comments keep their prefix code; tool directives like# type: ignore,// @ts-ignore,// @phpstan-ignore-linere-injectraw_textverbatim even if a translator mutated theirtextfield.
The check-comments.sh validator (run automatically after each file)
enforces three invariants: non-comment regions of src and dst are
byte-identical; translatable comment count matches; every untranslatable
directive is byte-preserved.
Auth: works headlessly with a Claude Max or Pro subscription (claude setup-token) or with an API key (export ANTHROPIC_API_KEY=sk-...). API mode bills per token; subscription mode is quota-bounded but flat-fee — pick what fits your usage.
The shipped extractors cover Python, JS/TS, and PHP — the popular web / data / PHP stacks. Adding a new language (Go, Rust, Java, Ruby, …) is a small, well-scoped contribution :
- Write a new extractor at
profiles/code-comments/hooks/extract-comments-<lang>.pythat emits the same JSONL v2 record schema as the existing ones. The smallest reference isextract-comments-py.py(uses Python'stokenize) ;extract-comments-js.pyshows a hand-rolled state machine for non-token-aware languages. - Add the file extensions to the resolver's profile globs in
profiles/code-comments/profile.yaml. - Add bats coverage at
tests/extract-comments-<lang>.batsfollowing the existing TDD-RED-GREEN structure (seeextract-comments-py.bats).
The reinjector and validators are language-agnostic — no changes
needed there. Open an issue with the lang-extractor label if your
stack isn't covered, and feel free to send a draft PR — feedback on
work-in-progress is welcome.
See the Setup checklist and Running the migration below for details on each step, validators, recovery, and tips.
This tool is designed for projects where:
- Documentation is markdown-heavy (hundreds of files, hundreds of thousands of words) AND / OR has substantive in-code comments in Python, JS/TS, PHP.
- Consistent terminology matters (a given source-language term must always map to the same target-language term across the whole repo).
- Internal references are critical (slash commands, file paths, anchors, frontmatter keys must be preserved character-for-character).
- The work fits in 2-4 overnight headless runs (Claude Max / Pro subscription) or fits a known API budget.
For smaller projects (< 50 files), manual translation is faster than setting up the tool. For very large projects (millions of words), professional translation services are likely cheaper than running this end to end.
┌──────────────────────────────────────────────────────────────────┐
│ │
│ scan-project.sh ───▶ i18n-manifest.yaml (per-file confidence │
│ + profile, user-editable) │
│ │ │
│ build-glossary-draft.sh ───▶ glossary.draft.yaml (freq-ranked │
│ top-N source terms, user fills) │
│ │ │
│ ▼ │
│ glossary.yaml ──┐ translate-batch.sh │
│ blacklist.txt ──┼──▶ ├── recovery.sh init/list-pending │
│ profiles/<x>/ ┘ ├── per file, resolve profile via │
│ prompt.md │ lib/resolve_profile.py │
│ ├── if markdown : build-prompt.py + │
│ │ claude --print (single call) │
│ ├── if code-comments : extract- │
│ │ comments-{py,js,php}.py ▸ claude ▸ │
│ │ reinject-comments.py (round-trip) │
│ ├── validate-translation.sh │
│ │ ├── core/validators/check-refs.sh │
│ │ ├── core/validators/check-glossary.sh │
│ │ ├── core/validators/check-comments.sh │
│ │ └── profiles/<x>/hooks/ │
│ │ └── check-structure.sh (markdown) │
│ └── recovery.sh mark-done + git commit │
│ │ │
│ ▼ │
│ create-tier-pr.sh ──▶ GitHub PR │
│ │
└──────────────────────────────────────────────────────────────────┘
A logical grouping of files translated as one batch and shipped as one PR. Tiers are ordered by criticality:
- Tier 1: showcase content (README, top-level docs, rules) — reviewed by a human before merge, sets the tone for the rest
- Tier 2-N: incremental content — opened as draft PRs after the overnight run, sampled in the morning, merged after sampling
The split into 4 tiers in the case study was chosen to match natural boundaries (audience-facing vs internal config vs auto-generated). Adapt to your project's structure.
A YAML file (glossary.yaml) listing canonical translations. The schema is
generic — source: / target: are free-text language labels, so the same
file shape works for any pair (FR → EN, EN → ES, …) :
source: French
target: English
locked_at: null # set by lock-glossary.sh
terms:
- source: chantier
target: project
forbidden: [worksite, jobsite]
context: "named work item — never the literal building site"
- source: boucle
target: loop
forbidden: [cycle, iteration]
context: "cycle is reserved for Red-Green-Refactor"The glossary is locked before the second night run. Locking prevents
silent drift if someone modifies it mid-migration. Lock = setting
locked_at and per-term locked: true + creating a git tag.
A flat-text file (blacklist.txt) of substrings that must appear
character-for-character in the translated output:
- All slash commands (
/work:work-explore,/dev:dev-tdd, …) - All paths visible in prose (
docs/guides/PROMPTING-GUIDE.md) - All frontmatter keys (
name:,type:,description:) - All identifier-like strings (
SKIP_PROMPT_CONTEXT, env vars, technical identifiers)
Word-boundary matching is used: /work:work-explore is correctly
distinguished from /work:work-explorer.
Polysemic words — words that are simultaneously a technical token AND
a common prose word (main as a git branch vs main feature, develop
as a branch vs the verb develop, ADR as an acronym vs adr in some
romance languages) should not be listed as bare tokens. The
count-match check is literal-substring and cannot reason about context :
a legitimate prose translation (main feature → característica principal) drops one occurrence and trips the validator.
The fix is to list the context instead. For each polysemic token, write out the git/code patterns rather than the bare word:
# bad
main
develop
ADR
# good
git checkout main
branch main
from main
branch develop
ADR-
an ADR
Cost : a handful of false-negatives if you forget a context. Benefit : zero false-positives on legitimate prose translation, no manual triage per tier.
A JSON file per tier (state-tier-N.json) tracking translation status
per file:
{
"tier": 1,
"files": [
{ "path": "README.md", "status": "draft",
"checksum_source": "abc123…" }
]
}Statuses: todo → draft → reviewed → merged. The runner only
processes todo files. Recovery from a crashed batch is automatic: re-run
the batch, it picks up where it left off.
Three validators run after each translation, before commit. They split into core (format-agnostic) and profile-specific:
core/validators/check-refs.sh— verifies blacklisted terms, slash commands, and file paths appear in the translated output as often as in the source.core/validators/check-glossary.sh— verifies forbidden translations don't appear (with code-block exclusion to avoid false positives).core/validators/check-comments.sh— code-comments profile: enforces the non-comment diff is empty, the translatable comment count matches, and untranslatable directives (# type: ignore,// @ts-ignore, …) are byte-preserved.profiles/markdown/hooks/check-structure.sh— markdown profile: verifies heading counts (H1-H6), code fence count, and frontmatter keys are identical.
Validators are TDD-tested. Their bats tests live in tests/.
-
Inventory your scope
Run
repoglot scan --root <repo>to produce ani18n-manifest.yamlwith one entry per translatable file (per-fileconfidence: high | medium | low, pre-assigned profile). Review everyconfidence: lowentry before going further — the runner refuses to start otherwise. Thenrepoglot inventory --from-manifest --root <repo>derives the tier-shapedinventory.jsonthatbatchconsumes. -
Seed the glossary
Run
repoglot glossary --root <repo>for a frequency-ranked draft (top-N source terms with_frequency+ sample sources). Edit the draft : fill thetarget:field of every term you want locked, strip the_*metadata, rename toglossary.yaml. Addforbidden:alternatives on terms that translators often mis-render. 50-100 terms is typically enough. -
Build the blacklist
List every slash command, file path pattern, frontmatter key, and identifier that must NOT be translated. The validator catches missing blacklist items in the diff between source and translated output.
-
Adapt the prompt template
profiles/markdown/prompt.mddefines the instructions Claude sees. The hard rules section is critical: it tells Claude what to preserve verbatim (frontmatter keys, slash commands, paths, code blocks). Test the prompt on 3-5 sample files before launching the full batch. -
Pre-flight
- Verify
claude --print "Translate to English: Hello world"works headlessly on your VM (you should get a translation back). Useclaude setup-tokenfor headless auth on Max/Pro, or setANTHROPIC_API_KEYfor API auth. - Ensure
bats tests/is green. - Adjust your target repo's CI workflow to tolerate the bilingual phase while a tier is in flight (e.g., a counter-validation script that accepts mixed source/target-language content).
- Verify
-
Branch strategy
We chose 4 PRs successive on
main(one per tier). Each tier branches frommain, after the previous tier merges. Rollback is granular. Avoid the "long branch merged at the end" pattern — it creates an unmergeable PR and defeats the purpose of incremental delivery.
Always run a single-file test before the overnight batch. This validates the prompt and the tool end-to-end, with the smallest possible blast radius.
# Render the prompt for a specific file (no Claude call):
repoglot batch --print-prompt README.md > /tmp/prompt.txt
# Read /tmp/prompt.txt to verify it looks right.
# Then translate one file with --no-commit so you can inspect manually:
repoglot batch --tier 1 --limit 1 --no-commit
git diff README.md # inspect the translation
# If unsatisfactory: git restore README.md, adjust prompt, retry.# Launch in background with logging:
nohup repoglot batch --tier 1 --root /path/to/your/project \
> tier-1.log 2>&1 &
echo $! > tier-1.pid
# Monitor from another shell:
repoglot status --root /path/to/your/project --watch # dashboard
tail -f tier-1.log # raw logIf a batch crashes mid-night, <root>/specs/i18n/state-tier-N.json retains
the per-file progress. Simply relaunch :
repoglot batch --tier 1 --root /path/to/your/project
# Picks up from where it stopped (skips files already in "draft" status).Dry-run vs real-run state :
--dry-runwrites to a separatestate-tier-N.dryrun.jsonso a previous rehearsal never blocks the next real batch. Both files live under<root>/specs/i18n/.
These notes summarize what worked and what we'd do differently next time.
- TDD on validators: writing the bats tests before the validator scripts caught several bugs in the regex word-boundary logic.
- State file granularity per file: easy to reason about, easy to recover, easy to display progress.
- Tier 1 with hybrid review: 1.5-2h human read-through of the showcase content (README + main docs + one in-depth section), automated validators on the rest. Catches the "reads weird in the target language" cases that automated tools can't see, while keeping the review budget reasonable.
- Glossary lock before tier 2: the human review of tier 1 may uncover terminology choices to revise. Locking after tier 1 review ensures all subsequent tiers use the validated vocabulary.
- Smaller initial scope estimate: we estimated 410k words from a
naïve
find+wc. Reality was 254k (~38% less) because website docs are auto-generated from.claude/*and didn't need separate translation. Always verify what's auto-generated before counting. - Earlier safety guard for --dry-run: the runner initially wrote
DRY-RUN markers into real files when
--rootdefaulted to the script's own repo. The fix was a one-line check, but caught only after the bug triggered. - Pre-built prompt rendering mode:
--print-prompt <file>was added late. It's invaluable for prompt iteration; should be in the tool from day 1.
- Don't use
--barewithsetup-token(Max/Pro) unless you've verified it works.--bareis API-key-only by design. - Don't translate identifiers (variable names, function names) even if they're in the source language. Renaming them is a separate refactor and can break downstream code/tests. The default rule: translate comments, keep identifiers.
- Don't auto-merge tier PRs. Even tiers 2-4 (draft) deserve a sample before merge. Drift can creep in subtly.
Most of the pipeline is project-agnostic. The parts you'll customize:
i18n-manifest.yaml— produced byrepoglot scan, edit per-file confidence / profile / tier assignments before the first batch. The manifest then drivesrepoglot inventory --from-manifest(single tier by default ; multi-tier shape is left as a per-project edit).templates/glossary.fr-en.yaml— start here, edit terms for your domaintemplates/blacklist.txt— start here, add your project's do-not-translate stringsprofiles/markdown/prompt.md— domain-specific instructions (e.g., "preserve the exact phrasing of API error codes in section X")
The validators and runner are generic and can be lifted as-is.
CHANGELOG.md— release notesCONTRIBUTING.md— dev environment, workflow, PR checklistSECURITY.md— vulnerability reporting policyexamples/case-study/spec.md— original spec of the case studyexamples/case-study/plan.md— original planexamples/case-study/journal.md— execution log of the 4 nightstests/— TDD tests for the validators and the runner