Manga reader with bubble detection, OCR, translation, and typesetting.
Open a folder of pages, capture speech bubbles (rectangle, freehand, auto-detect, or whole page), OCR them, translate them, style each entry, then render clean typeset pages to a destination folder — all locally, with optional DeepL / JPDB / custom endpoints.
- Local-first: bubble detection, PaddleOCR, local MTL models (Gemma, VNTL), OpenCV / TextSeg text cleaning.
- Live SVG overlay preview that mirrors the backend renderer (wrapping, spacing, scale, tilt, spin, shift).
- Per-entry layers, backgrounds, fonts, outline, spacing, transforms, speaker tags, and detector tuning.
- Standalone executables built with Nuitka (no install needed on the target machine).
This is an early release. Expect bugs, rough edges, and ongoing changes as Fox Reader continues to develop.
Version: 0.1.0 (pyproject.toml). The C launcher reports its own FOX_LAUNCHER_VERSION "1.0.0" (packaging/launcher/launcher.c).
- Features
- Quick Start
- How to Use
- Per-Entry Options Reference
- Text Clean & Typeset Reference
- Translation Providers
- Custom Endpoints (User Translation APIs)
- Hotkeys
- Settings & Configuration
- Environment Variables
- Installation in Depth
- Building Executables
- Download
- Releases
- Development
- Project Structure
- How It Works
- Troubleshooting
- Credits & Third-Party
- Image gallery with lazy thumbnails, keyboard navigation, and saved-page badges.
- Work mode (edit) and Reader mode (read) with fit-to-height / fit-to-width, zoom, pan, fullscreen, and webtoon strip mode.
Original / Savedswitch per page: compare the source scan against the typeset copy in the destination folder.- Displayed at
http://127.0.0.1:7954by default (config/fox_config.yaml).
- Four capture flows: rectangle (
c), freehand (f), bubble auto-detect, whole-page capture. - Right-to-left reading order by default, switchable to left-to-right, with manual reorder + resort.
- Two OCR engines, switchable in Settings: classic PaddleOCR and PaddleOCR-VL (GGUF vision-language side job).
- Greyscale toggle, Auto-OCR-on-select, click-to-re-read per entry.
- Built-in buttons per language: DeepL (
/translate/deepl, JA/ZH/KO), JPDB (/translate/jpdb), local MTL (/translate/ml), plus user-defined custom endpoints (/translate/custom). - Local MTL models: Gemma Q6/Q8 (JA/KO/ZH, GGUF via llama.cpp), VNTL Llama-3 8B (JA-EN with translation context + character roster).
- Auto-translate switch, translation context (previous pairs for context-aware models), character/speaker metadata for VNTL models, custom endpoint editor with preview/test/activate.
- Entry list = reading order (cards numbered 1..n, same numbers stamped on the overlay).
- Per-entry layer stack 1 (bottom) – 10 (top), gap-free and stable-sorted like the backend composite.
- Backgrounds:
auto(region dominant colour),color,transparent,clean(repaint artwork via text-clean). - Fonts served with data URIs + live preview; per-glyph symbol fallback (NotoSansSymbols2) at render time.
- Geometry: word/line spacing (fit-time), X/Y font scale (fit-time), X/Y shift in image pixels + X/Y/Z rotation in degrees (post-fit, text-only — the plate never moves).
- Inline edit-in-place over the bubble (
Entercommits,Shift+Enternewline,Esccancels).
- Clean methods:
region(selected region),ppocr(PaddleOCR boxes),textseg(glyph UNet++). Reconstruction fills:hybrid-level,hybrid,level,pyramid,telea,ns. - Detector tuning for TextSeg: speed presets (
best/fast/fastest/single), tile sizes (Off/256/512/1024/2048with paired overlap),glow/tta/transportswitches. - Preview renders to
cache/with progress polling (/api/progress/{name}), then Save promotes it (warns409 exists/stale_previewinstead of overwriting blindly). - Live overlay preview (toggleable) mirrors
src/fox_reader/typeset.py: same wrap/hyphen rules, same binary-search font fit, explicit per-word spacing, CSS-3D perspective tilt withcosfallback.
- Split flow: draw a cut across a region to divide it; pieces inherit styling and slot.
- Characters roster (VNTL speaker tags) with CSV import/export.
- Workspaces, layouts (
basic/default), dark/light themes. - Session + folder state APIs, single-instance launcher, graceful shutdown (
POST /api/shutdown).
- Python
>=3.12,<3.13 uvinstalled- For the UI build: Node.js
22+andpnpm(pnpm@11, CI uses Node 24)
Pick the PyTorch backend matching your system. <backend> is one of: cpu, macos, cu118, cu126, cu129, cu129_win, cu130, rocm71, rocm72 (ROCm wheels are Linux-only).
uv sync --extra dev --extra cpu
uv run --extra cpu fox-readerOr with the shortcuts (CPU backend):
# Windows
run.bat
# Linux / macOS
./run.shThen open http://127.0.0.1:7954. First launch redirects to /setup, which downloads the required models (bubble detector, OCR, MTL) from Hugging Face and marks each with a _completed file.
llama-cpp-python ships no wheels — every install compiles llama.cpp, and the CMAKE_ARGS decide whether the result has a GPU backend. It is therefore installed on its own via packaging/llamacpp.py, never through uv sync --extra gguf:
uv run --extra <backend> python packaging/llamacpp.py --device <backend> --install<backend> mirrors the torch extras (cpu, cu129, rocm72, macos). On CPU it looks for OpenBLAS (offers a download into nc/); on CUDA/ROCm it needs the matching toolkit (nvcc / ROCm 6.1+). Preview without installing, or install by hand (note --no-cache-dir is mandatory — uv's cache key ignores CMAKE_ARGS):
python packaging/llamacpp.py --device cu129 --print-args
CMAKE_ARGS="<printed line>" uv pip install --no-cache-dir llama-cpp-pythonWindows notes: hand installs need CMAKE_GENERATOR too (MinGW Makefiles + MinGW compilers, or Ninja inside a VS Developer Prompt for MSVC). MinGW builds also need -D_WIN32_WINNT=0x0A00 (printed as CFLAGS/CXXFLAGS) and copy libgcc/libstdc++/libwinpthread DLLs next to the built lib — --install handles all of this automatically.
Useful flags: --openblas {auto,download,off}, --cpu-baseline {avx2,avx,sse42}, --cuda-arch / --hip-arch, --python PATH, --find-openblas.
- Load a folder. Type a source folder path (or pick it in the folder modal) and optionally a destination folder. Loading sets the gallery, page dimensions, and which pages already have a typeset copy (
savedbadges). When source and destination are the same folder, saving overwrites and theOriginal / Savedswitch hides. - Pick a page. Click a gallery thumbnail or use
ArrowRight/ArrowLeft(htoggles the gallery). The viewer fits the page; zoom with+/-/0, pan by dragging. - Capture regions. Press
c(rectangle),f(freehand), the bubble button (bubble auto-detect), or page capture (whole-page OCR+translate). Drag on the page;Esccancels. Each capture becomes an entry card (reading order = list order). - OCR. Click an entry (or its OCR button) to read the region into the editor.
Auto OCRreads entries with no text on select; greyscale and language (languageSelect: japanese/chinese/korean) apply. - Translate. With
Auto TLon, translation follows OCR automatically. Otherwise click the entry's Translate button (DeepL / JPDB / MTL / custom, per language). Results land in the translation panel (#translatedText). - Confirm.
Confirm Entrystores editor + panel text onto the captured region. Re-confirming the same region replaces text but keeps styling. - Style entries. Expand an entry (
expand_more) for fonts, colours, layer, alignment, speaker, background, clean settings, spacing/scale/shift/rotate (see reference below). Every control repaints the overlay immediately (refresh); structural clean changes rebuild the panel (rebuild). - Preview & save.
Previewrenders the current page tocache/in a modal (progress polled);Savewrites it to the destination folder (Savein the modal commits the preview token).Generaterenders straight to the destination. Existing files need overwrite consent (409). - Compare. Flip
Original / Savedon pages with a saved copy. Editing locks while viewing Saved (entries describe the original).
Tips: right-click a card or region to hide it from the render; layers decide paint order (10 on top); empty text + clean/colour background erases a bubble; transparent draws text with no plate.
| Group | Controls | Values |
|---|---|---|
| Layer / Align / Speaker | Layer stepper, alignment segmented, speaker select | Layer 1–10 (ceiling = highest other + 1); align center/right/left; speaker None + character roster (meta_id) |
| Font / Outline | Family select, size select, font colour; outline width + colour | Size Auto + 6–72; outline Auto + 0–32 (auto ≈ round(size×0.09), min 1); colours Auto (detected) or hex |
| Spacing / Scale | W word spacing, L line spacing, X/Y stretch |
W Auto + 0.0–3.0×, L Auto + 0.5–3.0× (tenths); scale 0.5–2.0× (neutral 1.0, part of the font fit) |
| Shift / Rotate | X/Y nudge (image px), X/Y/Z rotation (deg) |
Shift bounded by page when loaded, backend clamps to ink; angles −180–180 (X/Y tilt plane, Z spins clockwise, all about block centre, text-only) |
| BG | Mode select + colour | auto / color (+hex) / transparent / clean |
Clean (when BG=clean) |
Method + fill + tweaks | Method region/ppocr/textseg; fill hybrid-level/hybrid/level/pyramid/telea/ns; speed best/fast/fastest/single; tile Off/256/512/1024/2048; glow/tta/transport per method knobs |
Defaults for untouched entries reproduce the pre-geometry render exactly (word=None, line=None, scale=1, shift=0, angle=0).
- What runs: clean-mode entries repaint artwork first (detection shared across frames), then entries composite bottom-layer-first (
typeset._ordered), each with its plate + fitted text + post-fit transforms. - Typeset fit: largest
6–72px that fits the padded type area (12% horizontal / 10% vertical padding), binary-searched then verified; spacing + scale participate in the fit, shift/rotation apply afterwards to the lettering only.
One detection pass runs per page, not per entry: entries sharing a detector (and, for TextSeg, the same tuning) share one result over the union of their regions. Animated pages detect once and reuse it across frames.
| Method | What it removes | Cost | Knobs |
|---|---|---|---|
region (default) |
Your selection verbatim — everything inside the shape, no model, cannot miss | Instant | None (fill + transport still apply) |
ppocr |
PP-OCRv6 detector quads, one per text line (~2 s/page, model already loaded) | Fast | None — box coverage has no halo for Glow to absorb (fill + transport still apply) |
textseg |
Glyph-level UNet++ segmentation + stroke-radius-relative mask refinement (counters filled, specks dropped, glow absorbed) | Slowest, most precise (20–40 s/page on CPU; runs on the padded union crop, not the full page) | glow, speed, tta, tile |
Preview marking: region = blurred plate, ppocr = diagonal hatch, textseg = dots. Unknown method names from stale projects fall back to region rather than failing the page.
Each text cluster is rebuilt inside its own neighbourhood (local crop, grown until ≥45% known pixels), largest first, so nearby paint — not distant art — is the source. Flat-paper holes are pre-filled with the ring median (avoids grey ghosts on white bubbles). Ordered best-first in INPAINT_METHODS; the default is the first non-transport entry because transport costs seconds per cluster. What the panel offers depends on the OpenCV build (caps probes xphoto.inpaint + flag, since OpenCV 5 ships an empty xphoto module that imports but cannot run): pure-xphoto fills hide when unavailable (degrading would leave Telea wearing their name), while hybrids degrade their detail term to Telea and keep their pyramid shading intact. Unknown fills fall back to the default.
| Fill | How it works | Best for / watch out |
|---|---|---|
hybrid-level (default) |
Pyramid shading + FSR high-frequency detail, level-fixed | Painted artwork; ~1 s/cluster trade that holds up |
hybrid |
Pyramid shading + transported edge detail (Telea when xphoto missing) | Soft art needing edge continuation |
hybrid-fsr |
Pyramid shading + FSR detail variant | Line art / hatching over shading |
hybrid-patch |
Pyramid shading + patch-synthesis detail (real paint texture) | Textured paint where grain matters |
level |
Pyramid solver with boundary-implied re-levelling | Smooth gradients with tone steps |
pyramid |
Push-pull interpolation + matched grain (no invention) | Soft airbrushed fills; hazes on very wide holes |
patch |
Multiscale patch synthesis | Real texture; can mismatch shading / ridge on wide holes |
transport |
Offset (shiftmap) cover with whole displaced pieces + tone fix | Structure at every scale (hair, folds); can invent content; ~20 s |
fsr / fsr-fast |
Frequency-selective reconstruction (BEST / FAST) | Straight edges continuing as themselves; hidden without xphoto flags |
shiftmap |
xphoto shiftmap in Lab | Offset fill; hidden without its flag |
telea / ns |
OpenCV diffusion inpaint (radius 7) | Always available; Telea is the universal fallback |
| Parameter | Applies to | Meaning |
|---|---|---|
method |
All | Detection source above (region/ppocr/textseg) |
fill |
All | Reconstruction above (hybrid-level default) |
glow (default on) |
textseg only |
Absorb the feathered glow around text into the mask (glow_k 2.0 vs 0.0 = CLI --no-glow); distances in stroke radii so cover titles and furigana behave |
transport (default on) |
All fills | Allow escalation: wide holes (half-width ≥ 60 px) in textured paint (median high-freq ≥ 2.0 in the 2–10 px shell) escalate smear-prone fills to offset transport; wide holes in flat paint correctly stay smooth. Needs xphoto transport or escalation stays off. Reconstruction groups by (fill, transport) |
speed (default fastest) |
textseg only |
best: flips + scales (1.0, 0.6, 0.4) + inverted pass (≤24 passes) · fast: no flips (≤6) · fastest: no inverted pass either (≤3; within 6.2% of best for ⅛ the cost — the interactive default) · single: one scale (1 pass, cheapest, no cross-scale vote to filter artwork). Inverted pass auto-skips on pages with no dark ground. Stale names → fastest |
tta (default on) |
textseg only |
Flip test-time averaging (4 views). Only best ships flips, so only there does off save 4×; elsewhere the switch is stored but inert (dimmed, still clickable) |
tile (default Off) |
textseg only |
Off (whole image, only whole-page view) / 256+48 / 512+96 / 1024/2048+192 overlap, cosine-blended. Small tiles exist for machines that cannot hold page activations, not for better detection. Only tabulated sizes accepted |
Switching method never loses tuning: speed/tta/tile ride along so toggling back restores them.
| Provider | Endpoint | Notes |
|---|---|---|
| DeepL | POST /translate/deepl |
Needs API token in Settings; buttons hidden until /api/translate/status allows; :fx keys use the free host |
| JPDB | POST /translate/jpdb |
Japanese parser API, token in Settings |
| Local MTL | POST /translate/ml |
Gemma Q6/Q8 (JA/KO/ZH, GGUF); VNTL Llama-3 8B (JA-EN, context + characters). Load/unload + VRAM check via /ml/* |
| Custom | POST /translate/custom |
User-templated endpoints (editor at /user_endpoints, preview/test/activate per language) |
Context pairs and character links ride only on MTL/custom requests, gated by the Context/Characters switches.
Any HTTP translation API can become a translator button. Open the editor at /user_endpoints (linked from /settings as “Manage endpoints →”). Each language holds up to 20 endpoints, of which 3 can be active at once — active endpoints become translator buttons in the reader next to DeepL/JPDB/MTL. Changes save immediately; endpoints persist in config/user_endpoints.yaml.
- Add Endpoint — starts a blank draft (default request
{text, src_lang, target_lang}, default response{text, error}). - Fill the blocks below (connection → wire format → languages → request schema → response schema → query/headers → encryption).
- Preview request —
POST /api/user_endpoints/preview. No network: the server validates the draft and returns exactly what would be sent (method/url/headers/body) plus schema examples. List fields are filled with samples ([[こんにちは, Hello], …], one roster entry + links) so you see their shape. When editing a saved endpoint, the form has no copy of the stored key, so the preview borrows it rather than failing. - Save Endpoint —
POST(new) /PUT(edit). Validation errors come back aserrorwith dotted paths (e.g.request_schema.children.text: …). The key is write-only and stripped from every browser-visible response. - Test — per-endpoint Test sends one real request with
こんにちは(or your text), empty history, and showssent / translated / alternatives / messageplus the exact request. Use it before activating. - Activate per language —
PUT /api/user_endpoints/{id}/activewith{"language": "japanese"}(DELETE deactivates). The reader fetchesGET /api/user_endpoints/activeand renders one button per active language. - Translate — in the reader, select text and click your endpoint’s button (or Auto TL if it routes there). History and roster ride only when the schema carries those nodes (and the Context/Characters switches are on). Selections over
max_text_lengthare refused client-side (413) before anything is billed or rate-limited. - Manage — edit, delete (
DELETE /api/user_endpoints/{id}), or swap actives at any time.
| Field | Meaning |
|---|---|
| Name | Button label — max 6 chars, sits beside the built-in engines |
| Method | GET/POST/PUT/PATCH/DELETE |
| Hostname / URL | Host with optional path; a pasted https:// prefix is fine. <TEXT>, <SOURCE_LANG>, <TARGET_LANG> are substituted — e.g. api.example.com/<SOURCE_LANG>/translate |
| Scheme / Port / HTTP | http/https, optional port 1–65535, HTTP version Default/2 (HTTP/2 via httpx) |
| Timeout | Seconds per attempt (> 0, default 20) |
| Retries | 0–5. Retries timeouts and 408/425/429/500/502/503/504 with exponential backoff (0.4s → 4s cap, honours Retry-After) |
| Max text length | 100–10000 chars (default 2000). Longer selections fail before sending |
| Target language | Sent wherever the schema uses Target Language (default en) |
- Request body (
body_format):json(JSON body),form(urlencoded body),query(payload appended to URL, no body),text(raw body).Auto=queryforGET/DELETE, elsejson.form/queryneed a flat object at the request root. - Response body (
response_format):json(decoded + read through the response schema) ortext(the whole body is the translation; the response schema is then unused). - Repeat list keys (doseq):
falsesends a list as one joined value;truerepeats the key (dt=t&dt=at). Applies toform/querypayloads. - Join response chunks with (
text_join): separator when a repeat template yields several fragments (\n,\tunderstood, else""). - Static query parameters: always appended to the URL (placeholders substituted); a
query-format body overrides same-named keys. - Request headers: sent as entered (placeholders substituted); an explicit
Content-Typeoverrides the body-format default (application/json,application/x-www-form-urlencoded,text/plain; charset=utf-8).
Tick the languages this endpoint serves (editor suggests japanese/korean/chinese; the list itself is free-form lowercase). The optional per-language code is what the API expects (japanese → ja); it is sent wherever the schema uses Source Language. One translator button is created per active language.
Typed JSON nodes: dict (named children), list (items and/or response-only each), and primitives string/int/float/bool/null plus uuid (request-only; mints a fresh id per request — variant uuid1/uuid2/uuid4, prefer uuid4 since uuid1 embeds this machine’s MAC).
Dynamic sources (the only moving parts; everything static is sent verbatim):
| Source | Sends |
|---|---|
text |
The selection (exactly one required) |
src_lang / target_lang |
Language codes (at most one each, optional) |
context |
Earlier [[source, english], …] pairs, oldest first (at most one; only when the node exists — otherwise requests are byte-identical with or without history) |
character_info + context_character_links |
Roster [{meta_id, name_en, name_ja, gender, alias_en, alias_ja}, …] + one speaker id (or null) per context pair. Always both or neither |
error / alt |
Response-only — rejected in a request schema |
Rules enforced on save: exactly one Text; repeat templates (each) and uuid only in requests where documented; Text/src_lang/target_lang must be string nodes; list-array sources must be bare list nodes (they consume the whole array, no items/each inside).
Placeholders work inside any static string, header, query value, or the hostname itself:
| Token | Expands to |
|---|---|
<TEXT> |
The selection (sealed first when request encryption is on — even in the URL, which is the most-logged part) |
<SOURCE_LANG> |
Source code for the ticked language |
<TARGET_LANG> |
Target language field |
- Exactly one Translated Text node (any shape — string, chunk list, or
[text, meta]pairs are flattened and joined withtext_join). - Optional Error node (string/null): an API error message beats a bare HTTP status.
- Optional Alternatives node (string/list): extra candidates, deduplicated, minus the main text.
- Repeat template (
each) on a list collects an unknown number of chunks (e.g. per-sentence arrays) into one translation. - With
response_format: text, skip the tree: non-2xx returnsHTTP <code> — <snippet>, empty bodies and undecryptable replies are502s, never silent.
Off unless a key is given. When on, the selection is sealed before it leaves this machine and the reply is opened on arrival — proxies, middleboxes, and request logs carry opaque tokens instead of the text. The endpoint holds the same key and reverses it: this hides content from the network, not from the API itself. Token length still tracks plaintext length, and timing/count are visible.
- Key types:
fernet(paste the 44-charFernet.generate_key()string) orpassphrase(words stretched with scrypt, fixed saltfox-reader/user-endpoint/v1,N=16384, R=8, P=1— a compatible server must use these verbatim, then base64url the 32 bytes). - Directions are independent: encrypt-before-sending and/or decrypt-on-arrival (at least one must be chosen). Names/aliases seal per-field (
meta_id/genderstay plaintext so links still join); context pairs seal per side; response fragments open one by one, then join. - Rotation: up to
4keys, whitespace/comma-separated — first encrypts, all are tried for decrypt. Add new-first, migrate the server, drop the old. - Replay guard (
max_age,30–86400s, blank = off): sealed replies carry a timestamp; refuse stale ones only when both clocks are trusted. - Requires the
cryptographypackage; the editor disables the switch without it. The key never reaches the browser except for encryption (edit previews borrow the stored secret server-side).
1. Minimal JSON POST (https://api.example.com/v2/translate):
// request root (dict): {"text": Text, "source": Source Language, "target": Target Language}
{"text": {"type": "string", "source": "text"},
"source": {"type": "string", "source": "src_lang"},
"target": {"type": "string", "source": "target_lang"}}
// response root (dict): {"translatedText": Translated Text, "message": Error}
{"translatedText": {"type": "string", "source": "text"},
"message": {"type": "string", "source": "error"}}2. GET query API (method GET, body query): same request tree; add static query {"key": "<YOUR_KEY>"} and a header if needed. The payload rides in the URL; over-long selections hit max_text_length first.
3. Chunked reply with alternatives: response root dict → data: dict → sentences: list(each: string→Translated Text) plus candidates: list→Alternatives, err: string→Error, text_join: "\n".
20endpoints / language,3active / language, name ≤6chars,max_text_length 100–10000,retries 0–5,timeout > 0,max_age 30–86400or blank.400empty text ·404unknown endpoint ·409never (endpoints overwrite by id) ·413selection too long ·500unbuildable request ·502bad JSON / missing Translated Text / undecryptable ·503/504unreachable / timeout ·429/5xxretried per policy.- Automatable:
GET /api/user_endpoints,GET …/active,POST …/preview,POST /,PUT /{id},DELETE /{id},POST /{id}/test,PUT|DELETE /{id}/active.
Typing in inputs/textareas (or the editable translation panel) disables hotkeys. Loading overlay, backend-status overlay, and folder modal also suppress them.
Work mode (default):
| Key | Action |
|---|---|
c |
Rectangle capture |
f |
Freehand capture |
h |
Toggle gallery panel (re-fits viewer) |
ArrowRight / ArrowLeft |
Next / previous page |
Reader mode (workspace-reader):
| Key | Action |
|---|---|
ArrowRight / ArrowDown / Space |
Next page |
ArrowLeft / ArrowUp |
Previous page |
f |
Fullscreen toggle |
s |
Strip (webtoon) mode toggle |
+ / = , -, 0 |
Zoom in / out / refit |
Esc |
Back to work mode |
Editors & dialogs:
| Context | Key | Action |
|---|---|---|
| Inline bubble editor | Enter / Shift+Enter / Esc |
Commit / newline / cancel |
| Colour hex box | Enter |
Commit parsed hex |
| Eyedropper / capture stroke | Esc / right-click |
Cancel pick / cancel stroke |
| Preview modal, character/appearance modals | Esc |
Close |
| Folder path input | Enter |
Load folder |
On-disk state (repo root, gitignored except fox_config.yaml):
config/fox_config.yaml—host: 127.0.0.1,port: 7954,mtl_dir: models,theme: dark|light,layout: basic|default.host/portare immutable at runtime.config/settings.yaml—mtl_defaultsper language (japanese/korean/chinese),devices(paddleocr/paddleocr_vl/bubble/translator/textseg: auto|<device>,mtl_threads),deepl_api_token/jpdb_api_token,ocr_engine: paddleocr|paddleocr-vl.config/user_endpoints.yaml— custom translation endpoints.config/workspaces.yaml— panel layout sizes.models/— weights (bubblemodel.safetensors, PaddleOCR det/rec + VL GGUF, manga seg safetensors, MTL Gemma/VNTL). Downloaded once via/setup; each completed download touches a_completedmarker.fonts/— your.ttf/.otffaces (scanned + cached in.fonts_cache.json); two fallbacks are compiled into the binary.NotoSansSymbols2is a render-time per-glyph fallback, never listed. Guide (adding faces, licensing, suggestions):/FONT.md.cache/— scratch previews/masks; cleared at startup and shutdown.
Pages:
/setup— first-run wizard: required + MTL download state, per-file progress, optional OCR-VL side job, finish/close./settings— devices (per-slot AUTO + thread count), per-language MTL defaults, API tokens (validated), OCR engine switch./user_endpoints— custom endpoint CRUD, mustache-style preview, test call, per-language activation.- Characters modal — VNTL speaker roster CRUD + CSV import/export; entry Speaker select tags
character_id.
| Variable | Default | Effect |
|---|---|---|
FOX_READER_LOG_LEVEL |
INFO |
Log verbosity |
FOX_READER_ROOT |
exe dir / source root | Override project root (models/fonts/config resolution) |
FOX_READER_DISK_ASSETS |
off | Read embedded frontend/fonts from disk (dev) |
FOX_READER_MTL_THREADS |
settings heuristic | Override MTL thread count |
FOX_READER_IGNORE_MTL_MEMORY |
off | Skip the VRAM gate before loading local MTL |
FOX_READER_MTL_VERBOSE |
off | Verbose local-MTL (llama.cpp) logging |
FOX_READER_VL_VERBOSE |
off | Verbose PaddleOCR-VL logging |
FOX_READER_SHUTDOWN_TOKEN |
open | Gate POST /api/shutdown (launcher sets it) |
USERNAME/USER fall back to fox-reader for endpoint rendering.
| Backend extra | Torch index | Platform |
|---|---|---|
cpu |
CPU | Win/Linux/macOS |
macos |
PyPI | Apple Silicon |
cu118 / cu126 / cu129 / cu130 |
matching CUDA index | Linux (CUDA toolkit needed for GGUF builds) |
cu129_win |
CUDA index (Windows split) | Windows |
rocm71 / rocm72 |
ROCm index | Linux only (no ROCm wheels on Windows/macOS) |
uv sync --extra dev --extra <backend>
uv run --extra <backend> fox-readerFrontend (TypeScript + Vite + Preact):
cd frontend
pnpm install
pnpm dev # watch mode (Terminal 1) + `uv run fox-reader` (Terminal 2)
pnpm build # → static/app.js (gitignored IIFE bundle)Windows pnpm install may fail with UNKNOWN ... symlink: enable Developer Mode (Settings → System → For developers) or pnpm install --node-linker=hoisted. packaging/build.py retries with that flag automatically.
New to this? Start with
packaging/build.md— the copy-paste guide (what to download, which--devicefor your GPU, build, run) for Windows and Linux.
Standalone builds compile the backend with Nuitka, embed the frontend, and ship a small C launcher (the user-visible entry point).
Result layout:
fox-reader/
launcher(.exe) starts backend, opens UI, closes both on exit
bin/ compiled backend + linked libs (+ MSVC runtime DLLs + .dist-info metadata)
cache/ scratch (empty, created at build)
config/ fox_config.yaml
fonts/ your fonts
models/ weights (downloaded on first run)
uv,pnpm, Node.js22+(CI pins 24 + pnpm 11), Python3.12- C compiler for the launcher (verified by compiling+linking a smoke program;
python packaging/toolchain.py --whichshows the winner):- Windows: VS Build Tools (“Desktop development with C++”) or MinGW-w64 (
gcconPATH; use winlibs/MSYS2mingw-w64-x86_64-gcc, not MinGW.org baregcc). With no compiler, the build asks Nuitka to fetch its MinGW64. - Linux:
gcc/clang(apt install build-essential); CUDA variants pin GCC per toolkit (11.8→11, 12.6/12.9→13, 13.0→14). - macOS: Xcode CLT (Metal builds only on macOS).
- Windows: VS Build Tools (“Desktop development with C++”) or MinGW-w64 (
- GPU builds: matching CUDA toolkit (
nvcc) or ROCm 6.1+; CPU GGUF builds want OpenBLAS (apt install libopenblas-devon Linux, or--openblas download).
python packaging/build.py
python packaging/build.py --clean # wipe previous artifacts first
python packaging/build.py --skip-frontend # reuse prebuilt static/app.js
python packaging/build.py --keep-venv # reuse build venv next time
python packaging/build.py --deep-compile # compile deps too (hours)
python packaging/build.py --device cpu # cpu|macos|cu118|cu126|cu129|cu129_win|cu130|rocm71|rocm72
python packaging/build.py --nuitka-arg ARG # repeatable, forwarded to Nuitka
python packaging/build.py --no-blas # alias for --openblas offGGUF / CPU tuning:
python packaging/build.py --no-gguf
python packaging/build.py --gguf-device cpu # CUDA torch + CPU llama.cpp
python packaging/build.py --openblas download
python packaging/build.py --cpu-baseline sse42 # avx2 (default) | avx | sse42
python packaging/build.py --cuda-arch 75-real,86-real,120-virtual
python packaging/build.py --hip-arch gfx1100,gfx1201Defaults compile only fox_reader + entry (5–15 min); --deep-compile is the slow release path. Nothing targets the build machine specifically (GGML_NATIVE=OFF, pinned ISA, virtual GPU arches for JIT fallback). Archives are named fox-reader-v{version}-{device}[-gguf[-ggufDevice]]-{platform} (zip on Windows, tar.gz on Linux).
Shared libraries always travel as DLLs, never as data files: --noinclude-data-files covers *.so*/*.dll/*.dylib/*.pyd, because a versioned Linux library such as torch's libgomp.so.1 slips through Nuitka's built-in filter, gets collected as data while the DLL scan claims the same path, and stops the build with a data-file-vs-DLL conflict.
Linux then still ships every shared library twice into bin/: once under its package (torch/lib/libtorch_cpu.so, where the package's own RUNPATH finds it — the layout Windows ships only, and works) and once flat at the top level (Nuitka's dependency scan, e.g. for torchvision). Both copies are live, so the build replaces the flat copies with relative symlinks instead of deleting them — every lookup path keeps resolving, the bytes ship once. Version triplets (libllama.so, .so.0, .so.0.20.0, symlinks in the wheel that arrive as three full copies) collapse the same way after a SHA-256 check. Also dropped, since nothing the app runs imports, spawns, or dlopens them: torch/test, torch/include, the test executables and protobuf compiler in torch/bin (torch_shm_manager stays), and *test* stubs in torch/lib. Unique transitive dependencies (e.g. hashed libopenblas-*.so) are never touched, ambiguous cases are kept with a warning, and any refused symlink keeps the original file. Set FOX_BUILD_KEEP_DUP_LIBS=1 to ship the tree as Nuitka produced it.
--device rocm* is refused on Windows (no ROCm torch wheels). AMD GPUs can still run GGUF models via llama.cpp HIP:
python packaging/build.py --device cpu --gguf-device rocm72Needs the HIP SDK ($HIP_PATH / hipconfig / %ProgramFiles%\AMD\ROCm); the build drives ROCm clang + Ninja itself (MSVC cannot compile the device code).
Grab a ready-made build from
GitHub releases when one
matches your machine — CPU builds (...-cpu-gguf-...) run anywhere, GPU
builds need that CUDA driver:
| You need | Get it |
|---|---|
| Windows or Linux, no NVIDIA GPU | Download the cpu archive for your OS |
| Windows + older NVIDIA GPU | Download a cu118/cu126 archive if one is listed |
| Linux + NVIDIA GPU, RTX 40/50 series, or CUDA 12.9/13.0 | Not provided — build it locally (one command, see below) |
| AMD GPU | Not provided yet (ROCm builds are coming soon) — use cpu |
Missing your build? Make it yourself — no programming needed, just copy and paste:
# Linux NVIDIA GPU, e.g. CUDA 12.9
python3 packaging/build.py --device cu129Full walkthrough (what to install, which --device for your card, then
run): packaging/build.md.
- GitHub releases:
https://github.com/iamokish/foxreader/releases - Produced by
.github/workflows/build.ymlonv*tags (or manually): native matrix builds → one release with all active archives + auto notes. - Active variants today:
cpu-gguf-linux,cpu-gguf-windows,cu118-gguf-windows,cu126-gguf-windows(e.g.fox-reader-v0.1.0-cpu-gguf-windows-x64.zip). Furthercu129/cu130/Linux-CUDA rows exist commented-out in the workflow — uncomment to enable. - Pick the archive matching your OS + GPU (CPU builds run anywhere; CUDA builds need that CUDA driver; GGUF models run through the bundled llama.cpp with no toolchain needed).
cd frontend
pnpm lint / pnpm lint:fix
pnpm fmt / pnpm fmt:check
pnpm typecheckTests (default suite):
uv run --locked pytest tests/ -v
uv run pytest tests/test_backend.py -v
uv run pytest tests/test_frontend.py -v
uv run pytest tests/test_build.py -vGated/extended checks:
FOX_TEST_FULL=1 FOX_TEST_SLOW=1 FOX_TEST_DIST=1 uv run pytest tests/ -v(FOX_TEST_FULL external tools, FOX_TEST_SLOW slow cases, FOX_TEST_DIST built-distribution checks: executable, DLLs, metadata, OpenCV, launcher single-instance, shutdown, assets/fonts, layout. Legacy tests/manual/ scripts are non-strict xfail until updated.)
foxreader/
src/fox_reader/ FastAPI backend (routes/, services/, translate/, clean/, models/)
frontend/ Preact + Vite UI (src/, templates/, static/)
src/components/ Toolbar, Viewer, Gallery, Translation, Entries, FontSelector,
SVGOverlay, Folder, Appearance, Characters, common
src/layouts/ Basic / Default (+ ActiveLayout switch)
src/workspace/ work/reader modes, strip, overlays
src/themes/ dark / light
src/entries/ geometry, textFit (preview layout), shapes (overlay),
split, inpaint, pageTl
packaging/ Nuitka build, toolchain + launcher (C), llama.cpp helper, asset embed
config/ fox_config.yaml, settings.yaml, user_endpoints.yaml, workspaces.yaml
fonts/ / models/ / cache/ faces / weights / scratch
tests/ backend + frontend + build suites (+ manual legacy scripts)
.github/workflows/ build.yml (matrix + tagged releases)
Pages → [folder] → gallery
→ capture (rect/free/bubble/page) → entries (reading order)
→ OCR (PaddleOCR classic or VL) → editor text
→ translate (DeepL/JPDB/local MTL/custom) → translation panel
→ confirm → styled entries (layers, fonts, clean, geometry)
→ preview (/inpaint/preview → cache/) → save (/inpaint/save-preview or /generate → dest/)
- Backend (
src/fox_reader/app.pylifespan): clearscache/, loads config/settings/endpoints/workspaces, gates local-MTL defaults, configures devices/threads, stands up session/font/progress singletons; serves embeddedstatic/+templates/with no-cache (except/font). - Bubble split rasterises the drawn stroke, rebuilds pieces from the mask, and hands them back as new regions.
- Clean groups jobs by detector/fill so detection runs once per page; typesetting (
typeset.py) then composites layers bottom-up with per-entry plates, auto-fit sizes, symbol fallback, and text-only post-fit transforms. - Frontend overlay (
entries/textFit.ts+entries/shapes.ts) ports that layout (paragraphs, hyphen-only breaks, alignment, binary-search fit) and draws plate/text/clean-FX/outline/number badges in layer order.
| Symptom | Fix |
|---|---|
UNKNOWN ... symlink on pnpm install (Windows) |
Enable Developer Mode, or pnpm install --node-linker=hoisted |
| No C compiler for build | Install VS Build Tools (C++) or MinGW-w64 (winlibs/MSYS2); check python packaging/toolchain.py --which |
| CUDA build fails | Install exactly the toolkit matching --device; Linux CUDA pins GCC (see workflow); Windows CUDA drives MSVC |
| ROCm on Windows refused | Expected — use --device cpu --gguf-device rocm72 + HIP SDK for GGUF |
| GGUF segfault after install | Reinstall with printed CMAKE_ARGS + --no-cache-dir (uv reuses wrong-device wheels otherwise) |
| Page shows setup wizard | Required models missing — complete /setup downloads (needs disk + network) |
| MTL will not load | Check /ml/memory and device pins in /settings; FOX_READER_IGNORE_MTL_MEMORY=1 bypasses the gate for testing |
| Port in use | Change port in config/fox_config.yaml (host/port are file-only) |
| Stale UI after save | Flip Original / Saved; gallery badges mark pages with a dest copy |
- OCR: PaddleOCR (classic + VL)
- Detection: PyTorch bubble segmentation (SafeTensors) + OpenCV (
opencv-contrib-python) - Local MT: Gemma / VNTL via llama.cpp (
llama-cpp-python) - Online MT: DeepL API (official v2 client), JPDB (official v1 client); custom user endpoints supported
- Packaging: Nuitka, Vite, Preact, Hugging Face Hub, FastAPI/Uvicorn, Pydantic
- Fonts: ComicMono (bundled fallback) + NotoSansSymbols2 (render-time symbol fallback) — see
fonts/licences - Models download from Hugging Face on first run (
/setup); GPU wheels from the PyTorch indexes per--device