Skip to content

Latest commit

 

History

15 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

pdf-translator

On-device PDF translator for macOS (EN→KO by default, reflow-based). Instead of reproducing the original layout, it restores reading order and pours the translated text into a clean new PDF — the bar is "as readable as Chrome's reader mode." Works fully offline with no API keys. The original handoff spec is preserved in git history (PDF-Translator-Spec.md, initial commit).

PDF ─ has text layer ─▶ extract (PDFKit: lines + bbox + font)
   │                      └▶ geometry-based paragraph assembly ─▶ Vision cell structure for table pages only
   └─ scanned ────────▶ structure (RecognizeDocumentsRequest: paragraphs/titles/tables/lists)
                                   │
             URL & glossary masking ─▶ translate (Apple on-device | Gemini) ─▶ restore
                                   │
                render (block layout + table grids, OS-bundled Korean fonts)

Requirements

  • macOS 26.0+ — uses the direct TranslationSession(installedSource:target:) entry point and RecognizeDocumentsRequest
  • Node 18+, pnpm. No Swift toolchain — the pdf-cli and translate-cli binaries arrive prebuilt (universal, arm64 + x86_64) inside the swiftx packages
  • Translation language packs preinstalled: System Settings > General > Language & Region > Translation Languages (the run stops with install instructions if missing)

Build & Usage

pnpm install       # includes the prebuilt Swift CLIs from swiftx
pnpm build         # TypeScript → dist/

node dist/cli.js input.pdf
Option Description
-o, --output <path> Output path (default: input.ko.pdf)
--source, --target Language pair (default: en → ko)
--pages N[-M] Translate only the given pages (useful for long documents at ~1.5s per paragraph)
--glossary terms.json Pin terminology: {"sub-threshold": "서브 임계"} — consistent terms across the document
--engine apple|gemini Translation engine. gemini requires the GEMINI_API_KEY env var (default: apple, on-device)
--extractor apple|pdfjs Text-layer extraction backend. pdfjs uses unpdf (PDF.js), a cross-platform alternative to the Swift/PDFKit path — no bold detection (see design notes). Default: apple
--renderer swift|js PDF rendering backend. js uses pdf-lib, a cross-platform alternative to the Swift/Core Graphics path. Korean output needs a CJK font via --font (see design notes). Default: swift
--font <path.otf> CJK font (.otf/.ttf) for --renderer js to draw Korean. Also read from PDF_TRANSLATOR_FONT. e.g. Noto Sans KR
--ocr vision|tesseract Scanned-document recognizer. tesseract uses tesseract.js, a cross-platform alternative to the Swift/Vision path — needs local traineddata via --tessdata, no table structure (see design notes). Default: vision
--tessdata <dir> Directory of <lang>.traineddata for --ocr tesseract. Also read from PDF_TRANSLATOR_TESSDATA. e.g. eng.traineddata, kor.traineddata

Cross-platform path (no macOS)

The default pipeline uses Apple frameworks (PDFKit, Vision, TranslationSession, Core Graphics) and needs macOS 26.0+. Each of those stages also has a pure-JS backend, so the four flags below let the tool run on Linux, Windows, or serverless. The swiftx packages still install there — their bundled binaries are macOS-only, but nothing spawns them unless an Apple backend is selected.

pnpm install
pnpm build          # TypeScript → dist/

Text-layer PDF (has selectable text) — extract with PDF.js (pdfjs), translate with Gemini, render with pdf-lib (js):

export GEMINI_API_KEY=...
node dist/cli.js input.pdf \
  --extractor pdfjs \
  --engine gemini \
  --renderer js \
  --font NotoSansKR-Regular.otf

Scanned PDF (image-only, no text layer) — add tesseract.js OCR (--ocr tesseract) with a traineddata directory:

export GEMINI_API_KEY=...
node dist/cli.js scanned.pdf \
  --engine gemini \
  --renderer js \
  --font NotoSansKR-Regular.otf \
  --ocr tesseract \
  --tessdata ./tessdata

Prerequisites you supply once (the tool ships none of these, to stay offline and keep the repo light):

  • GEMINI_API_KEY — the only online piece; the Apple engine is on-device but macOS-only.
  • A CJK font for --font — a fontkit-embeddable single-file .otf/.ttf such as Noto Sans KR. .ttc collections don't embed. Without it, --renderer js falls back to Helvetica and errors on Korean.
  • traineddata for --tessdata — <lang>.traineddata files (e.g. from tessdata_fast); en/ko map to eng/kor.

Flags are independent — mix cross-platform and Apple backends freely (e.g. --extractor pdfjs alone on macOS). Any stage you don't override keeps its Swift/Apple default.

Verification status: the extract (pdfjs) and render (js) backends are measured end-to-end on Linux; the Gemini engine and tesseract OCR are wired and unit-tested but their live API / live OCR passes were not exercised in the build environment (org egress policy blocks the key and traineddata download). Details and limits per backend live in docs/ — unpdf-notes.md, js-renderer-notes.md, tesseract-ocr-notes.md.

Layout

Two Swift CLIs sit at the Apple framework boundaries, installed as prebuilt binaries from swiftx (@cbcruk/pdf-cli, @cbcruk/translate-cli); the Node orchestrator (src/cli.ts) drives them in pipeline order. Every arrow between stages is a JSON contract, so any stage can be swapped independently.

Stage Module Role
Detect pdf-cli info (Swift) does a text layer exist?
Extract pdf-cli extract (Swift) — or extract-pdfjs.ts via --extractor pdfjs per-line text + bbox + font size/bold
— scanned path pdf-cli structure (Swift) — or structure-tesseract.ts via --ocr tesseract — + structure-blocks.ts RecognizeDocumentsRequest paragraphs/titles/tables/lists → blocks directly
Assemble assemble.ts / assemble.utils.ts lines → paragraph/heading/table blocks; header/footer & ToC removal, URL rejoining
Enrich tables enrich-tables.ts + pdf-cli structure (Swift) attach Vision cell structure to geometry-detected tables (table pages only)
Protect protect.ts mask URLs/emails/glossary terms as ⟦U0⟧ tokens; restore after translation
Translate apple-translator.ts | llm-translator.ts engine seam (translator.types.ts): translate-cli (Swift) on-device | Gemini API
Render pdf-cli render (Swift) — or render-js.ts via --renderer js block-by-block layout, table grids, pagination

Shell-out plumbing — process spawn, JSON parsing, binary lookup, exit-code errors — lives in @cbcruk/swift-bridge, not here. Failures surface as SwiftCliError with the CLI's exit code.

How it works

Paragraph assembly — Translation quality comes from paragraph-level context, so grouping lines into paragraphs correctly is the core job. Using document-wide statistics (median line gap, common left-edge alignments, dominant body font size), paragraphs break on vertical gaps (1.35×), font changes, indentation, and bullet markers. Headers/footers, page numbers, rotated text (arXiv stamps), and ToC dot leaders are stripped before assembly.

Tables — A run of 3+ consecutive lines sharing an x-offset away from the body alignment marks a table's position; Vision structure then runs on those pages only to recover cell structure. Only cells containing words are translated, and the renderer draws real grids (borders, automatic column widths, in-cell wrapping). If Vision matching fails, the table degrades gracefully to monospaced text.

Protected spans — URLs, emails, and bare domains are swapped for ⟦U0⟧ tokens before translation and restored afterwards (empirically verified that Apple NMT passes the tokens through untouched). URLs split across lines by typesetting (including right after https:) are rejoined at line and block boundaries. Glossary terms are masked the same way on the Apple path; on the Gemini path they are passed as prompt instructions instead (see below), so the model can inflect them naturally rather than substituting a fixed string.

Gemini context batching — The --engine gemini path translates in batches (≤20 units, split on page boundaries) rather than one segment at a time. Each request carries a read-only CONTEXT block — the section heading plus the preceding source paragraphs — so short headings and boundary paragraphs are disambiguated by their surroundings (this is what makes Abstract → 초록 beat the isolated Abstract → 추상). Context is drawn from source text, not prior translations, so batches stay independent and run concurrently (4 at a time). The glossary is injected as an instruction (term → 번역). Count-mismatch or API failures split the batch in half and retry, falling back to the source string for a single unit that still fails, so a partial failure never sinks the whole document.

Scanned documents — Without a text layer, pages are rasterized at 3x and RecognizeDocumentsRequest returns pre-grouped paragraphs/titles/tables/lists that map directly to blocks (no line assembly).

Design notes (deviations from the spec, with measurements)

  • No SwiftUI hosting needed: the design's biggest unknown (TranslationSession's SwiftUI coupling) is resolved by macOS 26.0+'s init(installedSource:target:). The language-pack download UI is still SwiftUI-only, so packs must be preinstalled
  • Swift CLIs moved out to swiftx: pdf-cli and translate-cli used to live in this repo's swift/ directory and were built locally with pnpm build:swift. They now come from swiftx as prebuilt universal binaries inside npm-shaped tarballs, so this repo no longer needs a Swift toolchain. The JSON contract is unchanged — the same 23-page fixture renders byte-identical text before and after the move
  • OCR via pdf-cli subcommand instead of node-swift in-process binding: assembly is geometry-based and needs per-line bboxes, which the existing vision-ocr module doesn't provide (it returns a flat transcript). The JSON contract is the seam, so switching back is cheap
  • Parallelism doesn't help on-device, but it does over the API: running 2–3 translate-cli processes concurrently takes exactly as long as one — Apple's on-device translation is serialized at the system daemon level (~1.5s per paragraph ceiling). The Gemini path has no such serialization, so it runs batches concurrently — the real throughput fix
  • Render fonts: Helvetica Neue with an Apple SD Gothic Neo cascade. Drawing Latin with SD Gothic alone loses doubled letters (ll → l) on extraction round-trips, and NSFont.systemFont embeds private font names that fall back to Times in other viewers
  • --extractor pdfjs (unpdf) as a cross-platform spike: PDF.js emits text items (runs), not lines, so extract-pdfjs.ts regroups them into lines by hasEOL + baseline jumps — work PDFKit does for free. It shares the same bottom-left coordinate system, so the ExtractedLine contract is a drop-in seam. Verified end-to-end on Linux (extract → assemble → blocks) with no Swift. The catch: unpdf exposes only a generic fontFamily ("sans-serif"), so bold is always false — size-based heading detection survives, bold-only headings don't. Kept behind a flag, not the default, because structure (Vision) and render (Core Graphics) still pin the tool to macOS. See docs/unpdf-notes.md
  • --renderer js (pdf-lib) completes the cross-platform path: with --extractor pdfjs --engine gemini --renderer js, a text-layer PDF translates EN→KO with no macOS. render-js.ts mirrors the Swift renderer's metrics (Letter, 64pt margin, 11/16/9pt fonts, grid tables, bottom-left pagination). Korean needs a fontkit-embeddable single-file CJK font (.otf/.ttf) via --font; Noto Sans KR with subset:true embeds in ~30KB and round-trips exactly (verified end-to-end on Linux, headings/body/tables/URLs). Limits: single font weight (headings differ by size only, no bold cascade), .ttc/unifont not embeddable, and the fallback Helvetica is WinAnsi-only (errors on Korean without a font). See docs/js-renderer-notes.md
  • --ocr tesseract (tesseract.js) extends the cross-platform path to scanned PDFs: structure-tesseract.ts rasterizes pages with unpdf (renderPageAsImage + @napi-rs/canvas), OCRs them, and maps tesseract's blocks/paragraphs to the StructuredPage contract — converting pixel/top-left bboxes to point/bottom-left so blocksFromStructure and its heading detection work unchanged. Rasterization, mapping, and coordinate conversion are verified on Linux; the live OCR pass was not exercised here because org egress policy blocks the traineddata download and no system tesseract is present (like --engine gemini vs the live API). Needs local traineddata via --tessdata (offline-first), and doesn't recover table/list structure (tables/lists come back empty). See docs/tesseract-ocr-notes.md

Known limitations

  • Multi-column (2-column) layouts depend on PDFKit/Vision reading order — not validated
  • Vision cell recognition sometimes merges dense header rows, and tables spanning a page break render as two grids
  • Mistranslation of short headings/cells ("Abstract" → "추상") is a segment-level NMT ceiling on the Apple path — mitigate with --glossary, or switch to --engine gemini, whose context batching translates each segment against its section heading and surrounding text
  • --engine gemini's batching, context assembly, and glossary injection are covered by a mocked-transport test, but the path has not yet been exercised against the live Gemini API (requires a key)

About

On-device PDF translator for macOS that restores reading order and reflows translated text into a clean new PDF (Apple Translation, Vision, PDFKit; offline, no API keys)

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages