Understand. Translate. Audit. Publish.
A sovereign, production-grade engine for extracting, verifying, translating, and publishing academic dissertations, legal corpuses, and multi-volume books with zero data loss.
Quickstart β’ Architecture β’ Concordance Benchmark β’ CLI Commands β’ Roadmap β’ Changelog
Generic translation utilities treat books as unformatted strings, throwing text into LLMs and outputting mutilated layouts with dropped footnotes, mangled citations, and zero accountability.
DocuVerb is built on an opposing architectural foundation:
Translation is NOT the whole product. Translation is merely one transformation stage inside a verified, audited document lifecycle.
DocuVerb Lifecycle
β
Document Intelligence
(Docling + PyMuPDF + docx)
β
βββββββββββββββββΌββββββββββββββββ
βΌ βΌ βΌ
Structure Translation Terminology
(H1..H6, AST) (Agnostic LLM) (Invariant)
β β β
βββββββββββββββββΌββββββββββββββββ
βΌ
Empirical QA & Concordance
(1:1 Mathematical Verification)
βΌ
Editorial-Grade Publishing
(Qahera UI HTML β’ Word Gutter β’ MD)
| Capability | What Generic Translators Do | What DocuVerb Delivers |
|---|---|---|
| Document Understanding | Strip text into raw txt, destroying hierarchy. | Normalized Document Model (DVND) preserving headings (H1..H6), quotes, footnotes, and metadata. |
| Provider Independence | Hardcoded to OpenAI or one specific local tool. | Agnostic LLMProvider interface decoupled from orchestration. Plug in LM Studio, Ollama, or Cloud APIs seamlessly. |
| Resilience & State | In-memory loops that crash on timeouts. | SQLite WAL Transaction Engine with atomic batch locks and instant resume from power failures. |
| Quality Assurance | Subjective "AI says it looks good". | Empirical Concordance Audit: mathematical proof of segment counts, zero omissions, and glossary fidelity. |
| Editorial Publishing | Raw txt or plain markdown download. | Multi-target Luxury Output: interactive web books (featuring Qahera UI 48px header), Word docs with 26mm gutter binding margins, and clean Markdown. |
DocuVerb was verified by migrating and compiling the complete doctoral dissertation of Doctoral Dissertation Corpus: 82,196 French words across 1,638 structural segments.
================================================================================
DOCUVERB EMPIRICAL AUDIT REPORT
================================================================================
Corpus : Large-Scale Academic Dissertation Corpus
Total Source Segments : 1,638
Translated Segments : 1,638
Missing Segments : 0
Empty Translations : 0
Terminology Invariants : 399 occurrences checked (43 terms)
Completeness Rate : 100.00%
Concordance Score : 100.00% (Bit-for-Bit SHA-256 Verified)
Verification Status : PASSED (VERIFIED)
================================================================================
Want to see DocuVerb in action immediately without needing external documents? Run the built-in demo seeder:
`οΏ½ash
python app/cli.py demo
python app/cli.py audit --project demo_book
python app/cli.py compile --project demo_book --format all ` Open projects/demo_book/output/book_publication.html or οΏ½ook_publication.docx to inspect the results immediately!
# Clone or navigate to the repository
cd c:/wamp64/www/DocuVerb
# Install dependencies
pip install -r requirements.txtpython app/cli.py ingest --project my_book --file "path/to/source.pdf"python app/cli.py audit --project my_book# Compile HTML, Word DOCX (with binding margins), and Markdown in one command
python app/cli.py compile --project my_book --format allOutputs are immediately generated in projects/<project_id>/output/:
book_publication.html: Interactive web edition featuring Qahera UI (48px sticky header, reading progress bar, dark/sepia modes).book_publication.docx: Word document with 26mm right gutter margin for Arabic binding andTraditional Arabictypography.master_complete_book.md: Unified single-file Markdown alongside per-chapter files.
DocuVerb/
βββ core/ # Models (DVND), Database, Project lifecycle, Checkpoints
β βββ models.py # Segment, SegmentType, AuditMetrics, NormalizedDoc
β βββ database.py # SQLite ACID engine (WAL mode, transactions)
β βββ project.py # Isolated project directory & config management
β βββ checkpoints.py # Batch locks, state recovery, and fault tolerance
β
βββ ingestion/ # Document parsers (PyMuPDF, python-docx, TXT, MD)
βββ segmentation/ # Structural classifier (H1..H6, quotes, footnotes)
β
βββ translation/ # Agnostic Provider abstraction & Translation Runner
β βββ provider.py # Abstract LLMProvider interface (generate)
β βββ lmstudio.py # Local LM Studio REST client
β βββ openai_compatible.py # Generic OpenAI v1/chat/completions adapter
β βββ mock_provider.py # Deterministic test provider
β βββ prompt_builder.py # Scholarly prompt synthesis & glossary injection
β βββ runner.py # Batching, exponential backoff retries, and persistence
β
βββ quality/ # Empirical QA, completeness auditor, consistency checker
β βββ completeness.py # Segment counting & zero-omission detection
β βββ terminology.py # Invariant glossary compliance scanner
β βββ auditor.py # Tabular audit report generator
β
βββ publishing/ # Multi-target publication engine
β βββ html.py # Luxury academic HTML + Qahera UI compact reader
β βββ docx.py # Publication Word DOCX with RTL & 26mm binding gutters
β βββ markdown.py # Per-chapter and consolidated master book exporter
β
βββ templates/ # CSS stylesheets (book_editorial_print.css) & templates
βββ importers/ # Non-destructive legacy migration engines
βββ projects/ # Project workspaces (demo_academic, etc.)
βββ app/ # Unified CLI (app/cli.py) & configuration
βββ tests/ # Automated unit and integration test suite (9 tests)
βββ docs/ # Architecture specs, competitive matrix, roadmap
βββ CHANGELOG.md # Single Source of Truth (Keep a Changelog standard)
βββ README.md # You are here
Run the complete regression and unit test suite with zero external dependencies:
python -m unittest discover -s tests -p "test_*.py"Output:
.........
----------------------------------------------------------------------
Ran 9 tests in 4.945s
OK
DocuVerb follows Semantic Versioning 2.0.0. See the comprehensive docs/ROADMAP.md for full milestones.
| Version | Focus Area | Status |
|---|---|---|
v0.1.0 |
Foundation Baseline: Core Engine, WAL DB, Provider Abstraction, Auditor, HTML/DOCX/MD Publishing, CLI. | β Released |
v0.2.0 |
Document Intelligence: IBM Docling integration, Figure/Image extraction & rebinding, Footnote superscript linkage. | β³ Next |
v0.3.0 |
Bilingual Concordance: Side-by-side synchronized reader, Auto-glossary extraction, Segment post-editing CLI. | π Planned |
v0.4.0 |
Prepress Engine: Playwright headless W3C paged print-to-PDF with recto/verso mirror margins. | π Planned |
v1.0.0 |
DocuVerb Studio: Frozen OpenAPI 3.1 contract, FastAPI local backend, Qahera UI Web Control Center. | π― Major Target |
v1.1.0 |
Translation Memory: SQLite FTS5 translation memory, hybrid local/cloud provider routing. | π Future |
Released under the MIT License.
Architected and developed by Alwkala Digital360 for high-stakes academic, literary, and legal translation workflows.