Skip to content
alwkalaPublic

Latest commit

Β 

History

1 Commit

Folders and files

Repository files navigation

πŸ“š DocuVerb (DocuVerbatim)

Modular Document Intelligence, Translation & Publishing Engine

Release Tests Concordance Python License

Understand. Translate. Audit. Publish.

A sovereign, production-grade engine for extracting, verifying, translating, and publishing academic dissertations, legal corpuses, and multi-volume books with zero data loss.

Quickstart β€’ Architecture β€’ Concordance Benchmark β€’ CLI Commands β€’ Roadmap β€’ Changelog


πŸ’‘ The DocuVerb Philosophy

Generic translation utilities treat books as unformatted strings, throwing text into LLMs and outputting mutilated layouts with dropped footnotes, mangled citations, and zero accountability.

DocuVerb is built on an opposing architectural foundation:

Translation is NOT the whole product. Translation is merely one transformation stage inside a verified, audited document lifecycle.

                     DocuVerb Lifecycle
                             β”‚
                  Document Intelligence
                (Docling + PyMuPDF + docx)
                             β”‚
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β–Ό               β–Ό               β–Ό
         Structure      Translation     Terminology
        (H1..H6, AST)   (Agnostic LLM)   (Invariant)
             β”‚               β”‚               β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             β–Ό
                 Empirical QA & Concordance
                (1:1 Mathematical Verification)
                             β–Ό
                 Editorial-Grade Publishing
             (Qahera UI HTML β€’ Word Gutter β€’ MD)

✨ Key Pillars & Differentiators

Capability What Generic Translators Do What DocuVerb Delivers
Document Understanding Strip text into raw txt, destroying hierarchy. Normalized Document Model (DVND) preserving headings (H1..H6), quotes, footnotes, and metadata.
Provider Independence Hardcoded to OpenAI or one specific local tool. Agnostic LLMProvider interface decoupled from orchestration. Plug in LM Studio, Ollama, or Cloud APIs seamlessly.
Resilience & State In-memory loops that crash on timeouts. SQLite WAL Transaction Engine with atomic batch locks and instant resume from power failures.
Quality Assurance Subjective "AI says it looks good". Empirical Concordance Audit: mathematical proof of segment counts, zero omissions, and glossary fidelity.
Editorial Publishing Raw txt or plain markdown download. Multi-target Luxury Output: interactive web books (featuring Qahera UI 48px header), Word docs with 26mm gutter binding margins, and clean Markdown.

πŸ›οΈ Empirical Concordance Benchmark

DocuVerb was verified by migrating and compiling the complete doctoral dissertation of Doctoral Dissertation Corpus: 82,196 French words across 1,638 structural segments.

================================================================================
                    DOCUVERB EMPIRICAL AUDIT REPORT
================================================================================
  Corpus                  : Large-Scale Academic Dissertation Corpus
  Total Source Segments   : 1,638
  Translated Segments     : 1,638
  Missing Segments        : 0
  Empty Translations      : 0
  Terminology Invariants  : 399 occurrences checked (43 terms)
  Completeness Rate       : 100.00%
  Concordance Score       : 100.00% (Bit-for-Bit SHA-256 Verified)
  Verification Status     : PASSED (VERIFIED)
================================================================================

πŸš€ Quickstart

⚑ 5-Second Test-Drive (Instant Demo)

Want to see DocuVerb in action immediately without needing external documents? Run the built-in demo seeder:

`οΏ½ash

1. Initialize realistic scholarly demo project (philosophical & translation theory corpus)

python app/cli.py demo

2. Run Empirical QA Concordance Audit

python app/cli.py audit --project demo_book

3. Compile into publication formats (HTML with Qahera UI, Word DOCX with gutter margins, Markdown)

python app/cli.py compile --project demo_book --format all ` Open projects/demo_book/output/book_publication.html or οΏ½ook_publication.docx to inspect the results immediately!


πŸ“₯ Production Ingestion Workflow

1. Requirements & Setup

# Clone or navigate to the repository
cd c:/wamp64/www/DocuVerb

# Install dependencies
pip install -r requirements.txt

2. Ingest a Document

python app/cli.py ingest --project my_book --file "path/to/source.pdf"

3. Run the Empirical Concordance Audit

python app/cli.py audit --project my_book

4. Compile Publication Outputs

# Compile HTML, Word DOCX (with binding margins), and Markdown in one command
python app/cli.py compile --project my_book --format all

Outputs are immediately generated in projects/<project_id>/output/:

  • book_publication.html: Interactive web edition featuring Qahera UI (48px sticky header, reading progress bar, dark/sepia modes).
  • book_publication.docx: Word document with 26mm right gutter margin for Arabic binding and Traditional Arabic typography.
  • master_complete_book.md: Unified single-file Markdown alongside per-chapter files.

πŸ“¦ Project Architecture

DocuVerb/
β”œβ”€β”€ core/                  # Models (DVND), Database, Project lifecycle, Checkpoints
β”‚   β”œβ”€β”€ models.py          # Segment, SegmentType, AuditMetrics, NormalizedDoc
β”‚   β”œβ”€β”€ database.py        # SQLite ACID engine (WAL mode, transactions)
β”‚   β”œβ”€β”€ project.py         # Isolated project directory & config management
β”‚   └── checkpoints.py     # Batch locks, state recovery, and fault tolerance
β”‚
β”œβ”€β”€ ingestion/             # Document parsers (PyMuPDF, python-docx, TXT, MD)
β”œβ”€β”€ segmentation/          # Structural classifier (H1..H6, quotes, footnotes)
β”‚
β”œβ”€β”€ translation/           # Agnostic Provider abstraction & Translation Runner
β”‚   β”œβ”€β”€ provider.py        # Abstract LLMProvider interface (generate)
β”‚   β”œβ”€β”€ lmstudio.py        # Local LM Studio REST client
β”‚   β”œβ”€β”€ openai_compatible.py # Generic OpenAI v1/chat/completions adapter
β”‚   β”œβ”€β”€ mock_provider.py   # Deterministic test provider
β”‚   β”œβ”€β”€ prompt_builder.py  # Scholarly prompt synthesis & glossary injection
β”‚   └── runner.py          # Batching, exponential backoff retries, and persistence
β”‚
β”œβ”€β”€ quality/               # Empirical QA, completeness auditor, consistency checker
β”‚   β”œβ”€β”€ completeness.py    # Segment counting & zero-omission detection
β”‚   β”œβ”€β”€ terminology.py     # Invariant glossary compliance scanner
β”‚   └── auditor.py         # Tabular audit report generator
β”‚
β”œβ”€β”€ publishing/            # Multi-target publication engine
β”‚   β”œβ”€β”€ html.py            # Luxury academic HTML + Qahera UI compact reader
β”‚   β”œβ”€β”€ docx.py            # Publication Word DOCX with RTL & 26mm binding gutters
β”‚   └── markdown.py        # Per-chapter and consolidated master book exporter
β”‚
β”œβ”€β”€ templates/             # CSS stylesheets (book_editorial_print.css) & templates
β”œβ”€β”€ importers/             # Non-destructive legacy migration engines
β”œβ”€β”€ projects/              # Project workspaces (demo_academic, etc.)
β”œβ”€β”€ app/                   # Unified CLI (app/cli.py) & configuration
β”œβ”€β”€ tests/                 # Automated unit and integration test suite (9 tests)
β”œβ”€β”€ docs/                  # Architecture specs, competitive matrix, roadmap
β”œβ”€β”€ CHANGELOG.md           # Single Source of Truth (Keep a Changelog standard)
└── README.md              # You are here

πŸ§ͺ Test Suite Execution

Run the complete regression and unit test suite with zero external dependencies:

python -m unittest discover -s tests -p "test_*.py"

Output:

.........
----------------------------------------------------------------------
Ran 9 tests in 4.945s

OK

πŸ—ΊοΈ SemVer Development Roadmap

DocuVerb follows Semantic Versioning 2.0.0. See the comprehensive docs/ROADMAP.md for full milestones.

Version Focus Area Status
v0.1.0 Foundation Baseline: Core Engine, WAL DB, Provider Abstraction, Auditor, HTML/DOCX/MD Publishing, CLI. βœ… Released
v0.2.0 Document Intelligence: IBM Docling integration, Figure/Image extraction & rebinding, Footnote superscript linkage. ⏳ Next
v0.3.0 Bilingual Concordance: Side-by-side synchronized reader, Auto-glossary extraction, Segment post-editing CLI. πŸ“‹ Planned
v0.4.0 Prepress Engine: Playwright headless W3C paged print-to-PDF with recto/verso mirror margins. πŸ“‹ Planned
v1.0.0 DocuVerb Studio: Frozen OpenAPI 3.1 contract, FastAPI local backend, Qahera UI Web Control Center. 🎯 Major Target
v1.1.0 Translation Memory: SQLite FTS5 translation memory, hybrid local/cloud provider routing. πŸ”­ Future

πŸ“„ License & Attribution

Released under the MIT License.
Architected and developed by Alwkala Digital360 for high-stakes academic, literary, and legal translation workflows.