Get started · macOS app · Release 1.8.0 · Contribute
Search your own scientific papers from Claude, Codex or another MCP-compatible assistant.
Ragdoc is a server for the Model Context Protocol (MCP), a standard way for assistants to call external tools. Once connected, your assistant can use Ragdoc to search the articles in your personal literature library, read their content and retrieve supporting passages with source context and provenance information.
Ragdrop is the companion macOS application that prepares that library: import PDFs from Finder or Zotero, compare extracted text with the original, then approve documents for indexing. Ragdoc searches the articles you have added to this library; you ask questions through your connected assistant.
For example: “Which papers in my library compare field measurements with satellite estimates?” Then: “Show me the source passages describing their limitations.” Traceable passages help you check an answer; they do not guarantee scientific accuracy.
Actual English interface with original synthetic examples. The demonstration controls along the bottom are not part of the normal app.
- Bring in PDFs from Finder or Zotero. Queue several articles and detect duplicates using PDF fingerprints.
- Review before indexing. Compare the original PDF with rendered extraction, Markdown source, and extracted tables or figures. Follow page links when the conversion provides matching locators.
- Keep track of each article. Conversion, human review, transfer, indexing and verification remain distinct; failures and pending decisions stay visible.
- Search beyond exact wording. Ragdoc combines lexical and vector retrieval, with optional reranking, source filters and structured evidence results.
- Read the supporting context. Retrieve passages and canonical document snapshots with version hashes and provenance coverage rather than relying on a search snippet alone.
Human review is a separate step. Approval makes an extraction eligible for indexing; successful indexing does not certify its scientific accuracy.
Ragdrop offers light, dark and system appearance under Settings → Appearance. See screenshot provenance and reproduction.
Prepare your library with Ragdrop
flowchart LR
A[PDF / Zotero] --> B[OCR conversion]
B --> C[Human review in Ragdrop]
C -->|Approve and add| D[Your indexed article library]
Search it from your assistant
flowchart LR
A[Claude / Codex / compatible client] -->|MCP request| B[Ragdoc server]
B -->|Search and read| C[Your indexed article library]
C -->|Passages and provenance| B
B -->|MCP results| A
| Your goal | Start here |
|---|---|
| Explore the macOS interface without a backend or API keys | Build the isolated synthetic demo |
| Build Ragdrop and connect your own backend | macOS application guide |
| Run the search backend or connect an MCP client | Backend installation |
| Understand provenance, migration and evaluation | Scientific reliability guide |
Backend setup is required for real imports. Ragdrop requires macOS 14 or later and Swift 6.2 to build. Real imports need a configured SSH-accessible Linux backend and an OCR service credential. The release includes an Apple silicon development build, signed ad hoc and not notarized. Building from source remains available. The demo works without those services.
Ragdrop reads local PDFs and the Zotero Desktop local API. Mistral OCR is the default converter; MinerU is an alternative. These converters upload selected PDFs to their respective services before the human review step. Approved Markdown, metadata and visual artifacts are transferred to your backend over SSH.
Ragdoc stores Chroma vectors, a SQLite FTS5 lexical index and versioned Markdown snapshots on your infrastructure. Voyage AI receives document text during embedding and queries during semantic search. Cohere, when configured, receives the query and candidate passages for reranking. Your MCP client receives the retrieved content; its own model and data handling are separate from Ragdoc.
Lexical retrieval and canonical reads can run locally once the library is indexed.
With alpha=0, search skips Voyage; to avoid reranking API calls, leave Cohere
unconfigured too. Building a new vector index uses Voyage. Model identifiers and
configuration are described in the backend guide.
Storage and indexes stay on your infrastructure; OCR, embeddings and optional
reranking use the services listed above.
- OCR can lose or misread equations, table structure, units and reading order. Check the PDF before relying on extracted content.
- An exact match to a canonical snapshot establishes textual provenance, not scientific truth or correct OCR. PDF page links depend on available, verified locator coverage and may be missing.
- Search scores rank candidates; they are not confidence probabilities. No result does not establish that evidence is absent from the literature.
- Evaluation tooling separates human-reviewed judgments from provisional assistant diagnostics. Current diagnostics do not establish a general retrieval-quality score. See the evaluation requirements.
- Back up source Markdown, canonical snapshots and the index before migration. Index replacement has recovery checks, but is not a database-wide transaction.
See CONTRIBUTING.md for offline checks and contribution scope. Useful next steps include simpler backend onboarding, broader extraction tests, accessibility, and language support. Use original synthetic documents in issues and tests; do not commit articles, personal libraries, credentials or generated indexes.
MIT license. Service credentials and any rights needed to process your own documents remain your responsibility.




