Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -96,3 +96,16 @@ OPENAI_API_KEY=sk-your-api-key-here
# OPENCMO_MODEL_DEFAULT=gpt-4o
# OPENCMO_MODEL_CMO=gpt-4o
# OPENCMO_MODEL_SEO=gpt-4o-mini

# === Optional: account-scoped knowledge retrieval (RAG) ===
# Install .[rag] or .[all]. Enable per account in Knowledge > Models and retrieval.
# Embedding/rerank URLs, model names and keys are saved per account in that UI.
# OPENCMO_RAG_ENABLED=1 # Master switch; 0 disables retrieval for every account
# OPENCMO_QDRANT_URL=http://127.0.0.1:6333
# OPENCMO_QDRANT_API_KEY= # Shared by the app and the Compose Qdrant service
# OPENCMO_QDRANT_PREFIX=opencmo # Unique for each independent app/database deployment
# OPENCMO_RAG_STORAGE_PATH= # Defaults to knowledge/ next to OPENCMO_DB_PATH
# OPENCMO_RAG_MAX_DOCUMENTS=10000
# OPENCMO_RAG_MAX_CHUNKS=500000
# OPENCMO_KNOWLEDGE_CONCURRENCY=1
# Docker: docker compose --profile rag up -d --build
2 changes: 2 additions & 0 deletions .gitattributes
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
/scripts/qdrant-entrypoint.sh text eol=lf
/tests/fixtures/*.pdf binary
8 changes: 8 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,13 @@ on:
jobs:
backend:
runs-on: ubuntu-latest
services:
qdrant:
image: qdrant/qdrant:v1.19.1
ports:
- 6333:6333
env:
QDRANT__TELEMETRY_DISABLED: "true"
strategy:
matrix:
python-version: ["3.11", "3.12"]
Expand All @@ -29,6 +36,7 @@ jobs:
DATAFORSEO_LOGIN: ""
DATAFORSEO_PASSWORD: ""
XUEQIU_COOKIE: ""
OPENCMO_RAG_TEST_QDRANT: http://127.0.0.1:6333
run: pytest tests/ -v --tb=short

frontend:
Expand Down
6 changes: 6 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -121,6 +121,12 @@ OpenCMO includes a report system inside each project workspace. Open the **Repor
- **Multi-agent pipeline**: human-facing reports use a 6-phase pipeline instead of a single prompt.
- **Graceful fallback**: if the deep pipeline fails, OpenCMO falls back to simpler generation paths so reports stay available.

## Knowledge library (RAG)

Project documents and account-shared cases can now be indexed with recursive parent/child splitting, Qdrant dense retrieval, BM25/title recall, RRF fusion and model reranking. Chat, reports and content drafts share the retrieval service and persist source citations. Supported inputs include historical reports, text PDFs, DOCX, Markdown, TXT and public web pages.

Enable the optional Qdrant service with `docker compose --profile rag up -d`, then configure your account's embedding and rerank APIs under the project's Knowledge library. See [RAG setup, permissions, APIs and evaluation](docs/rag.md). Real model quality must be evaluated with your configured providers.

## Quick Start

OpenCMO works with OpenAI-compatible APIs, including OpenAI, DeepSeek, NVIDIA NIM, Kimi-compatible gateways, and Ollama.
Expand Down
6 changes: 6 additions & 0 deletions README_zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -121,6 +121,12 @@ OpenCMO 已经内置正式报告系统。你可以在项目中打开 **Reports**
- **多智能体管线**:面向人的报告使用 6 阶段管线,而不是单次 prompt。
- **优雅降级**:深度管线失败时,会自动回退到更简单的生成路径,确保报告始终可用。

## 知识库与 RAG

支持把历史报告、客户资料和营销案例导入项目知识库:递归父子切分、Qdrant 语义召回、BM25 与标题召回、RRF 融合、模型重排及可定位原文的引用。聊天、报告和内容草稿共用检索服务;资料默认按项目隔离,可主动共享给同账号其他项目,对外内容另有用途开关。

支持文本 PDF、DOCX、Markdown、TXT、粘贴文本和指定网页。使用 `docker compose --profile rag up -d` 启动可选的 Qdrant 服务,再到项目的“知识库 → 模型与检索设置”配置 Embedding 和 Rerank API。详见 [部署、权限、接口与评测说明](docs/rag.md)。真实模型质量需使用实际配置的服务评测。

## 快速开始

OpenCMO 兼容 OpenAI 协议 API,包括 OpenAI、DeepSeek、NVIDIA NIM、Kimi 兼容网关、Ollama 等。
Expand Down
19 changes: 19 additions & 0 deletions docker-compose.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,26 @@ services:
- opencmo_data:/data
env_file:
- .env
environment:
OPENCMO_QDRANT_URL: http://qdrant:6333
OPENCMO_QDRANT_API_KEY: ${OPENCMO_QDRANT_API_KEY:-}
restart: unless-stopped

qdrant:
image: qdrant/qdrant:v1.19.1
profiles: [rag]
ports:
- "127.0.0.1:6333:6333"
environment:
OPENCMO_QDRANT_API_KEY: ${OPENCMO_QDRANT_API_KEY:-}
QDRANT__TELEMETRY_DISABLED: "true"
volumes:
- qdrant_data:/qdrant/storage
- ./scripts/qdrant-entrypoint.sh:/qdrant/opencmo-entrypoint.sh:ro
entrypoint: ["/bin/sh", "/qdrant/opencmo-entrypoint.sh"]
command: []
restart: unless-stopped

volumes:
opencmo_data:
qdrant_data:
54 changes: 54 additions & 0 deletions docs/rag-validation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,54 @@
# RAG validation

Local verification performed on 2026-09-07/08. No production deployment was performed.

## Automated checks

- Full backend suite: **700 passed, 3 skipped**. The skipped cases are legacy Jinja2 routes.
- RAG suite: **37 passed**, including an actual Qdrant server test.
- Ruff: passed.
- Frontend TypeScript and Vite production build: passed; the existing large-bundle advisory remains.
- Browser: desktop and 390-pixel mobile layouts, document upload, retrieval, model configuration, source preview, centered citation dialog and Unicode/emoji source offsets passed with no page errors.
- GitHub CI passed for Python 3.11, Python 3.12 and the frontend. A subsequent tracing hardening change also passed 123 targeted tests locally.

Tests cover parent/child budgets and exact spans, PDF pages and image-only rejection, DOCX tables, HTML cleanup, URL restrictions, tenant/project isolation, sharing revocation, outbound-use restrictions, tombstones, shadow versions, profile migration failure, rebuilding a missing vector collection, account index leases, orphan reconciliation, persisted citations and non-reuse of revoked evidence.

The broader suite needed test isolation fixes: asynchronous worker completion now waits for completion instead of a 50 ms sleep; route tests avoid unintended live scans; settings tests restore environment defaults. Windows file reads explicitly use UTF-8.

Model calls in automated and browser tests are fixtures. An actual Qdrant engine test is not evidence that the configured embedding model meets semantic quality targets.

The existing Trustabl workflow remains gated on findings in unchanged tools, including timeout detection, dynamic outbound URLs and email-report idempotency. Its file-specific findings did not point to files changed by the RAG implementation. RAG-backed generation explicitly disables hosted SDK tracing; this does not constitute a security audit of every existing tool.

## Capacity result

An isolated native Qdrant 1.19.1 instance and SQLite were tested with synthetic documents/vectors.

| Measurement | Result |
| --- | ---: |
| Documents | 10,000 |
| Child passages | 500,000 |
| Vector dimensions | 1,024 |
| Concurrent retrievals | 5 |
| Sample retrievals | 50 |
| Indexing time | 873.7 seconds |
| Ingestion throughput | 572.3 child passages/second |
| Retrieval P50 | 62 ms |
| Retrieval P95 | 2,828 ms |
| SQLite size | 3,867,803,648 bytes |
| Qdrant status | green |
| Indexed vectors | 499,200; remaining points are searchable in the smaller segment |
| Qdrant peak working set | Approximately 17.1 GiB |

Environment: Windows host with 32 logical processors and approximately 31.3 GiB physical memory; Qdrant search threads configured to 4. The same Qdrant process also hosted small UI/evaluation collections, so the process memory figure is not a per-collection measurement.

Latency combines dense, BM25 and title/entity retrieval, including cold initialization in the sample. It excludes embedding, query-rewrite and rerank API latency. It is not an end-to-end chat latency claim.

An earlier Docker-backed run stopped at approximately 425,000 passages with an underlying storage I/O error. That run is not reported as a capacity pass. The completed run used a separate test disk with adequate free space. The capacity tool now checks estimated local disk requirements and documents the need to check the vector-storage disk separately.

## Evaluation corpus

The committed corpus contains **40 fictional documents and 120 Chinese/English questions**, split into development and holdout sets. The ablation runner completed BM25, dense, fusion and fusion-plus-rerank runs in synthetic wiring mode.

These synthetic scores are deliberately not used to claim Recall@20 or nDCG@10 acceptance for real models. No actual Embedding/Rerank credentials were available during this implementation. **Real model quality remains unverified.**

To evaluate it, configure the provider credentials and run live evaluation as described in [the RAG guide](rag.md). Live evaluation records degradation and fails its quality gate when holdout thresholds are not met.
Loading
Loading