Skip to content

feat: collab 批次改动 —— OCR 检索 / 50k 压测 harness / Swagger UI / LLM-judge - #1

Closed
wjs480 wants to merge 9 commits into
FPSZ:collabfrom
wjs480:collab-batch
Closed

wjs480 wants to merge 9 commits into
FPSZ:collabfrom
wjs480:collab-batch

Conversation

@wjs480

@wjs480 wjs480 commented Sep 14, 2026 •

Copy link
Copy Markdown

本 PR 包含 collab 上的 8 个提交:

  • feat(parser): OCR 图片/扫描件可检索(tesseract 集成 + settings 持久化)
  • perf(scale): 压测 harness 支持续跑与断言,新增 50k CI 复跑 workflow
  • feat(server): 新增 Swagger UI 交互式 API 文档页
  • docs(qa): 新增作答层 LLM-judge 首个基线(126 题实跑)
  • fix(core): 修复 /v1 端点重复拼接与回归探活冷启动假失败
  • fix(server): 服务器模式应用 index_filter 索引筛选(对齐桌面端行为)
  • fix: 修复 workspace clippy 警告
  • feat(server): ask 链路补默认级别 trace 事件并关联审计 request-id

说明:已合并上游最新 collab(含 vite ^7.3.6 安全修复),不会回退该修复。

Sourcery 总结

提升协作检索覆盖率、API 易用性、可观测性和大规模质量验证能力,同时增强运行时稳定性。

新功能:

  • 支持基于 OCR 的图像文件、扫描版 PDF 和嵌入式文档图像索引与搜索,并可配置 Tesseract 支持。
  • 添加离线 Swagger UI 页面,用于交互式浏览服务器的 OpenAPI 规范。
  • 添加覆盖 126 个检索回归问题的答案级 LLM 评审基线。
  • 添加可恢复的 50k 规模性能运行,支持竞争断言和定期 CI 报告构件。

Bug 修复:

  • 防止重复添加 /v1 端点片段,并避免模型冷启动期间出现错误的回归健康检查失败。
  • 在服务器模式下应用索引过滤,使其行为与桌面端一致。
  • 将 ask 审计事件和完成日志与请求 ID 关联,并添加默认的 trace 级请求可观测性。

改进:

  • 通过应用设置持久化 OCR 可执行文件配置,同时保留显式环境变量的优先级。
  • 扩展规模验证文档,加入 10k 和 50k 结果以及答案质量解读。

CI:

  • 添加可手动配置并按周定时运行的 50k 性能规模 CI,支持基于阈值的通过/失败验证和报告上传。

文档:

  • 记录大规模检索性能结果、限制,以及受限环境下的 LLM 评审基线。

测试:

  • 添加对 OCR 配置、端点规范化、索引过滤设置、请求 ID 传递和 Swagger UI 资源的覆盖测试。

杂项:

  • 解决工作区中的 Clippy 警告,并保留上游最新的依赖安全更新。
Original summary in English

Summary by Sourcery

Improve collab retrieval coverage, API usability, observability, and large-scale quality validation while hardening runtime behavior.

New Features:

  • Enable OCR-based indexing and search for image files, scanned PDFs, and embedded document images with configurable Tesseract support.
  • Add an offline Swagger UI page for interactively exploring the server's OpenAPI specification.
  • Add an answer-level LLM judge baseline covering 126 retrieval-regression questions.
  • Add resumable 50k-scale performance runs with contention assertions and scheduled CI report artifacts.

Bug Fixes:

  • Prevent duplicate /v1 endpoint segments and avoid false regression health failures during model cold starts.
  • Apply index filtering in server mode to match desktop behavior.
  • Associate ask audit events and completion logs with request IDs and add default trace-level request observability.

Enhancements:

  • Persist OCR executable configuration through application settings while preserving explicit environment-variable precedence.
  • Expand scale-validation documentation with 10k and 50k results and answer-quality interpretation.

CI:

  • Add manually configurable and weekly scheduled 50k performance-scale CI runs with threshold-based pass/fail validation and uploaded reports.

Documentation:

  • Document large-scale retrieval performance results, limitations, and the constrained-environment LLM-judge baseline.

Tests:

  • Add coverage for OCR configuration, endpoint normalization, index-filter settings, request-ID propagation, and Swagger UI assets.

Chores:

  • Resolve workspace Clippy warnings and retain the latest upstream dependency security updates.

FPSZ added 9 commits August 14, 2026 20:06
- ask_handler 完成/失败分别落 info/warn 事件,默认日志级别即可见,
  事件显式携带 request_id 与各阶段耗时(doc_recall/chunk_dense/merge/answer)
- 中间件把 request-id 注入 RequestId 扩展,审计 metadata 增加 request_id 字段
- 新增 extensions 注入端到端测试(http_tests)
- memori-parser: 行结束事件里嵌套 if 折叠为 match guard(行为等价)
- memori-core: format! 参数去掉多余引用(行为等价)
- 修复后 workspace clippy -D warnings 全绿
- AppSettings 新增 index_filter 字段(此前 serde 静默丢弃该配置)
- replace_engine 在 set_indexing_config 后应用筛选配置
- 新增 2 个反序列化单测固定解析行为
- build_openai_url:endpoint 已显式带 /v1(无尾斜杠)时不再追加,
  修复拼成 /v1/v1 导致 404 的问题;新增幂等单测
- 回归 harness 探活:单次 10s 超时改为 3 次重试(每次 20s + 5s 间隔),
  本地模型冷载(实测 12s+)不再被误判为服务不可用
- 报告入库:retrieval_regression_v2_judge_report.json(106 应答题全部判分)
- answer_correct_rate 0.401(受限配置下限:1.5B 作答 + 无 rerank)
- RETRIEVAL_BASELINE_V2.md 新增作答层基线章节,含环境、数字与解读
- GET /api/docs:Swagger UI 页面,数据源 /api/openapi.json
- CSS/JS 静态资源构建期内置(include_str!/include_bytes!),离线可用,
  符合本地优先原则,不引入 utoipa 依赖
- 新增端到端测试:页面 HTML、静态资源 content-type 与非空校验
- ocr.rs:tesseract 外部进程封装(30s 超时强杀、PSM 4 单列排版、
  路径探测 OnceLock 缓存、临时文件原子序号防并发冲突)
- 扫描 PDF:无文本层回退提取 XObject 图片 OCR,支持 DCTDecode 直写
  与 [ASCII85Decode FlateDecode] 链式解码(含 ASCII85/Hex 解码器)
- docx 内嵌图片(word/media)OCR 追加,20MB 大小上限
- 独立图片文件(png/jpg/jpeg)进白名单直接 OCR
- 找不到 tesseract 时全链路静默降级,既有行为不变
- settings.json 新增 ocr_tesseract_path,server/desktop 启动时注入
  环境变量(显式 env 优先);ocr_smoke 调试工具自动读 settings
- v2 套件图片/扫描 4 题:文档召回 0/4 → 4/4
- perf_scale.rs: --start-doc 断点续跑参数(50k 压测期间遗留未提交)

- perf_scale.rs: --max-contention-factor 断言,超标非零退出(供 CI 门槛,目标小于 2x)

- 新增 .github/workflows/perf-scale.yml:手动触发 + 每周日 02:34 UTC 全量 50k,报告 JSON 上传 artifact

- README/PERF_SCALE_50K.md:同步 50k 验证结论与 CI 复跑说明
@sourcery-ai

sourcery-ai Bot commented Sep 14, 2026 •

Copy link
Copy Markdown

审查者指南

本 PR 将 collab 的 OCR 检索、规模性能验证、Swagger UI、LLM-judge 评测和多项服务器/运行时修复整合到同一批次:OCR 通过可选的外部 Tesseract 处理图片及扫描文档并接入 settings,perf harness 支持续跑和 CI 阈值断言,服务器提供离线内置 Swagger UI 与 request-id 关联的 ask 可观测性,同时修正 endpoint、探活和 index_filter 行为并补充基线文档与测试。

OCR 文档索引序列图

sequenceDiagram
    participant Indexer
    participant Parser
    participant Tesseract
    participant SearchIndex

    Indexer->>Parser: extract_document_text(file_path)
    alt Image file
        Parser->>Parser: ocr_available()
        Parser->>Tesseract: ocr_image_file(path)
        Tesseract-->>Parser: Recognized text
    else Scanned PDF
        Parser->>Parser: extract_pdf_images(path)
        Parser->>Tesseract: ocr_image_file(image_path)
        Tesseract-->>Parser: Recognized text
    end
    Parser-->>Indexer: Extracted text
    Indexer->>SearchIndex: Index OCR text
Loading

请求关联的 ask 可观测性序列图

sequenceDiagram
    participant Client
    participant Middleware
    participant AskHandler
    participant Engine
    participant Audit

    Client->>Middleware: HTTP ask request
    Middleware->>Middleware: request_id_trace_middleware(req, next)
    Middleware->>AskHandler: RequestId extension
    AskHandler->>Engine: ask request
    Engine-->>AskHandler: AskResponseStructured
    AskHandler->>Audit: append_audit_event(request_id)
    AskHandler-->>Client: Ask response with x-request-id
    alt Ask failure
        Engine-->>AskHandler: Engine error
        AskHandler-->>Client: ApiError
    end
Loading

文件级变更

变更 详情 文件
为图片和扫描件增加可选的 Tesseract OCR 检索链路,并支持配置持久化。
  • 新增图片文件、DOCX 内嵌图片和无文本层 PDF 的 OCR 回退。
  • 实现 PDF 图片流解码、临时文件管理、超时、大小限制和静默降级。
  • 在桌面端和服务器端从 settings 注入 Tesseract 路径,并扩展图片文件索引类型。
  • 增加 OCR 冒烟工具及基础单元测试。
memori-parser/Cargo.toml
memori-parser/src/lib.rs
memori-parser/src/ocr.rs
memori-parser/examples/ocr_smoke.rs
memori-desktop/src/dto.rs
memori-desktop/src/model_runtime.rs
memori-server/src/dto.rs
memori-server/src/main.rs
memori-vault/src/lib.rs
memori-core/src/model_config.rs
建立可续跑、可断言的 50k 规模性能压测与定期 CI 验证。
  • 为 perf_scale harness 增加 --start-doc 断点续跑和争用系数阈值非零退出。
  • 新增每周定时及手动可调参的 50k GitHub Actions workflow,并上传 JSON 报告 artifact。
  • 补充 1k/10k/50k 性能结果、阶段分解、已知长尾和后续优化方向。
.github/workflows/perf-scale.yml
memori-core/examples/perf_scale.rs
docs/qa/PERF_SCALE_50K.md
README.md
补充作答层 LLM-judge 基线,区分检索命中与最终答案质量。
  • 新增 126 题真实问答和 judge 评测报告,输出逐题判定及理由。
  • 记录受限模型、无 rerank 环境及指标口径,明确结果仅作为下限基线。
docs/qa/RETRIEVAL_BASELINE_V2.md
docs/qa/retrieval_regression_v2_judge_report.json
增强服务器 API 可发现性和请求级可观测性。
  • 将构建期内置的 Swagger UI 静态资源挂载到 /api/docs,并连接 /api/openapi.json。
  • 在请求 extensions 注入 request-id,ask 成功/失败日志和审计事件携带该 ID 及阶段耗时。
  • 增加 Swagger 页面、资源服务和 request-id 注入测试。
memori-server/assets/swagger_ui.html
memori-server/assets/swagger-ui.css
memori-server/assets/swagger-ui-bundle.js
memori-server/assets/swagger-ui-bundle.js.LICENSE.txt
memori-server/assets/swagger-ui-standalone-preset.js
memori-server/assets/swagger-ui-standalone-preset.js.LICENSE.txt
memori-server/src/routes/swagger_ui.rs
memori-server/src/routes/mod.rs
memori-server/src/middleware.rs
memori-server/src/routes/ask.rs
memori-server/src/http_tests.rs
修复服务器和模型运行时配置行为,使其与桌面端保持一致。
  • 避免已包含 /v1 的 endpoint 被重复拼接。
  • 服务器启动时应用 index_filter 配置,补充 settings 反序列化测试。
  • 为嵌入探活增加冷启动重试,避免模型加载延迟造成回归失败。
  • 修复检索输出字段传递及 workspace clippy 警告。
memori-core/src/model_config.rs
memori-core/src/retrieval_output.rs
memori-core/examples/retrieval_regression.rs
memori-server/src/dto.rs
memori-server/src/model_runtime.rs
更新依赖以支持 OCR,并合入上游安全修复。
  • 引入 image 和 flate2 以完成 PNG 编码及 PDF Flate 解码。
  • Cargo.lock 同步依赖锁定;PR 说明保留上游 Vite 7.3.6 安全修复。
memori-parser/Cargo.toml
Cargo.lock

提示与命令

与 Sourcery 交互

  • 触发新的审查: 在 pull request 中评论 @sourcery-ai review。
  • 继续讨论: 直接回复 Sourcery 的审查评论。
  • 从审查评论生成 GitHub issue: 回复审查评论,请 Sourcery 根据该评论创建 issue。你也可以使用 @sourcery-ai issue 回复审查评论,以便从中创建 issue。
  • 生成 pull request 标题: 在 pull request 标题的任意位置写入 @sourcery-ai,即可随时生成标题。你也可以在 pull request 中评论 @sourcery-ai title,以便随时重新生成标题。
  • 生成 pull request 摘要: 在 pull request 正文中任意位置写入 @sourcery-ai summary,即可在指定位置随时生成 PR 摘要。你也可以在 pull request 中评论 @sourcery-ai summary,以便随时重新生成摘要。
  • 生成审查者指南: 在 pull request 中评论 @sourcery-ai guide,即可随时重新生成审查者指南。
  • 解决所有 Sourcery 评论: 在 pull request 中评论 @sourcery-ai resolve,即可解决所有 Sourcery 评论。如果你已经处理完所有评论且不想再看到它们,此功能会非常有用。
  • 忽略所有 Sourcery 审查: 在 pull request 中评论 @sourcery-ai dismiss,即可忽略所有现有的 Sourcery 审查。如果你想从新的审查开始,这项功能尤其有用——别忘了评论 @sourcery-ai review 以触发新的审查!

自定义使用体验

访问你的控制面板,即可:

  • 启用或禁用审查功能,例如 Sourcery 生成的 pull request 摘要、审查者指南等。
  • 更改审查语言。
  • 添加、移除或编辑自定义审查说明。
  • 调整其他审查设置。

获取帮助

Original review guide in English

Reviewer's Guide

本 PR 将 collab 的 OCR 检索、规模性能验证、Swagger UI、LLM-judge 评测和多项服务器/运行时修复整合到同一批次:OCR 通过可选的外部 Tesseract 处理图片及扫描文档并接入 settings,perf harness 支持续跑和 CI 阈值断言,服务器提供离线内置 Swagger UI 与 request-id 关联的 ask 可观测性,同时修正 endpoint、探活和 index_filter 行为并补充基线文档与测试。

Sequence diagram for OCR document indexing

sequenceDiagram
    participant Indexer
    participant Parser
    participant Tesseract
    participant SearchIndex

    Indexer->>Parser: extract_document_text(file_path)
    alt Image file
        Parser->>Parser: ocr_available()
        Parser->>Tesseract: ocr_image_file(path)
        Tesseract-->>Parser: Recognized text
    else Scanned PDF
        Parser->>Parser: extract_pdf_images(path)
        Parser->>Tesseract: ocr_image_file(image_path)
        Tesseract-->>Parser: Recognized text
    end
    Parser-->>Indexer: Extracted text
    Indexer->>SearchIndex: Index OCR text
Loading

Sequence diagram for request-correlated ask observability

sequenceDiagram
    participant Client
    participant Middleware
    participant AskHandler
    participant Engine
    participant Audit

    Client->>Middleware: HTTP ask request
    Middleware->>Middleware: request_id_trace_middleware(req, next)
    Middleware->>AskHandler: RequestId extension
    AskHandler->>Engine: ask request
    Engine-->>AskHandler: AskResponseStructured
    AskHandler->>Audit: append_audit_event(request_id)
    AskHandler-->>Client: Ask response with x-request-id
    alt Ask failure
        Engine-->>AskHandler: Engine error
        AskHandler-->>Client: ApiError
    end
Loading

File-Level Changes

Change Details Files
为图片和扫描件增加可选的 Tesseract OCR 检索链路,并支持配置持久化。
  • 新增图片文件、DOCX 内嵌图片和无文本层 PDF 的 OCR 回退。
  • 实现 PDF 图片流解码、临时文件管理、超时、大小限制和静默降级。
  • 在桌面端和服务器端从 settings 注入 Tesseract 路径,并扩展图片文件索引类型。
  • 增加 OCR 冒烟工具及基础单元测试。
memori-parser/Cargo.toml
memori-parser/src/lib.rs
memori-parser/src/ocr.rs
memori-parser/examples/ocr_smoke.rs
memori-desktop/src/dto.rs
memori-desktop/src/model_runtime.rs
memori-server/src/dto.rs
memori-server/src/main.rs
memori-vault/src/lib.rs
memori-core/src/model_config.rs
建立可续跑、可断言的 50k 规模性能压测与定期 CI 验证。
  • 为 perf_scale harness 增加 --start-doc 断点续跑和争用系数阈值非零退出。
  • 新增每周定时及手动可调参的 50k GitHub Actions workflow,并上传 JSON 报告 artifact。
  • 补充 1k/10k/50k 性能结果、阶段分解、已知长尾和后续优化方向。
.github/workflows/perf-scale.yml
memori-core/examples/perf_scale.rs
docs/qa/PERF_SCALE_50K.md
README.md
补充作答层 LLM-judge 基线,区分检索命中与最终答案质量。
  • 新增 126 题真实问答和 judge 评测报告,输出逐题判定及理由。
  • 记录受限模型、无 rerank 环境及指标口径,明确结果仅作为下限基线。
docs/qa/RETRIEVAL_BASELINE_V2.md
docs/qa/retrieval_regression_v2_judge_report.json
增强服务器 API 可发现性和请求级可观测性。
  • 将构建期内置的 Swagger UI 静态资源挂载到 /api/docs,并连接 /api/openapi.json。
  • 在请求 extensions 注入 request-id,ask 成功/失败日志和审计事件携带该 ID 及阶段耗时。
  • 增加 Swagger 页面、资源服务和 request-id 注入测试。
memori-server/assets/swagger_ui.html
memori-server/assets/swagger-ui.css
memori-server/assets/swagger-ui-bundle.js
memori-server/assets/swagger-ui-bundle.js.LICENSE.txt
memori-server/assets/swagger-ui-standalone-preset.js
memori-server/assets/swagger-ui-standalone-preset.js.LICENSE.txt
memori-server/src/routes/swagger_ui.rs
memori-server/src/routes/mod.rs
memori-server/src/middleware.rs
memori-server/src/routes/ask.rs
memori-server/src/http_tests.rs
修复服务器和模型运行时配置行为,使其与桌面端保持一致。
  • 避免已包含 /v1 的 endpoint 被重复拼接。
  • 服务器启动时应用 index_filter 配置,补充 settings 反序列化测试。
  • 为嵌入探活增加冷启动重试,避免模型加载延迟造成回归失败。
  • 修复检索输出字段传递及 workspace clippy 警告。
memori-core/src/model_config.rs
memori-core/src/retrieval_output.rs
memori-core/examples/retrieval_regression.rs
memori-server/src/dto.rs
memori-server/src/model_runtime.rs
更新依赖以支持 OCR,并合入上游安全修复。
  • 引入 image 和 flate2 以完成 PNG 编码及 PDF Flate 解码。
  • Cargo.lock 同步依赖锁定;PR 说明保留上游 Vite 7.3.6 安全修复。
memori-parser/Cargo.toml
Cargo.lock

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

嘿——我发现了 4 个问题

面向 AI Agent 的提示
请处理本次代码审查中的评论:

## 单独评论

### 评论 1
<location path="memori-parser/src/lib.rs" line_range="658" />
<code_context>
+    if ocr::ocr_available()
</code_context>
<issue_to_address>
**issue (bug_risk):** DOCX 嵌入图片的 OCR 在全新进程中总是失败,因为代码将临时图片写入 `temp_dir()/memori-ocr`,但没有创建该目录;`std::fs::write` 会返回错误,图片因此被跳过。

**触发条件:** 在任何 PDF OCR 创建共享临时目录之前,为包含嵌入式 PNG/JPEG 图片的 DOCX 建立索引。

**建议修复:** 在写入 DOCX 临时文件前创建 `memori-ocr` 目录,或将临时文件创建逻辑集中到一个会创建该目录的辅助函数中。

```suggestion
        let mut ocr_texts = Vec::new();
        let _ = std::fs::create_dir_all(std::env::temp_dir().join("memori-ocr"));
```
</issue_to_address>

### 评论 2
<location path="memori-parser/src/ocr.rs" line_range="82-122" />
<code_context>
+    // spawn + 轮询等待实现超时:tesseract 卡死时强制终止,不拖住索引线程。
</code_context>
<issue_to_address>
**issue (bug_risk):** 子进程的 stdout 已通过管道连接,但直到 `try_wait` 报告 tesseract 已退出后才读取。当 OCR 输出超过操作系统管道缓冲区容量时,tesseract 会在写入 stdout 时阻塞,`try_wait` 永远无法观察到进程退出,最终代码会在 30 秒超时后终止该进程。

**触发条件:** tesseract 生成的识别文本足以填满 stdout 管道时。

**建议修复:** 在等待期间并发排空 stdout,或在工作线程/任务中使用带超时的 `wait_with_output`,并在超时时终止子进程。
</issue_to_address>

### 评论 3
<location path="memori-parser/src/ocr.rs" line_range="87" />
<code_context>
+}
+
+/// 对单张图片执行 OCR(chi_sim 中文)。任何失败返回 None,调用方静默降级。
+pub fn ocr_image_file(path: &Path) -> Option<String> {
+    let tesseract = resolve_tesseract()?;
+    let started = std::time::Instant::now();
+    // spawn + 轮询等待实现超时:tesseract 卡死时强制终止,不拖住索引线程。
+    let mut child = match Command::new(&tesseract)
+        .arg(path)
+        .arg("stdout")
+        .arg("-l")
+        .arg("chi_sim")
+        .arg("--psm")
+        .arg(OCR_PSM)
</code_context>
<issue_to_address>
**issue (bug_risk):** 每次 OCR 调用都要求存在 `chi_sim` traineddata 文件,因此安装了 tesseract 可执行文件但缺少中文语言数据的环境会以非零退出状态失败并返回无文本;因此,所宣传的图片/PDF OCR 功能在仅安装英文语言包的标准 tesseract 环境中不可用。

**触发条件:** tesseract 安装时未包含可选的 `chi_sim` 语言包。

**建议修复:** 使 OCR 语言可配置,或回退到已安装的语言(例如 `eng`),而不是硬编码 `chi_sim`。

```suggestion
        .arg("eng")
```
</issue_to_address>

### 评论 4
<location path="memori-server/src/routes/ask.rs" line_range="5-8" />
<code_context>

 pub(crate) async fn ask_handler(
     State(state): State<ServerState>,
+    Extension(request_id): Extension<RequestId>,
     headers: HeaderMap,
     Json(payload): Json<AskRequest>,
</code_context>
<issue_to_address>
**issue (broader_impact):** 新的失败事件仅在 `ask_structured` 的错误映射中发出。诸如查询为空、策略拒绝、引擎缺失、会话失败或审计/设置错误等提前失败,会在没有预期的 `warn` 事件、也没有具有关联关系的失败审计记录的情况下返回。

**触发条件:** ask 请求在调用 `engine.ask_structured(...).await` 之前失败,或在该调用之外失败时。

**建议修复:** 围绕处理器结果集中记录/记录日志中的 ask 失败,或在每条提前返回的失败路径中添加相同的 request-id/错误/耗时处理。
</issue_to_address>

Sourcery 评估

需要人工审查。 需要优先处理 4 个发现;此外,此更改加入了外部进程 OCR 和新的解析器依赖,并且 OCR 输出在回滚后仍可能保留在搜索索引中。错误的提取结果或配置可能需要清除并重建索引,但其影响是有限且可修复的,并非天然不可逆。

阻塞性发现:memori-parser/src/lib.rs:658、memori-parser/src/ocr.rs:122、memori-parser/src/ocr.rs:87、memori-server/src/routes/ask.rs:8


Sourcery 对开源项目免费——如果您喜欢我们的审查,请考虑分享 ✨
Original comment in English

Hey - I've found 4 issues

Prompt for AI Agents
Please address the comments from this code review:

## Individual Comments

### Comment 1
<location path="memori-parser/src/lib.rs" line_range="658" />
<code_context>
+    if ocr::ocr_available()
</code_context>
<issue_to_address>
**issue (bug_risk):** DOCX embedded-image OCR always fails on a fresh process because the code writes the temporary image under `temp_dir()/memori-ocr` without creating that directory; `std::fs::write` returns an error and the image is skipped.

**Triggers:** When a DOCX containing embedded PNG/JPEG images is indexed before any PDF OCR has created the shared temporary directory.

**Suggested fix:** Create the `memori-ocr` directory before writing DOCX temporary files, or centralize temporary-file creation in a helper that creates it.

```suggestion
        let mut ocr_texts = Vec::new();
        let _ = std::fs::create_dir_all(std::env::temp_dir().join("memori-ocr"));
```
</issue_to_address>

### Comment 2
<location path="memori-parser/src/ocr.rs" line_range="82-122" />
<code_context>
+    // spawn + 轮询等待实现超时:tesseract 卡死时强制终止,不拖住索引线程。
</code_context>
<issue_to_address>
**issue (bug_risk):** The child process's stdout is piped but is not read until after `try_wait` reports that tesseract exited. When OCR output exceeds the OS pipe buffer, tesseract blocks while writing stdout, `try_wait` never observes exit, and the code kills the process at the 30-second timeout.

**Triggers:** When tesseract produces enough recognized text to fill the stdout pipe.

**Suggested fix:** Drain stdout concurrently while waiting, or use `wait_with_output` in a worker thread/task with a timeout and kill the child on timeout.
</issue_to_address>

### Comment 3
<location path="memori-parser/src/ocr.rs" line_range="87" />
<code_context>
+}
+
+/// 对单张图片执行 OCR(chi_sim 中文)。任何失败返回 None,调用方静默降级。
+pub fn ocr_image_file(path: &Path) -> Option<String> {
+    let tesseract = resolve_tesseract()?;
+    let started = std::time::Instant::now();
+    // spawn + 轮询等待实现超时:tesseract 卡死时强制终止,不拖住索引线程。
+    let mut child = match Command::new(&tesseract)
+        .arg(path)
+        .arg("stdout")
+        .arg("-l")
+        .arg("chi_sim")
+        .arg("--psm")
+        .arg(OCR_PSM)
</code_context>
<issue_to_address>
**issue (bug_risk):** Every OCR invocation requires the `chi_sim` traineddata file, so installations that have the tesseract executable but lack Chinese language data fail with a nonzero exit status and return no text; the advertised image/PDF OCR feature is therefore unavailable on standard English-only tesseract installations.

**Triggers:** When tesseract is installed without the optional `chi_sim` language package.

**Suggested fix:** Make the OCR language configurable or fall back to an installed language such as `eng` instead of hardcoding `chi_sim`.

```suggestion
        .arg("eng")
```
</issue_to_address>

### Comment 4
<location path="memori-server/src/routes/ask.rs" line_range="5-8" />
<code_context>

 pub(crate) async fn ask_handler(
     State(state): State<ServerState>,
+    Extension(request_id): Extension<RequestId>,
     headers: HeaderMap,
     Json(payload): Json<AskRequest>,
</code_context>
<issue_to_address>
**issue (broader_impact):** The new failure event is emitted only inside the `ask_structured` error mapping. Early failures such as an empty query, policy rejection, missing engine, session failure, or audit/setup errors return without the promised `warn` event and without a correlated failure audit record.

**Triggers:** When an ask request fails before or outside the `engine.ask_structured(...).await` call.

**Suggested fix:** Centralize ask failure recording/logging around the handler result, or add the same request-id/error/elapsed handling to each early-return failure path.
</issue_to_address>

Sourcery assessment

Needs a human reviewer. 4 findings to address first, and the change adds external-process OCR and new parser dependencies, and OCR output can remain persisted in the search index after a revert. A bad extraction or configuration can require clearing and rebuilding the index, but the impact is bounded and repairable rather than inherently irreversible.

Blocking findings: memori-parser/src/lib.rs:658, memori-parser/src/ocr.rs:122, memori-parser/src/ocr.rs:87, memori-server/src/routes/ask.rs:8


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

Comment thread memori-parser/src/lib.rs
.then(|| name.to_string())
})
.collect();
let mut ocr_texts = Vec::new();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

issue (bug_risk): DOCX 嵌入图片的 OCR 在全新进程中总是失败,因为代码将临时图片写入 temp_dir()/memori-ocr,但没有创建该目录;std::fs::write 会返回错误,图片因此被跳过。

触发条件: 在任何 PDF OCR 创建共享临时目录之前,为包含嵌入式 PNG/JPEG 图片的 DOCX 建立索引。

建议修复: 在写入 DOCX 临时文件前创建 memori-ocr 目录,或将临时文件创建逻辑集中到一个会创建该目录的辅助函数中。

Suggested change
let mut ocr_texts = Vec::new();
let mut ocr_texts = Vec::new();
let _ = std::fs::create_dir_all(std::env::temp_dir().join("memori-ocr"));
Original comment in English

issue (bug_risk): DOCX embedded-image OCR always fails on a fresh process because the code writes the temporary image under temp_dir()/memori-ocr without creating that directory; std::fs::write returns an error and the image is skipped.

Triggers: When a DOCX containing embedded PNG/JPEG images is indexed before any PDF OCR has created the shared temporary directory.

Suggested fix: Create the memori-ocr directory before writing DOCX temporary files, or centralize temporary-file creation in a helper that creates it.

Suggested change
let mut ocr_texts = Vec::new();
let mut ocr_texts = Vec::new();
let _ = std::fs::create_dir_all(std::env::temp_dir().join("memori-ocr"));

Comment thread memori-parser/src/ocr.rs
Comment on lines +82 to +122
// spawn + 轮询等待实现超时:tesseract 卡死时强制终止,不拖住索引线程。
let mut child = match Command::new(&tesseract)
.arg(path)
.arg("stdout")
.arg("-l")
.arg("chi_sim")
.arg("--psm")
.arg(OCR_PSM)
.stdout(std::process::Stdio::piped())
.spawn()
{
Ok(child) => child,
Err(_) => {
warn!(path = %path.display(), "tesseract 启动失败,跳过 OCR");
return None;
}
};
let deadline = Duration::from_secs(OCR_TIMEOUT_SECS);
loop {
if let Ok(Some(status)) = child.try_wait() {
if !status.success() {
warn!(
path = %path.display(),
status = %status,
"tesseract 识别失败,跳过 OCR"
);
return None;
}
break;
}
if started.elapsed() > deadline {
let _ = child.kill();
let _ = child.wait();
warn!(
path = %path.display(),
timeout_secs = OCR_TIMEOUT_SECS,
"OCR 超时已终止"
);
return None;
}
std::thread::sleep(Duration::from_millis(100));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

issue (bug_risk): 子进程的 stdout 已通过管道连接,但直到 try_wait 报告 tesseract 已退出后才读取。当 OCR 输出超过操作系统管道缓冲区容量时,tesseract 会在写入 stdout 时阻塞,try_wait 永远无法观察到进程退出,最终代码会在 30 秒超时后终止该进程。

触发条件: tesseract 生成的识别文本足以填满 stdout 管道时。

建议修复: 在等待期间并发排空 stdout,或在工作线程/任务中使用带超时的 wait_with_output,并在超时时终止子进程。

Original comment in English

issue (bug_risk): The child process's stdout is piped but is not read until after try_wait reports that tesseract exited. When OCR output exceeds the OS pipe buffer, tesseract blocks while writing stdout, try_wait never observes exit, and the code kills the process at the 30-second timeout.

Triggers: When tesseract produces enough recognized text to fill the stdout pipe.

Suggested fix: Drain stdout concurrently while waiting, or use wait_with_output in a worker thread/task with a timeout and kill the child on timeout.

Comment thread memori-parser/src/ocr.rs
.arg(path)
.arg("stdout")
.arg("-l")
.arg("chi_sim")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

issue (bug_risk): 每次 OCR 调用都要求存在 chi_sim traineddata 文件,因此安装了 tesseract 可执行文件但缺少中文语言数据的环境会以非零退出状态失败并返回无文本;因此,所宣传的图片/PDF OCR 功能在仅安装英文语言包的标准 tesseract 环境中不可用。

触发条件: tesseract 安装时未包含可选的 chi_sim 语言包。

建议修复: 使 OCR 语言可配置,或回退到已安装的语言(例如 eng),而不是硬编码 chi_sim。

Suggested change
.arg("chi_sim")
.arg("eng")
Original comment in English

issue (bug_risk): Every OCR invocation requires the chi_sim traineddata file, so installations that have the tesseract executable but lack Chinese language data fail with a nonzero exit status and return no text; the advertised image/PDF OCR feature is therefore unavailable on standard English-only tesseract installations.

Triggers: When tesseract is installed without the optional chi_sim language package.

Suggested fix: Make the OCR language configurable or fall back to an installed language such as eng instead of hardcoding chi_sim.

Suggested change
.arg("chi_sim")
.arg("eng")

Comment on lines +5 to 8
Extension(request_id): Extension<RequestId>,
headers: HeaderMap,
Json(payload): Json<AskRequest>,
) -> Result<Json<AskResponseStructured>, ApiError> {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

issue (broader_impact): 新的失败事件仅在 ask_structured 的错误映射中发出。诸如查询为空、策略拒绝、引擎缺失、会话失败或审计/设置错误等提前失败,会在没有预期的 warn 事件、也没有具有关联关系的失败审计记录的情况下返回。

触发条件: ask 请求在调用 engine.ask_structured(...).await 之前失败,或在该调用之外失败时。

建议修复: 围绕处理器结果集中记录 ask 失败并记录日志,或在每条提前返回的失败路径中添加相同的 request-id/错误/耗时处理。

Original comment in English

issue (broader_impact): The new failure event is emitted only inside the ask_structured error mapping. Early failures such as an empty query, policy rejection, missing engine, session failure, or audit/setup errors return without the promised warn event and without a correlated failure audit record.

Triggers: When an ask request fails before or outside the engine.ask_structured(...).await call.

Suggested fix: Centralize ask failure recording/logging around the handler result, or add the same request-id/error/elapsed handling to each early-return failure path.

@FPSZ FPSZ left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

感谢这批工作量不小的提交。Swagger UI、trace/request-id、index_filter 对齐、clippy 这几块质量不错。但 OCR 部分(69304d8)目前端到端不工作,另外有两处会往知识库写脏数据 / 拖住 async runtime 的问题,所以先打回。下面按严重度列,都带了文件行号。


🔴 阻断项(必须修)

1. 图片 OCR 在真实索引链路里没接上 — 功能实际不生效

memori-vault/src/lib.rs:29 把 png/jpg/jpeg 加进了 SUPPORTED_CONTENT_EXTENSIONS,memori-parser/src/lib.rs:573 也正确 dispatch 到了 ocr::ocr_image_file。

但 memori-core/src/indexing_rebuild.rs:229 的 read_document_text 的二进制格式白名单没有同步加图片:

if matches!(ext.as_str(), "docx"|"pdf"|"pptx"|"xlsx"|"doc"|"ppt"|"xls") {
    // → memori_parser::extract_document_text
} else {
    tokio::fs::read_to_string(path).await   // ← png/jpg 落到这里
}

结果是图片被当 UTF-8 文本读 → InvalidData → OCR 永远不会执行;而且每张图片都会写一次 indexing_runtime.last_error,UI 上会持续报错。

请把 "png" | "jpg" | "jpeg" 加进 read_document_text 的 matches 分支,并补一个走真实索引链路(而不是只调 parser)的集成测试。

2. tesseract stdout 管道死锁 — 文字越多越必然触发

memori-parser/src/ocr.rs:101:stdout 用了 Stdio::piped(),但轮询循环只做 try_wait() + sleep(100ms),全程不读管道,要等子进程退出后才 wait_with_output()。

管道缓冲区一满(Windows 匿名管道通常 4KB,中文 UTF-8 约 1300 字就满),tesseract 会阻塞在 write 上永不退出 → try_wait() 永远返回 None → 30s 后被超时 kill → 返回 None。

也就是说文字密集的页面 100% 失败,而这恰恰是 OCR 唯一有价值的场景;同时每次都要白白烧掉 30 秒超时。

建议改成起一个线程读 stdout,或直接用 wait_with_output() 配合独立的超时看门狗线程;别在不读管道的情况下轮询 try_wait。


🟠 严重

3. DOCX 内嵌图 OCR 临时目录没创建

memori-parser/src/lib.rs:674 往 temp_dir()/memori-ocr 写文件,但只有 PDF 那条路径调了 create_dir_all。干净机器上 DOCX 内嵌图 OCR 会全部静默失败。

4. PDF OCR 兜底同步阻塞 tokio worker

memori-parser/src/lib.rs:720 的 PDF OCR 回退可以从 build_reference_excerpt → build_citations 触发,而这条链是同步跑在 async engine_search 里的。一个被引用的扫描件最多会用 std::thread::sleep 轮询阻塞一个 tokio worker 达 20×30s,而且每次 ask 都重来一遍(没有缓存)。

问答链路上不应该做 OCR。建议只在索引期做 OCR 并把结果落库,ask 期直接读已抽取文本。

5. 解压无上限,可 OOM

memori-parser/src/ocr.rs:249 的尺寸上限只卡了压缩流,ZlibDecoder::read_to_end 是无界的。一个 20MB 的 flate 图可以膨胀到几十 GB,直接把索引进程打爆。请用 take(limit) 限制解压后大小。

6. 忽略 /BitsPerComponent,会把 OCR 噪声当文档内容索引

memori-parser/src/ocr.rs:335 没有校验 /BitsPerComponent,16-bit 图能通过长度检查并被当成 8-bit 重新解释,生成一张乱码 PNG。OCR 出来的噪声文本会被当作真实文档内容写进知识库。

这条对本项目是原则性问题——证据链的可信度是核心卖点,宁可跳过也不能索引噪声。请显式只接受 BitsPerComponent == 8,其余跳过。

7. 无文本层 PDF 的返回语义变了

memori-parser/src/lib.rs:721 由 Some("") 改成了 None。没装 tesseract 的环境下,每个扫描件都会报 "文件读取失败(可能被占用)" 并污染 last_error——这个错误信息会误导用户去查文件占用。建议保持 Some(""),或给出"无文本层且未配置 OCR"的专门错误分类。


🟡 次要

  • memori-parser/src/ocr.rs:41 — tesseract 路径锁进 OnceLock 且 apply_ocr_path_to_env 不覆盖已有 env,改错了配置必须重启整个应用才能生效;env 为空字符串时还会顺带禁掉 PATH 回退。
  • .github/workflows/perf-scale.yml:66 — artifact 上传缺 if: always(),恰恰在 --max-contention-factor 断言 FAIL 时丢掉报告,而那正是唯一需要看报告的时候。
  • memori-parser/src/lib.rs:665 — read_to_end 失败时会穿过尺寸护栏,对截断的字节做 OCR,应该直接跳过。
  • memori-core/examples/perf_scale.rs:377 — --start-doc 低于实际进度时 resume 会重复计数,虚报 indexed_chunks 和 chunks/s,压测数字会偏乐观。

✅ 这些部分没问题

  • request-id extension 在 build_router 里层级正确,Extension<RequestId> 保证存在。
  • server 端 index_filter 的应用语义与桌面端一致。
  • build_openai_url 的 /v1 去重逻辑正确,尾斜杠和 # 的场景都覆盖到了。
  • finish_full_rebuild 只改元数据,压测 resume 路径不会误删数据。
  • Swagger UI(离线资源)、LLM-judge 126 题基线文档、clippy 修复都 OK。

建议的处理方式

把这个 PR 拆开:

  1. 可以先合:083c16e Swagger UI、cac8442 trace/request-id、f764abe index_filter、8d7eaa2 clippy、3ccae63 /v1 修复、3fcbbbd LLM-judge 文档。这几块独立且干净。
  2. OCR(69304d8)单独开一个 PR 返工,至少修掉 #1、#2、#5、#6,并补一个跑真实索引链路的端到端测试(现在的测试只覆盖到 parser 层,所以 #1 才没被发现)。
  3. 压测 harness(af6ba4f) 修掉 if: always() 和 resume 计数即可合。

另外提醒一下:docs/qa/retrieval_regression_v2_judge_report.json 有 7223 行,vendored 的 swagger-ui 压缩产物也一并入库了。judge 报告建议只留摘要,明细走 CI artifact;swagger-ui 资源如果确定要 vendor,请在 README 或该目录下注明版本号和来源,方便后续跟安全更新。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants