Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 1 addition & 6 deletions docs/ir-and-source-map.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,12 +46,7 @@ content.json -> element_ir.json -> source_map.json -> notes.md / GUI / coverage

## 构建时机

IR 会在两个阶段写入:

1. `export_content` 阶段写基础 IR,供后续 prompt、source map 和 coverage 使用。
2. coverage 阶段结束后刷新最终 IR,合入 `covered`、`missing`、`marker-only` 等实际状态。

这样最终 `element_ir.json` 不是只停留在前置状态,而是反映生成后的覆盖结果。
构建过程中,prompt、coverage 和 source map 会直接从当前 `Deck` 构造需要的元素视图。coverage 阶段结束后写入一次最终 `element_ir.json`,合入 `covered`、`missing`、`marker-only` 等实际状态,避免重复生成中间文件。

## source_map.json

Expand Down
12 changes: 8 additions & 4 deletions docs/pipeline.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,15 +8,19 @@ Ingest -> Understand -> Write -> Guard -> Export

底层模块可以保持细粒度,方便缓存、调试和局部刷新;用户侧和 LLM 工作流应该看到清楚的阶段边界。

实现中由 `BUILD_PHASES` 按这五个阶段组织步骤。构建开始时会根据 preset 和选项排除未启用的 OCR、Vision、图裁剪等步骤,并把实际计划写入 `progress.json` 的 `planned_stages`;`current_phase` 表示当前产品阶段。逐步耗时仍保留在 `run_summary.json`,方便定位慢点。

## 阶段总览

| 阶段 | 目标 | 典型产物 |
| --- | --- | --- |
| Ingest | 把 PPT/PDF 变成稳定、可追溯、可复现的结构化材料。 | `content.json`、`element_ir.json`、`source_map.json`、截图、图片资产、parser adapter |
| Understand | 理解课件主题、章节结构、页面角色、图表含义和关键元素。 | `deck_understanding.json`、`page_understanding.json`、`sections.json`、`deck_brief.json`、`semantic_layout.json`、`table_understanding.json`、`figure_grounding.json` |
| Ingest | 解析 PPT/PDF,提取页面元素和资源。 | 内存中的 `Deck`、截图、图片资产 |
| Understand | 理解课件主题、章节结构、页面角色、图表含义和关键元素。 | `content.json`、`page_modalities.json`、`deck_understanding.json`、`page_understanding.json`、`sections.json`、`deck_brief.json`、`semantic_layout.json`、`table_understanding.json`、`figure_grounding.json`、`content_guard.json` |
| Write | 生成可读学习笔记,而不是机械逐页搬运。 | `notes.md`、`page_notes.json`、`weave_report.json`、`teaching_enrichment.json` |
| Guard | 检查是否漏掉关键内容、是否有来源、是否像讲义。 | `coverage.json`、`coverage.md`、`content_guard.json`、`quality_report.json` |
| Export | 输出阅读和复习材料。 | `notes.toc.md`、`notes.docx`、`notes.pdf`、`notes.tex`、`review.md`、`exam.html` |
| Guard | 检查是否漏掉关键内容、是否有来源、是否像讲义。 | `coverage.json`、`coverage.md`、`element_ir.json`、`source_map.json`、`quality_report.json` |
| Export | 输出阅读材料和构建摘要。 | `notes.toc.md`、`notes.docx`、`notes.pdf`、`notes.tex`、`run_summary.json` |

`lecture` 的教学补充采用按章节判断:整合后的章节稿已有足够正文,并包含例子、易错点和自测线索时,跳过额外模型调用;`force` 仍会执行补充。

## 什么不交给 LLM

Expand Down
1 change: 1 addition & 0 deletions gui/README_GUI.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ SlideNote Studio is a Streamlit interface for `python -m slidenote build` and `p
- Enter Text / Vision / OCR API keys on the page. Keys are passed only through the child-process environment, not command-line flags.
- Select extra exports: Markdown ZIP, TOC Markdown, Word, PDF, or LaTeX.
- Keep progress, ETA, Doctor readiness, usage, and cost details in compact diagnostics panels.
- Saved page modality corrections apply to the next build of the same source file; a different file does not inherit them.
- Generate a study pack from the Notes workspace: `review.md`, `exam.md`, `exam.json`, `exam.html`, and related files.
- Download `notes.zip`, `notes.md`, `coverage.md`, export files, or the complete output ZIP.
- Switch to **Textbook library**, upload a PDF textbook, and build a RAG-ready corpus. The corpus is not connected to note generation yet.
Expand Down
1 change: 1 addition & 0 deletions gui/README_GUI.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ SlideNote Studio 是一个基于 Streamlit 的图形界面。它包装 `python -
- 在页面里临时填写 Text / Vision / OCR API key;key 只通过本次子进程环境变量传入,不写进命令行。
- 选择是否导出 `notes.zip`、目录 Markdown、Word、PDF 或 LaTeX。
- 进度、ETA、Doctor、用量和成本信息收在紧凑的诊断区里。
- 在页面里保存的模态修正会用于同一源文件的下一次构建;更换文件后不会沿用旧修正。
- 在 Notes workspace 基于已有输出目录生成复习包:`review.md`、`exam.md`、`exam.json`、`exam.html` 等。
- 下载 `notes.zip`、`notes.md`、`coverage.md`、导出文件或完整结果 ZIP。
- 切换到 **Textbook library**,上传 PDF 教材,构建 RAG-ready 教材库;该库当前不会自动参与笔记生成。
Expand Down
131 changes: 116 additions & 15 deletions gui/app.py
Original file line number Diff line number Diff line change
@@ -1,16 +1,20 @@
from __future__ import annotations

import io
import hashlib
import html
import io
import json
import os
import re
import shutil
import subprocess
import threading
import time
import zipfile
from datetime import datetime
from collections import deque
from datetime import datetime, timezone
from pathlib import Path
from queue import Empty, Queue
from typing import Any

import streamlit as st
Expand Down Expand Up @@ -180,6 +184,12 @@ def _run_simplified_app() -> None:
st.error(f"Could not prepare output folder: {exc}")
return
config = _clone_config_for_run(preview_config, input_path=input_path, output_dir=output_dir, progress_json=progress_json)
if _carry_modality_overrides(
Path(st.session_state["last_output_dir"]) if st.session_state.get("last_output_dir") else None,
input_path,
output_dir,
):
st.caption("Saved page modality corrections will be used for this build.")
_run_build(config)
st.session_state["last_output_dir"] = str(output_dir)

Expand Down Expand Up @@ -662,7 +672,8 @@ def _clone_config_for_run(config: StudioConfig, input_path: Path, output_dir: Pa


def _prepare_run_paths(uploaded, output_base: Path, timestamped_subfolder: bool) -> tuple[Path, Path, Path]:
run_name = f"{safe_run_name(uploaded.name)}_{int(time.time())}"
# Keep each uploaded source immutable so corrections can be checked against its bytes.
run_name = f"{safe_run_name(uploaded.name)}_{time.time_ns()}"
input_path = UPLOADS_DIR / f"{run_name}{Path(uploaded.name).suffix.lower()}"
input_path.write_bytes(uploaded.getbuffer())
output_base.mkdir(parents=True, exist_ok=True)
Expand All @@ -672,6 +683,65 @@ def _prepare_run_paths(uploaded, output_base: Path, timestamped_subfolder: bool)
return input_path, output_dir, progress_json


def _output_source_path(output_dir: Path) -> Path | None:
content = _read_json(output_dir / "content.json") or {}
if not isinstance(content, dict):
return None
source_name = content.get("source_path")
if not isinstance(source_name, str) or not source_name:
return None
source_path = Path(source_name)
if not source_path.is_absolute():
source_path = ROOT / source_path
return source_path


def _output_source_matches(output_dir: Path, input_path: Path) -> bool:
manifest = _read_json(output_dir / "page_modalities.overrides.json") or {}
source_hash = manifest.get("source_sha256") if isinstance(manifest, dict) else None
if source_hash is not None:
if not isinstance(source_hash, str) or len(source_hash) != 64 or any(char not in "0123456789abcdefABCDEF" for char in source_hash):
return False
try:
digest = hashlib.sha256()
with input_path.open("rb") as current_file:
while chunk := current_file.read(1024 * 1024):
digest.update(chunk)
return digest.hexdigest() == source_hash.lower()
except OSError:
return False
source_path = _output_source_path(output_dir)
if source_path is None:
return False
try:
if source_path.stat().st_size != input_path.stat().st_size:
return False
with source_path.open("rb") as old_file, input_path.open("rb") as new_file:
while old_chunk := old_file.read(1024 * 1024):
if old_chunk != new_file.read(len(old_chunk)):
return False
return not new_file.read(1)
except OSError:
return False


def _carry_modality_overrides(previous_output_dir: Path | None, input_path: Path, output_dir: Path) -> bool:
manifest_name = "page_modalities.overrides.json"
target = output_dir / manifest_name
if target.is_file():
if _output_source_matches(output_dir, input_path):
return True
# Preserve corrections for the old source without applying them to a different upload.
target.replace(output_dir / f"page_modalities.overrides.stale-{time.time_ns()}.json")
if previous_output_dir is None or previous_output_dir == output_dir:
return False
source = previous_output_dir / manifest_name
if not source.is_file() or not _output_source_matches(previous_output_dir, input_path):
return False
shutil.copy2(source, target)
return True


def _prepare_textbook_paths(uploaded) -> tuple[Path, Path]:
if Path(uploaded.name).suffix.lower() != ".pdf":
raise ValueError("Textbook library v1 only accepts PDF files.")
Expand All @@ -691,7 +761,7 @@ def _run_build(config: StudioConfig) -> None:
status_box = st.empty()
stage_box = st.empty()
log_box = st.empty()
logs: list[str] = []
logs: deque[str] = deque(maxlen=120)

process = subprocess.Popen(
cmd,
Expand All @@ -704,21 +774,38 @@ def _run_build(config: StudioConfig) -> None:
errors="replace",
bufsize=1,
)
while process.poll() is None:
output_queue: Queue[str] = Queue()

def read_output() -> None:
if process.stdout is not None:
line = process.stdout.readline()
if line:
logs.append(line.rstrip())
with process.stdout:
for line in process.stdout:
output_queue.put(line.rstrip("\r\n"))

reader = threading.Thread(target=read_output, daemon=True)
reader.start()

while True:
for _ in range(200):
try:
logs.append(output_queue.get_nowait())
except Empty:
break
_update_progress_ui(config.progress_json, progress_bar, status_box, stage_box)
log_box.code("\n".join(logs[-80:]) or "Running...", language="text")
log_box.code("\n".join(list(logs)[-80:]) or "Running...", language="text")
if process.poll() is not None:
break
time.sleep(0.25)

if process.stdout is not None:
rest = process.stdout.read()
if rest:
logs.extend(rest.splitlines())
reader.join(timeout=2.0)
while True:
try:
logs.append(output_queue.get_nowait())
except Empty:
break
process.wait()
_update_progress_ui(config.progress_json, progress_bar, status_box, stage_box)
log_box.code("\n".join(logs[-120:]) or "No console output.", language="text")
log_box.code("\n".join(logs) or "No console output.", language="text")

if process.returncode == 0:
_generate_cost_report(config.output_dir)
Expand Down Expand Up @@ -1121,7 +1208,21 @@ def _save_modality_override(output_dir: Path, slide_id: int, modality: str, note
path = output_dir / "page_modalities.overrides.json"
data = _read_json(path) or {"schema_version": 1, "pages": {}}
pages = data.setdefault("pages", {})
pages[str(slide_id)] = {"modality": modality, "note": note, "updated_at": datetime.utcnow().isoformat(timespec="seconds") + "Z"}
pages[str(slide_id)] = {
"modality": modality,
"note": note,
"updated_at": datetime.now(timezone.utc).isoformat(timespec="seconds").replace("+00:00", "Z"),
}
source_path = _output_source_path(output_dir)
if source_path is not None:
try:
digest = hashlib.sha256()
with source_path.open("rb") as source_file:
while chunk := source_file.read(1024 * 1024):
digest.update(chunk)
data["source_sha256"] = digest.hexdigest()
except OSError:
pass
path.write_text(json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8")


Expand Down
3 changes: 2 additions & 1 deletion gui/studio_core.py
Original file line number Diff line number Diff line change
Expand Up @@ -224,7 +224,8 @@ def progress_percent(progress: dict[str, Any]) -> float:
current = progress.get("current_stage") or {}
stages = progress.get("stages") or []
completed = len(stages)
total_known_stages = 13
planned = progress.get("planned_stages")
total_known_stages = len(planned) if isinstance(planned, list) and planned else 13
base = min(completed / total_known_stages, 0.95)
stage_total = current.get("total") or 0
stage_current = current.get("current") or 0
Expand Down
29 changes: 6 additions & 23 deletions slidenote/build/artifacts.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,43 +3,26 @@
from pathlib import Path
from typing import Any

from slidenote.pipeline import ArtifactRegistry, BuildContext, FunctionStage, StageResult, run_stage
from slidenote.pipeline import ArtifactRegistry


def _run_json_stage(
deck,
context: BuildContext,
state,
*,
name: str,
artifact_name: str,
artifact_path: str,
message: str,
complete_message: str,
runner,
dependencies: list[str] | None = None,
) -> dict[str, Any]:
progress = context.progress
progress = state.progress
progress.start_stage(name, message=message)

def stage_runner(stage_deck, stage_context: BuildContext) -> StageResult:
report = runner(stage_deck)
artifacts: dict[str, str] = {}
if stage_context.artifacts is not None:
stage_context.artifacts.write_json(artifact_name, artifact_path, report)
registered = stage_context.artifacts.relative_path(artifact_name)
if registered:
artifacts[artifact_name] = registered
return StageResult(name=name, report=report, artifacts=artifacts)

stage = FunctionStage(
name=name,
dependencies=dependencies or [],
artifacts=[artifact_name],
runner=stage_runner,
)
result = run_stage(deck, context, stage)
report = runner(deck)
state.artifacts.write_json(artifact_name, artifact_path, report)
progress.finish_stage(complete_message)
return result.report or {}
return report or {}


def _register_export_artifacts(artifacts: ArtifactRegistry, export_report: dict[str, Any]) -> None:
Expand Down
3 changes: 3 additions & 0 deletions slidenote/build/progress.py
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,9 @@ def callback(event: dict[str, Any]) -> None:

def _llm_progress(progress: ProgressReporter):
def callback(record: dict[str, Any]) -> None:
if record.get("event") == "total":
progress.set_total(record.get("total"))
return
label = record.get("context_id") or record.get("slide_id")
progress.advance(
message=f"LLM context {label}",
Expand Down
6 changes: 3 additions & 3 deletions slidenote/build/runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,9 +8,10 @@
_friendly_build_error,
)
from slidenote.build.errors import UserFacingConfigError
from slidenote.build.stages import BUILD_STAGES, _print_build_outputs
from slidenote.build.stages import BUILD_PHASES, _print_build_outputs
from slidenote.build.state import create_build_state
from slidenote.exporting import parse_export_formats
from slidenote.pipeline import run_build_plan


def run_build(args: argparse.Namespace) -> int:
Expand All @@ -23,8 +24,7 @@ def run_build(args: argparse.Namespace) -> int:

state = create_build_state(args, export_formats)
try:
for stage in BUILD_STAGES:
stage(state)
run_build_plan(state, BUILD_PHASES)
except Exception as exc:
friendly_message = _friendly_build_error(exc, args)
if friendly_message:
Expand Down
Loading
Loading