Skip to content

Latest commit

 

History

History
1103 lines (879 loc) · 42.5 KB

File metadata and controls

1103 lines (879 loc) · 42.5 KB

情绪告别仪式 - 技术方案

版本:v0.5 | 作者:[待填写] | 日期:2026-01-23

1. 概述

1.1 背景

基于 PRD 需求,构建一个帮助用户通过仪式感交互处理负面情绪的产品。产品核心是将心理咨询疗法(CBT)具象化为火葬仪式,配合 AI 情绪分析与内容生成,提供轻量级、即时可用的情绪疗愈体验。

1.2 MVP 范围

范围 MVP 后续迭代
平台 手机端 WebApp(响应式) -
输入方式 语音,文字输入(最多 500 字) -
仪式类型 后端返回3种仪式视角,前端实现火葬动画 土葬、水葬动画
用户体系 无需登录,device_id + SQLite -
浏览器 现代浏览器(Chrome/Safari/Edge 最新 2 版本) -

1.3 技术目标

  • 沉浸式体验:流畅的动画与交互,营造仪式感
  • 智能化:AI 驱动的情绪识别与个性化内容生成
  • 低延迟:情绪分析与内容生成响应时间 < 3s
  • 数据持久化:SQLite + Render Volume 存储用户数据和生成图片

2. 系统架构

sequenceDiagram
    participant U as 用户
    participant F as 前端
    participant B as 后端
    participant G as Gemini

    U->>F: 1. 输入烦恼
    F->>B: POST /api/analyze
    Note over F: [analyzing] Loading...
    B->>G: 内容路由 Agent (守卫 + 风险分类)
    alt 无效内容
        G-->>B: blocked=True
        B-->>F: 400 CONTENT_BLOCKED
    else 高风险 (自伤/自杀)
        G-->>B: is_high_risk=True
        Note over B: 跳过情绪分析,emotion="Crisis"
        B->>G: 图片生成 Agent
        G-->>B: image_path
        B-->>F: {record_id, "Crisis", "fire", image_path}
    else 正常烦恼
        G-->>B: is_high_risk=False
        B->>G: 情绪分析 Agent + 图片生成 Agent
        G-->>B: emotion_label, recommended_ritual, image_path
        B-->>F: {record_id, emotion_label, recommended_ritual, image_path}
    end

    U->>F: 2. 选择仪式
    F->>B: POST /api/ritual
    Note over F: [processing] Loading...
    alt 高风险记录
        B->>G: 危机支持 Agent
        G-->>B: 3 安全建议 + 紧急总结
    else 正常记录
        B->>G: 对应仪式 Agent (火/土/水)
        G-->>B: perspectives[3], summary
    end
    B-->>F: {perspectives, summary}

    U->>F: 3. 选择视角
    Note over F: [animation] 播放图腾动画
    F->>B: POST /api/complete
    B-->>F: {success}
    Note over F: [complete] 展示图腾+总结
Loading

3. 技术选型

3.1 前端

组件 选型 说明
运行时 Bun 快速的 JS 运行时和包管理器
框架 React 组件化开发
构建工具 Vite 极快的 HMR,Bun 原生支持
动画 Framer Motion 页面过渡 + 仪式动画
状态管理 Zustand 轻量、简洁
样式 Tailwind CSS 快速开发
国际化 i18next + react-i18next 多语言支持 (English, 后续中文)
设备标识 uuid UUID v7 生成 (时间排序)
时间格式 date-fns 相对时间显示 ("2 hours ago")
API 转换 camelcase-keys + snakecase-keys snake_case ↔ camelCase 自动转换

3.2 后端

组件 选型 说明
语言 Python 3.12+ AI 生态友好
框架 FastAPI 异步、自动 OpenAPI 文档
AI 框架 Google ADK Agent Development Kit,构建 AI Agent
AI 可观测性 Logfire Pydantic 出品,自动追踪 Google GenAI 调用
ORM SQLAlchemy (sync) 同步模式,SQLite 无需异步
DB 迁移 Alembic 数据库 schema 迁移,FastAPI 启动时自动执行
包管理 uv 快速的 Python 包管理器
容器化 Docker Compose 本地开发环境
部署 Render Blueprint (render.yaml) 定义前后端服务,GitHub 推送自动部署
数据库 SQLite 轻量级,文件存储在 Render Volume

3.3 AI / LLM

能力 选型 说明
主模型 gemini-3-flash-preview 快速且经济,用于情绪分析和仪式 Agent
守卫模型 gemini-2.5-flash-lite 超快,用于 Guardrail 内容过滤
路由模型 gemini-2.5-flash-lite 内容路由 Agent,风险分类
危机模型 gemini-3-flash-preview 危机支持 Agent,处理高风险内容
图片模型 gemini-2.5-flash-image 原生图片生成,1024x1024 PNG
图片提示词模型 gemini-3-flash-preview 图片 Prompt Agent,带 thinking 生成优质提示词
搜索模型 gemini-2.5-flash-lite Web 搜索 Agent,处理时效性烦恼

3.4 环境变量

后端环境变量

变量名 说明 示例
GOOGLE_API_KEY Google API 密钥 (ADK 需要) AIza...
GEMINI_MODEL Gemini 主模型 (情绪分析/仪式) gemini-3-flash-preview
GEMINI_GUARDRAIL_MODEL Gemini 守卫模型 (内容过滤) gemini-2.5-flash-lite
GEMINI_ROUTING_MODEL 内容路由 Agent 模型 (风险分类) gemini-2.5-flash-lite
GEMINI_CRISIS_MODEL 危机支持 Agent 模型 gemini-3-flash-preview
GEMINI_IMAGE_MODEL Gemini 图片模型 gemini-2.5-flash-image
GEMINI_PROMPT_MODEL 图片 Prompt Agent 模型 (带 thinking) gemini-3-flash-preview
GEMINI_ASR_MODEL Gemini ASR 模型 (语音转文字) gemini-2.5-flash-lite
GEMINI_SEARCH_MODEL Web 搜索 Agent 模型 gemini-2.5-flash-lite
GEMINI_SEARCH_MAX_TOKENS Web 搜索输出上限 (~200 words) 350
WORRY_HISTORY_LIMIT 仪式 Agent 历史上下文数量 5
IMAGE_PROMPT_MAX_ATTEMPTS 图片生成反馈循环最大次数 3
GEMINI_THINKING_LEVEL 思考级别 LOW / MEDIUM / HIGH
GEMINI_MAX_OUTPUT_TOKENS 最大输出 token 4096
LOGFIRE_TOKEN Logfire API token (可选) ...
RATE_LIMIT_PER_MINUTE 每分钟请求上限 100
CORS_ORIGINS_STR 允许的前端域名 https://xxx.onrender.com
DATA_DIR 数据存储目录 (Render Volume) /var/data
评估测试 (Giskard)
GISKARD_EVAL_MODEL Giskard 扫描评估 LLM (litellm provider/model 格式) gemini/gemini-3-flash-preview
GISKARD_EMBEDDING_MODEL Giskard 嵌入模型 (litellm provider/model 格式) gemini/gemini-embedding-001
评估测试 (pydantic-evals)
PYDANTIC_EVAL_MODEL pydantic-evals LLMJudge 模型 (pydantic-ai provider:model 格式) google-gla:gemini-3-flash-preview

前端环境变量

变量名 说明 默认值
VITE_API_URL 后端 API 地址 http://localhost:8000
VITE_API_TIMEOUT API 请求超时时间 (毫秒) 30000

3.5 数据存储

层级 方案 说明
前端 localStorage 仅存 device_id (UUID)
后端 SQLite + 文件系统 用户数据、仪式记录、生成图片
/var/data/                 # Render Volume 挂载点
├── db.sqlite              # SQLite 数据库 (Alembic 管理 schema 迁移)
└── {device_id}/           # 生成的情绪图片 (per device)
    └── {record_id}.png    # 图片以 record_id 命名

数据库迁移 (Alembic):

  • FastAPI 启动时自动运行 alembic upgrade head
  • 自动检测 pre-Alembic 数据库并 stamp baseline
  • Render 不能使用 preDeployCommand(因为无法访问 Persistent Disk),所以在 lifespan 中执行
  • 迁移脚本使用 batch_alter_table 支持 SQLite ALTER TABLE

3.6 开发规范

类别 工具 说明
前端 Lint Biome Rust 编写,Lint + Format 一体化
后端 Lint Ruff Python linter + formatter
前端类型检查 tsc TypeScript compiler (--noEmit)
后端类型检查 ty Python type checker
后端安全扫描 Bandit Python security linter
环境变量检查 dotenv-linter .env 文件检查
Markdown mdformat Markdown 格式化
任务运行 just 命令运行器,统一 lint/test/build 等任务
前端测试 Bun Test 原生测试运行器 + happy-dom
后端测试 pytest 核心逻辑单元测试
AI 安全测试 Giskard LLM 安全扫描 (harmfulness, sycophancy, prompt injection, hallucination for web search)
AI 质量评估 pydantic-evals LLM-as-judge 评估 (仪式质量、情绪分类、危机安全) + Logfire 追踪
AI 评估 LLM Gemini 3 Flash Preview Giskard 扫描评估器 (via LiteLLM gemini/ prefix) + pydantic-evals LLMJudge (via google-gla:)
测试策略 Selective TDD 关键逻辑 TDD,UI 组件手动测试
本地联调 Docker Compose 一键启动前后端
Mock MSW (Mock Service Worker) 开发时拦截请求,返回模拟数据
日志监控 Render 日志 MVP 阶段使用平台自带日志

3.7 LLM 评估测试 (Giskard + pydantic-evals)

使用 Giskard + pydantic-evals 测试 LLM 输出质量。与单元测试分离,不在 CI 中运行。Logfire 追踪 pydantic-evals + Google GenAI 调用。

三种测试类型:

  1. 直接断言测试 — 结构化输出 Agent (返回 bool/enum/JSON),使用真实 API 调用 + pytest 断言
  2. Giskard 扫描测试 — 自由文本输出 Agent (返回 perspectives/summary),使用 Giskard 检测器评估
  3. pydantic-evals 评估 — Dataset + Evaluator 模式,支持 LLMJudge 和确定性评估器,Logfire 追踪所有 spans
Agent 输出类型 测试方法
Guardrail (_call_judge) bool 直接断言
路由 Agent RoutingOutput (JSON) 直接断言 + prompt_injection 扫描
情绪 Agent EmotionOutput (enum) 直接断言 + prompt_injection 扫描
仪式 Agents (火/土/水) 自由文本 Giskard 扫描 (harmfulness, sycophancy)
危机 Agent 自由文本 直接断言 (紧急号码) + Giskard 扫描 (harmfulness)
Web 搜索 Agent 自由文本 直接断言 + Giskard 扫描 (hallucination)

Giskard 检测器分配:

检测器 适用 Agent 说明
harmfulness 仪式文本、危机文本 检测有害内容
hallucination Web 搜索 检测虚构事实 (仅适用于有外部数据源的 Agent)
sycophancy 仪式文本 检测过度迎合
prompt_injection 路由、情绪 (所有 Agent) 检测行为变化

运行命令:

just eval-giskard-safety  # Giskard 安全关键测试 (~3-5 min, ~$0.10)
just eval-giskard         # 全部 Giskard 测试
just eval-pydantic        # pydantic-evals 质量评估 (LLMJudge + 确定性)
just eval                 # 全部评估测试 (Giskard + pydantic-evals)
just eval-giskard-scan    # 完整 Giskard 扫描 (~30+ min, ~$3-5),生成 HTML 报告

评估 LLM 配置:

  • Giskard 使用 GISKARD_EVAL_MODEL (默认 gemini/gemini-3-flash-preview) 作为评估器,通过 LiteLLM 调用
  • pydantic-evals 使用 PYDANTIC_EVAL_MODEL (默认 google-gla:gemini-3-flash-preview) 作为 LLMJudge 模型
  • Logfire 追踪: conftest 统一配置 logfire.configure(send_to_logfire="if-token-present"),同时 instrument pydantic-ai 和 google-genai

3.8 错误处理与容错

场景 策略
Gemini API 失败 自动重试 2 次 → 失败后使用预设 fallback 内容
请求超时 前端设置 15s 超时,超时后提示用户重试
Rate Limiting 后端全局限制 100 次/分钟
磁盘空间不足 定期清理旧图片,保留最近 100 条记录

错误响应格式

{
  "error": "错误信息描述",
  "code": "ERROR_CODE"
}

常见错误码:

  • CONTENT_BLOCKED - 非烦恼内容被守卫拦截
  • CONTENT_TOO_LONG - 输入超过 500 字
  • RATE_LIMITED - 请求过于频繁
  • AI_SERVICE_ERROR - Gemini API 异常
  • INTERNAL_ERROR - 服务器内部错误

4. 技术决策记录

# 事项 决策 备注
1 燃烧动画实现 WebM 视频 使用带透明通道的 WebM 视频
2 UI 语言 English 支持 i18n,后续添加中文
3 博物馆详情 翻转卡片 点击卡片翻转显示详情,3D carousel 布局
4 语音输入 MVP 包含 后端 Gemini ASR (多语言 + 断句)
5 设备标识 UUID v7 时间排序,便于数据库索引
6 测试策略 Selective TDD Store/Utils/API 先写测试,UI 手动测试
7 导航方式 无 TabBar InputStep 有 Museum 按钮,CompleteStep 有两个按钮
8 DB 迁移 Alembic (lifespan) FastAPI 启动时自动执行,Render 无法用 preDeployCommand (Persistent Disk 不可访问)
9 危机内容路由 Query Routing Agent 守卫 + 风险分类前置于情绪分析,高风险内容跳过情绪检测直接进入危机支持

4.1 语音输入实现

使用后端 Gemini ASR 实现语音转文字(支持多语言):

// useAudioRecorder hook (src/hooks/useAudioRecorder.ts)
interface UseAudioRecorderReturn {
  isSupported: boolean;   // MediaRecorder API 是否可用
  isRecording: boolean;   // 是否正在录音
  error: string | null;   // 错误信息 ("denied" 等)
  start: () => Promise<void>;  // 请求麦克风权限并开始录音
  stop: () => Promise<{ blob: Blob; mimeType: string } | null>;  // 停止录音并返回音频
}

// useSpeechRecognition hook (src/hooks/useSpeechRecognition.ts)
interface UseSpeechRecognitionReturn {
  isSupported: boolean;   // MediaRecorder API 是否可用
  isListening: boolean;   // 是否正在录音
  isProcessing: boolean;  // 是否正在转录
  transcript: string;     // 转录文本
  error: string | null;   // 错误信息 ("denied" 等)
  start: () => void;      // 开始录音
  stop: () => void;       // 停止录音并调用后端 ASR
  reset: () => void;      // 重置状态
}

// 实现细节
- 使用 MediaRecorder API 录制音频 (支持 webm/mp4/ogg)
- 停止时调用 POST /api/transcribe 进行转录
- 支持多语言 (英文、中文、code-switching)
- 自动添加标点符号 (断句)
- 处理权限拒绝 (error: "denied")
- 不支持时 isSupported: false,显示禁用按钮 + tooltip

// VoiceInputButton 组件 (src/components/common/VoiceInputButton.tsx)
- 四种状态:idle (灰色麦克风)、listening (红色脉冲)、processing (三点动画)、unsupported (禁用+tooltip)
- 支持 reduced motion (跳过脉冲动画)
- Props: onTranscript(text) → 追加到输入框

4.2 响应式设计规范

断点 宽度 目标设备
默认 < 640px 手机 (主要目标)
sm 640px+ 大屏手机
md 768px+ 平板
lg 1024px+ 桌面 (次要)

移动端要求:

  • 安全区域: env(safe-area-inset-*) 处理刘海/Home 指示器
  • 触摸目标: 最小 44x44px
  • 输入框字号: 最小 16px (防止 iOS 缩放)
  • 视口高度: 使用 100dvh (动态视口高度)
  • 键盘处理: 输入框 onFocus 时 300ms 延迟后 scrollIntoView({ behavior: 'smooth', block: 'center' })

4.3 微交互动画

按钮反馈 (AnimatedButton):

  • whileTap={{ scale: 0.95 }} - 按下缩小
  • whileHover={{ scale: 1.02 }} - 悬停放大
  • 支持 reduced motion (禁用动画)

卡片选择 (SelectRitualStep, SelectPerspectiveStep):

  • whileHover={{ scale: 1.02, boxShadow: '0 0 20px rgba(124, 58, 237, 0.3)' }} - 悬停发光
  • whileTap={{ scale: 0.98 }} - 点击缩小

文字渐现 (CompleteStep):

  • staggerChildren: 0.15 - 子元素依次出现
  • itemVariants: { hidden: { opacity: 0, y: 10 }, visible: { opacity: 1, y: 0 } }
  • 支持 reduced motion (跳过 stagger 和位移)

4.4 火葬动画序列

阶段 1 (0.5s): 用户文字淡出、缩小
阶段 2 (2-3s): 火焰动画播放
阶段 3 (1-2s): 火焰淡出,晶体出现
阶段 4 (0.5s): 晶体稳定,准备进入完成页

注意: 阶段 3 期间后台调用 /api/complete,不阻塞动画


5. 页面结构

5.1 视图架构(零路由)

整个 App 为单页应用,通过状态切换视图,不使用路由库。

导航说明(无 TabBar):

  • InputStep:显示 "Museum" 按钮,可进入博物馆
  • CompleteStep:显示 "View Museum" 和 "Back Home" 两个按钮
  • Museum:显示 "Homepage" 链接,返回 InputStep

5.2 状态定义

interface AppState {
  view: 'ritual' | 'museum';
  step:
    | 'input'                   // 1. 用户输入烦恼
    | 'analyzingEmotion'        // 2. Loading: 分析情绪 + 生成图片
    | 'selectRitual'            // 3. 选择仪式
    | 'generatingPerspectives'  // 4. Loading: 生成视角 + 总结
    | 'selectPerspective'       // 5. 选择视角
    | 'totemAnimation'         // 6. 生成图腾动画
    | 'complete';               // 7. 展示图腾 + 总结
}

5.3 仪式流程

input → analyzingEmotion → selectRitual → generatingPerspectives → selectPerspective → totemAnimation → complete
  ↓            ↓                ↓                    ↓                     ↓                  ↓             ↓
[输入]     [分析情绪]        [选仪式]           [生成视角]             [选视角]          [生成图腾]    [图腾+总结]
                                                                                                           ↓
                                                                                                 [再来一次] [去博物馆]

5.4 心灵博物馆

布局: 3D 卡片轮播(Embla Carousel)

  • 中心卡片:原尺寸,正面朝向
  • 侧边卡片:缩小 (0.85),倾斜 (±15°),半透明

卡片设计: 翻转卡片

  • 正面: 图腾图标 + 日期、情绪标签、情绪图片、选择的视角文字
  • 背面: 原始烦恼、AI 总结、仪式类型
  • 交互: 点击翻转,滑动切换时自动重置

导航: "Homepage" 链接返回 InputStep

图腾类型: Crystal(火)、Tree(土)、Ripple(水)


6. 核心模块设计

6.0 内容路由 Agent (query_routing_agent)

/analyze 的第一步:Guardrail 守卫 + 风险等级分类。决定内容是否有效、是否高风险。

架构:

  • agent.py - Google ADK Agent + structured output + guardrail 回调
  • models.py - RoutingOutput (is_valid_worry, is_high_risk, blocked)
  • prompts.py - 风险分类指令 + few-shot 示例
  • Guardrail (before_model_callback) 从情绪分析 Agent 移至此处
routing_agent = Agent(
    name="query_router",
    model="gemini-2.5-flash-lite",          # 快速分类
    instruction=ROUTING_INSTRUCTION,
    output_schema=RoutingOutput,
    before_model_callback=worry_content_guardrail,  # 守卫回调 (复用)
)

分类逻辑:

  • Guardrail 先触发 → 无效内容 → blocked=True → 400 CONTENT_BLOCKED
  • 如果 Guardrail 通过 → Agent 判断风险等级 → is_high_risk: bool
  • 高风险指标:自伤、自杀意念、感到绝望想结束一切
  • 路由 Agent 失败时 → fail-open(is_high_risk=False,继续正常流程)

6.1 情绪分析 Agent (ritual_recommend_agent)

分析用户烦恼,推荐合适的仪式。仅处理非高风险内容(高风险由路由 Agent 拦截)。

架构:

  • agent.py - Google ADK LlmAgent + structured output
  • guardrail.py - Gemini-as-judge 守卫 (已移至路由 Agent 的 before_model_callback)
  • models.py - EmotionInput/EmotionOutput
  • prompts.py - XML 标签格式的提示词
emotion_agent = LlmAgent(
    name="emotion_detector",
    model="gemini-3-flash-preview",
    instruction=EMOTION_INSTRUCTION,
    output_schema=EmotionOutput,
    # 注意:guardrail 已移至 query_routing_agent
)

情绪标签 → 仪式映射:

  • 愤怒类 (anger, frustration) → fire (CBT)
  • 焦虑类 (anxiety, stress, worry) → fire (CBT)
  • 抑郁类 (depression, exhaustion) → earth (ACT)
  • 悲伤类 (sadness, grief, loss) → water (AEDP)
  • 危机 (Crisis) → fire (由路由 Agent 直接设置,跳过情绪分析)

6.2 Web 搜索 Agent (web_search_agent)

处理时效性烦恼,获取实时信息辅助仪式 Agent 生成更相关的视角。

架构:

  • agent.py - Google ADK Agent + 内置 google_search 工具
  • models.py - WebSearchInput/WebSearchOutput
  • prompts.py - 搜索指令 + 用户消息模板
# 使用 Google ADK 内置搜索工具
from google.adk.tools import google_search

web_search_agent = Agent(
    name="web_search_agent",
    model="gemini-2.5-flash-lite",
    instruction=WEB_SEARCH_INSTRUCTION,
    tools=[google_search],  # 内置 Google Search
    generate_content_config=types.GenerateContentConfig(
        max_output_tokens=350,  # 限制输出 ~200 words
    ),
)

注意事项:

  • google_search 内置工具不能与 output_schema 同时使用
  • 作为 AgentTool 提供给仪式 Agent 使用
  • 仪式 Agent 会根据烦恼内容判断是否需要搜索实时信息

6.4 图片生成服务 (image_gen)

根据用户烦恼生成治愈图片。使用 Gemini 原生图片生成 (非 ADK)。

架构:

  • generate.py - 直接调用 google-genai 客户端
  • models.py - ImageInput/ImageOutput (base64 编码)
  • prompts.py - 治愈图片提示词 + XML 标签
# 使用 google-genai 直接调用 (非 ADK)
from google import genai
from google.genai import types

async def generate_ritual_image(input_data: ImageInput) -> ImageOutput:
    client = genai.Client(api_key=settings.google_api_key)
    response = client.models.generate_content(
        model="gemini-2.5-flash-image",  # 原生图片生成模型
        contents=prompt,
        config=types.GenerateContentConfig(
            system_instruction=IMAGE_INSTRUCTION,
            response_modalities=["IMAGE"],  # 仅返回图片
        ),
    )
    # 返回 base64 编码的 PNG (1024x1024)
    return ImageOutput(image_base64=..., mime_type="image/png")

图片风格:

  • 宁静、治愈的氛围
  • 柔和的色调和光线
  • 抽象或象征性的情绪释放表达
  • 聚焦自然、风景或抽象元素

6.5 图片提示词 Agent (image_prompt_agent)

使用 thinking 能力生成高质量的图片提示词,支持反馈循环优化。

架构:

  • agent.py - Google ADK Agent + thinking + 反馈循环
  • models.py - ImagePromptInput/ImagePromptOutput
  • prompts.py - 提示词生成指令
from google.adk import Agent

image_prompt_agent = Agent(
    name="image_prompt_agent",
    model="gemini-3-flash-preview",  # 支持 thinking
    instruction=IMAGE_PROMPT_INSTRUCTION,
    output_schema=ImagePromptOutput,
    generate_content_config=types.GenerateContentConfig(
        thinking_config=types.ThinkingConfig(
            thinking_budget_tokens=1024,
        ),
    ),
)

反馈循环:

  • 最多 3 次尝试 (IMAGE_PROMPT_MAX_ATTEMPTS)
  • 图片生成失败时,将错误信息反馈给 Agent 重新生成提示词
  • 支持多种生成策略 (silhouette, color_emotion, nature_metaphor 等)

6.6 火仪式 Agent (CBT)

认知行为疗法:挑战非理性信念,认知重建

增强功能:

  • 接收用户过往烦恼历史 (最多 5 条) 作为上下文
  • 注入当前时间,支持处理时效性烦恼
  • 可调用 Web 搜索 Agent 获取实时信息
from google.adk.tools import AgentTool
from app.services.web_search_agent.agent import web_search_agent

fire_agent = Agent(
    name="fire_ritual",
    model="gemini-3-flash-preview",
    instruction=FIRE_INSTRUCTION,  # CBT 系统指令
    output_schema=RitualOutput,
    tools=[AgentTool(agent=web_search_agent)],  # Web 搜索作为工具
)

6.7 土仪式 Agent (ACT)

接纳承诺疗法:接纳当下,承诺行动

增强功能: 同火仪式 Agent (历史上下文、时间注入、Web 搜索工具)

earth_agent = Agent(
    name="earth_ritual",
    model="gemini-3-flash-preview",
    instruction=EARTH_INSTRUCTION,  # ACT 系统指令
    output_schema=RitualOutput,
    tools=[AgentTool(agent=web_search_agent)],
)

6.8 水仪式 Agent (AEDP)

加速体验性动力学治疗:允许情绪流动,温柔抚慰

增强功能: 同火仪式 Agent (历史上下文、时间注入、Web 搜索工具)

water_agent = Agent(
    name="water_ritual",
    model="gemini-3-flash-preview",
    instruction=WATER_INSTRUCTION,  # AEDP 系统指令
    output_schema=RitualOutput,
    tools=[AgentTool(agent=web_search_agent)],
)

6.9 危机支持 Agent (Crisis Support)

处理高风险内容(自伤/自杀意念)。提供即时安全建议和危机资源。

增强功能: 同火仪式 Agent (历史上下文、时间注入、Web 搜索工具)

crisis_agent = Agent(
    name="crisis_support",
    model="gemini-3-flash-preview",
    instruction=CRISIS_INSTRUCTION,
    output_schema=RitualOutput,       # 复用: 3 perspectives + summary
    tools=[AgentTool(agent=web_search_agent)],
)

3 个必需建议类型:

  1. Physical Grounding — 身体感官练习(冰水、深呼吸、5-4-3-2-1 感知)
  2. Professional Mobilization — 专业资源(988 Suicide & Crisis Lifeline、741741 Crisis Text Line、911)
  3. Immediate Safety Commitment — 短期安全承诺(移除危险物品、去共享空间、承诺安全 1 小时)

Summary 为直接紧急的安全呼吁,包含具体求助号码。

路由: 当 record.is_high_risk=True 时,/ritual 路由自动使用此 Agent,忽略用户选择的仪式类型。


7. 数据模型

7.1 SQLite 表结构

-- 仪式记录表 (Alembic 管理迁移: 001 + 002)
CREATE TABLE records (
    id              TEXT PRIMARY KEY,           -- UUID v7
    device_id       TEXT NOT NULL,              -- 用户设备标识 (UUID v7)
    content         TEXT NOT NULL,              -- 用户输入的烦恼
    emotion_label   TEXT NOT NULL,              -- 情绪标签 (Anger/Anxiety/Grief/Crisis)
    recommended_ritual TEXT NOT NULL,           -- AI 推荐的仪式 (fire/earth/water)
    selected_ritual TEXT,                       -- 用户选择的仪式
    image_path      TEXT,                       -- 生成的情绪图片路径
    perspectives    TEXT,                       -- JSON: 3个视角
    selected_perspective TEXT,                  -- 用户选择的视角
    summary         TEXT,                       -- AI 生成的总结
    is_high_risk    BOOLEAN DEFAULT 0 NOT NULL, -- 高风险标记 (自伤/自杀)
    status          TEXT DEFAULT 'pending',     -- pending/completed
    created_at      DATETIME DEFAULT CURRENT_TIMESTAMP,
    completed_at    DATETIME
);

CREATE INDEX idx_device_id ON records(device_id);
CREATE INDEX idx_created_at ON records(created_at);

7.2 TypeScript 类型定义

// 前端 localStorage 仅存 device_id (UUID v7)
interface LocalStorage {
  deviceId: string;
}

// 应用状态(Zustand store)
interface AppState {
  view: 'ritual' | 'museum';
  step:
    | 'input'
    | 'analyzingEmotion'
    | 'selectRitual'
    | 'generatingPerspectives'
    | 'selectPerspective'
    | 'totemAnimation'
    | 'complete';
}

// API 返回的记录
interface Record {
  id: string;
  content: string;
  emotionLabel: string;  // e.g., "Anger", "Anxiety", "Grief"
  recommendedRitual: 'fire' | 'earth' | 'water';
  selectedRitual: 'fire' | 'earth' | 'water';
  imagePath: string;
  perspectives: Perspective[];
  selectedPerspective: string;  // perspective id (e.g., "p1")
  summary: string;
  status: 'pending' | 'completed';
  createdAt: string;
  completedAt?: string;
}

interface Perspective {
  id: string;
  text: string;
}

8. API 设计

8.1 健康检查

GET /health

Response:
{
  "status": "ok"
}

8.2 通用约束

  • 所有请求需携带 X-Device-Id Header(UUID v7)
  • 所有请求的 content 字段最大 500 字,超出返回 400 错误
  • 后端设置全局 Rate Limiting(100 次/分钟),使用 slowapi 实现
  • CORS 配置允许前端域名访问
  • API 响应使用 snake_case(如 emotion_type),前端转换为 camelCase(如 emotionType)

8.3 情绪分析 + 仪式推荐

分析情绪、推荐仪式、生成情绪图片(并行)

POST /api/analyze
Header: X-Device-Id: <uuid>

Request:
{
  "content": "最近工作压力好大,老板总是在 deadline 前临时加需求..."
}

Response:
{
  "record_id": "550e8400-e29b-41d4-a716-446655440000",
  "emotion_label": "Anxiety",
  "recommended_ritual": "fire",
  "image_path": "/api/images/{device_id}/550e8400-e29b-41d4-a716-446655440000.png"
}

8.4 生成视角

根据选择的仪式,调用对应 Agent 生成 3 个视角 + 总结

POST /api/ritual
Header: X-Device-Id: <uuid>

Request:
{
  "record_id": "550e8400-e29b-41d4-a716-446655440000",
  "ritual_type": "fire"
}

Response:
{
  "perspectives": [
    {"id": "p1", "text": "这个困难是暂时的,不是永久的困境"},
    {"id": "p2", "text": "你无法控制他人的行为,但可以选择自己的回应方式"},
    {"id": "p3", "text": "即使结果不理想,你也从中学到了重要的经验"}
  ],
  "summary": "你今天释放了对工作压力的焦虑。通过火焰仪式,你认识到这个困难是暂时的。"
}

8.5 完成仪式

用户选择视角后,保存完整记录

POST /api/complete
Header: X-Device-Id: <uuid>

Request:
{
  "record_id": "550e8400-e29b-41d4-a716-446655440000",
  "selected_perspective": "p1"
}

Response:
{
  "success": true
}

8.6 获取历史记录

博物馆页面获取用户历史仪式记录

GET /api/records
Header: X-Device-Id: <uuid>

Response:
{
  "records": [
    {
      "id": "550e8400-e29b-41d4-a716-446655440000",
      "content": "最近工作压力好大...",
      "selected_ritual": "fire",
      "image_path": "/api/images/{device_id}/550e8400-e29b-41d4-a716-446655440000.png",
      "perspectives": [
        {"id": "p1", "text": "这个困难是暂时的..."},
        {"id": "p2", "text": "..."},
        {"id": "p3", "text": "..."}
      ],
      "selected_perspective": "p1",
      "summary": "你今天释放了对工作压力的焦虑...",
      "status": "completed",
      "created_at": "2026-01-26T10:30:00Z",
      "completed_at": "2026-01-26T10:35:00Z"
    }
  ]
}

8.7 获取图片

GET /api/images/{device_id}/{filename}

Response: image/png

8.8 语音转文字 (ASR)

使用 Gemini ASR 将音频转换为文字,支持多语言和自动断句

POST /api/transcribe
Header: X-Device-Id: <uuid>

Request:
{
  "audio_base64": "SGVsbG8gV29ybGQ=...",  // Base64 编码的音频数据
  "mime_type": "audio/webm"                // 音频格式 (audio/webm, audio/mp4, audio/ogg, etc.)
}

Response:
{
  "text": "你好,这是一个测试。",
  "languages": ["zh"]  // ISO 639-1 语言代码,按出现频率排序
}

支持的音频格式: audio/wav, audio/mp3, audio/aiff, audio/aac, audio/ogg, audio/flac, audio/webm

语言检测:

  • 单语言:["en"], ["zh"], ["ja"]
  • Code-switching:["zh", "en"] (按出现频率排序)
  • 无法识别:["und"]

9. 项目结构

gemini-hackathon-2026/
├── frontend/                      # 前端项目
│   ├── index.html                 # Vite 入口
│   ├── src/
│   │   ├── main.tsx               # React 挂载入口
│   │   ├── App.tsx                # 根组件
│   │   ├── global.css             # 全局样式
│   │   ├── api/                   # API 调用封装
│   │   ├── assets/                # 静态资源
│   │   │   ├── icons/             # SVG 图标
│   │   │   └── images/            # 图片
│   │   ├── components/            # React 组件
│   │   │   ├── Layout/            # 布局 + Tab 切换
│   │   │   ├── Ritual/            # 仪式流程
│   │   │   │   ├── InputStep.tsx            # 1. 输入烦恼
│   │   │   │   ├── AnalyzingStep.tsx        # 2. Loading
│   │   │   │   ├── SelectRitualStep.tsx     # 3. 选择仪式
│   │   │   │   ├── ProcessingStep.tsx       # 4. Loading
│   │   │   │   ├── SelectPerspectiveStep.tsx # 5. 选择视角
│   │   │   │   ├── AnimationStep.tsx        # 6. 图腾动画
│   │   │   │   └── CompleteStep.tsx         # 7. 展示结果
│   │   │   ├── Museum/            # 心灵博物馆
│   │   │   └── common/            # 共享组件
│   │   │       ├── AppLoader.tsx        # 应用加载状态
│   │   │       ├── AnimatedButton.tsx   # 按钮 (scale 0.95 按下反馈)
│   │   │       └── VoiceInputButton.tsx # 语音输入按钮
│   │   ├── hooks/                 # 自定义 hooks
│   │   │   ├── index.ts                 # barrel export
│   │   │   ├── useReducedMotion.ts      # 检测 prefers-reduced-motion
│   │   │   ├── useRitualAnimation.ts    # 仪式动画状态机
│   │   │   ├── useAudioRecorder.ts       # 音频录制 (MediaRecorder API)
│   │   │   └── useSpeechRecognition.ts  # 语音识别 (后端 Gemini ASR)
│   │   ├── i18n/                  # 国际化
│   │   │   ├── index.ts           # i18next 配置
│   │   │   └── locales/           # 语言文件 (en.json, zh.json)
│   │   ├── mocks/                 # MSW Mock 数据
│   │   ├── stores/                # Zustand 状态管理
│   │   ├── types/                 # TypeScript 类型
│   │   └── utils/                 # 工具函数
│   ├── tests/                     # Bun 测试
│   ├── package.json               # 依赖配置
│   ├── tsconfig.json              # TypeScript 配置
│   ├── vite.config.ts             # Vite 配置
│   ├── tailwind.config.ts         # Tailwind 配置
│   ├── biome.json                 # Biome Lint 配置
│   ├── Dockerfile                 # 容器配置
│   └── .env.example               # 环境变量模板
│
├── backend/                       # 后端项目
│   ├── app/
│   │   ├── main.py                # FastAPI 入口 (含 Alembic 自动迁移)
│   │   ├── core/
│   │   │   └── config.py          # pydantic-settings 配置
│   │   ├── services/              # AI Agents (Google ADK) + Services
│   │   │   ├── utils/             # 共享工具
│   │   │   │   ├── agent_runner.py # ADK Agent 运行 + Session 清理
│   │   │   │   ├── few_shot.py    # few-shot 事件构建
│   │   │   │   ├── retry.py       # 指数退避重试
│   │   │   │   └── tracing.py     # Logfire 初始化
│   │   │   ├── query_routing_agent/ # 内容路由 Agent (守卫 + 风险分类)
│   │   │   │   ├── agent.py       # ADK Agent + route_content()
│   │   │   │   ├── models.py      # RoutingOutput (blocked, is_high_risk)
│   │   │   │   └── prompts.py     # 风险分类指令 + few-shot
│   │   │   ├── ritual_recommend_agent/  # 情绪分析 + 仪式推荐 Agent
│   │   │   │   ├── agent.py       # ADK Agent + detect_emotion()
│   │   │   │   ├── models.py      # EmotionInput/EmotionOutput
│   │   │   │   ├── prompts.py     # 系统指令 + XML 标签
│   │   │   │   └── guardrail.py   # Gemini-as-judge 守卫 (被路由 Agent 使用)
│   │   │   ├── web_search_agent/  # Web 搜索 Agent
│   │   │   │   ├── agent.py       # ADK Agent + google_search 工具
│   │   │   │   ├── models.py      # WebSearchInput/WebSearchOutput
│   │   │   │   └── prompts.py     # 搜索指令
│   │   │   ├── ritual_agents/     # 仪式 Agents (含 Web 搜索工具)
│   │   │   │   ├── models.py      # RitualInput (含 worry_history) / RitualOutput
│   │   │   │   ├── fire/          # 火仪式 Agent (CBT)
│   │   │   │   │   ├── agent.py
│   │   │   │   │   └── prompts.py
│   │   │   │   ├── earth/         # 土仪式 Agent (ACT)
│   │   │   │   │   ├── agent.py
│   │   │   │   │   └── prompts.py
│   │   │   │   ├── water/         # 水仪式 Agent (AEDP)
│   │   │   │   │   ├── agent.py
│   │   │   │   │   └── prompts.py
│   │   │   │   └── crisis/        # 危机支持 Agent
│   │   │   │       ├── agent.py   # ADK Agent + run_crisis_support()
│   │   │   │       └── prompts.py # 安全建议 + 危机资源指令
│   │   │   ├── image_gen/         # 图片生成服务 (非 ADK)
│   │   │   │   ├── generate.py    # 直接调用 google-genai
│   │   │   │   ├── models.py      # ImageInput/ImageOutput (base64)
│   │   │   │   └── prompts.py     # 治愈图片提示词
│   │   │   ├── image_prompt_agent/ # 图片提示词 Agent (带 thinking)
│   │   │   │   ├── agent.py       # ADK Agent + 反馈循环
│   │   │   │   ├── models.py      # ImagePromptInput/ImagePromptOutput
│   │   │   │   └── prompts.py     # 提示词生成指令
│   │   │   └── asr/               # 语音转文字服务
│   │   │       ├── models.py      # TranscriptionInput/TranscriptionOutput
│   │   │       └── transcribe.py  # Gemini ASR (gemini-2.5-flash-lite)
│   │   ├── routers/               # API 路由
│   │   ├── schemas/               # Pydantic 请求/响应模型
│   │   ├── db/                    # 数据库
│   │   │   ├── database.py        # SQLite 连接
│   │   │   └── models.py          # 数据模型 (含 is_high_risk)
│   │   ├── middleware/            # 中间件
│   │   └── fallbacks/             # AI 失败时的兜底内容 (含危机兜底)
│   ├── alembic/                   # DB 迁移
│   │   ├── env.py                 # Alembic 环境 (render_as_batch for SQLite)
│   │   └── versions/              # 迁移脚本
│   │       ├── 001_create_records_table.py
│   │       └── 002_add_is_high_risk_column.py
│   ├── alembic.ini                # Alembic 配置
│   ├── tests/                     # pytest 测试 (154 unit + 67 eval)
│   │   └── evaluation/            # LLM 评估测试 (需 --run-eval)
│   │       ├── conftest.py        # 共享配置 (Giskard + pydantic-evals)
│   │       ├── datasets.py        # 共享测试数据
│   │       ├── giskard_evals/     # Giskard 扫描 + 直接断言测试
│   │       └── pydantic_evals/    # pydantic-evals LLM-as-judge 测试
│   ├── pyproject.toml             # uv 依赖配置
│   ├── Dockerfile                 # 容器配置
│   └── .env.example               # 环境变量示例
│
├── docs/                          # 文档
│   ├── prd.md                     # 产品需求文档
│   └── tech-spec.md               # 技术方案 (本文档)
├── docker-compose.yml             # 本地开发环境
├── render.yaml                    # Render 部署配置
├── justfile                       # 任务运行器 (lint/test/build)
├── .gitignore                     # Git 忽略配置
└── README.md                      # 项目说明

10. 技术难点与方案

难点 挑战 解决方案
AI 内容质量 生成内容可能空洞或不恰当 精细化 Prompt + 输出格式约束
AI 服务稳定性 Gemini API 可能超时或失败 自动重试 2 次 + 预设 fallback 内容
动画性能 手机端动画卡顿 Framer Motion 硬件加速 + will-change
冷启动 Render 冷启动慢 保持最小实例存活
费用控制 Gemini API 调用成本 后端全局 Rate Limiting + 输入长度限制 500 字

11. 里程碑

阶段 目标 交付物
M1 项目搭建 前后端脚手架、CI/CD
M2 核心流程 文字输入 → 情绪分析 → 新视角生成
M3 仪式动画 废纸团 → 燃烧 → 晶体生成
M4 完整闭环 图腾展示、心灵博物馆、SQLite 持久化
M5 优化上线 性能优化、部署上线

附录

A. 参考资料

B. 术语表

术语 说明
Valence 情绪效价,正面/负面程度
Arousal 唤醒度,情绪激烈程度
CBT Cognitive Behavioral Therapy,认知行为疗法
ACT Acceptance and Commitment Therapy,接纳承诺疗法
AEDP Accelerated Experiential Dynamic Psychotherapy,加速体验性动力学治疗
ADK Agent Development Kit,Google 的 AI Agent 开发框架