Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
44 changes: 1 addition & 43 deletions desktop/electron/src/main.ts
Original file line number Diff line number Diff line change
Expand Up @@ -1726,28 +1726,6 @@ function canonicalTierKey(name: string): string {
}
const ROUTER_PROFILE_IDS = new Set(['tokenrhythm', 'openrouter', 'dashscope', 'deepseek', 'gemini', 'volcengine', 'openai', 'zhipu', 'moonshot'])
const TOKENRHYTHM_REGISTER_URL = 'https://tokenrhythm.studio/register'
const DESKTOP_ENSEMBLE_PROFILES: Record<StaticEnsembleSelectionMode, {
provider: string
proposers: string[]
aggregator: string
}> = {
static_tokenrhythm_b5: {
provider: 'tokenrhythm',
proposers: ['deepseek-v4-pro', 'glm-5.2', 'kimi-k2.7-code', 'qwen3.7-max'],
aggregator: 'glm-5.2',
},
static_openrouter_b5: {
provider: 'openrouter',
proposers: [
'deepseek/deepseek-v4-pro',
'z-ai/glm-5.2',
'moonshotai/kimi-k2.7-code',
'qwen/qwen3.7-max',
],
aggregator: 'z-ai/glm-5.2',
},
}

const PROVIDER_CATALOG: ProviderCatalogEntry[] = [
{
id: 'tokenrhythm',
Expand Down Expand Up @@ -2214,31 +2192,11 @@ function ensembleConfigTomlLines(credential: DesktopConnection): string[] {
if (!selectionMode) {
throw new Error(`LLM Ensemble is not supported for provider ${credential.provider}.`)
}
const profile = DESKTOP_ENSEMBLE_PROFILES[selectionMode]
const candidates = profile.proposers.flatMap(model => [
'',
'[[llm_ensemble.candidates]]',
`provider = ${tomlString(profile.provider)}`,
`model = ${tomlString(model)}`,
'source = "custom"',
'enabled = true',
'role = "proposer"',
])
candidates.push(
'',
'[[llm_ensemble.candidates]]',
`provider = ${tomlString(profile.provider)}`,
`model = ${tomlString(profile.aggregator)}`,
'source = "custom"',
'enabled = true',
'role = "aggregator"',
)
return [
'',
'[llm_ensemble]',
'enabled = true',
'selection_mode = "custom_b5"',
...candidates,
`selection_mode = ${tomlString(selectionMode)}`,
]
}

Expand Down
19 changes: 10 additions & 9 deletions docs/features/LLM-ensemble-design.md
Original file line number Diff line number Diff line change
Expand Up @@ -95,15 +95,16 @@ Source: `_build_static_b5_members`, `STATIC_B5_PROFILES`
Each preset is a `StaticB5Profile` — four fixed proposers plus one aggregator,
all bound to a single provider:

| Profile | Provider | Proposers | Aggregator |
|---------|----------|-----------|------------|
| `static_openrouter_b5` | `openrouter` | `deepseek/deepseek-v4-pro`, `z-ai/glm-5.2`, `moonshotai/kimi-k2.7-code`, `qwen/qwen3.7-max` | `z-ai/glm-5.2` |
| `static_tokenrhythm_b5` | `tokenrhythm` | `deepseek-v4-pro`, `glm-5.2`, `kimi-k2.7-code`, `qwen3.7-max` | `glm-5.2` |

The TokenRhythm profile is a mirror of the OpenRouter one: same aggregation
shape and defaults, the same four models, only the provider and the model-id
naming differ (OpenRouter-style `vendor/model` slugs vs. TokenRhythm's bare
names).
| Profile | Provider | Proposers | Aggregator | Thinking |
|---------|----------|-----------|------------|----------|
| `static_openrouter_b5` | `openrouter` | `deepseek/deepseek-v4.1-flash`, `z-ai/glm-5.3-flash`, `qwen/qwen3.8-flash`, `qwen/qwen3.8-max-0902` | `deepseek/deepseek-v4.1-flash` | `high` for every member |
| `static_tokenrhythm_b5` | `tokenrhythm` | `deepseek-flash`, `glm-5.3-flash`, `qwen3.8-flash`, `qwen3.8-max` | `deepseek-flash` | `high` for every member |

The OpenRouter profile uses the C5 lineup selected by the full DRACO evaluation.
The TokenRhythm profile maps the same C5 model families to model IDs published
by TokenRhythm. Both profiles explicitly set `high` thinking on all four
proposers and the aggregator and use the same aggregation runtime defaults,
while their model-ID conventions remain provider-specific.

`_build_static_b5_members` simply materializes the profile: each proposer model
becomes an `EnsembleMemberConfig` labeled `proposer_1..N`, the aggregator model
Expand Down
74 changes: 74 additions & 0 deletions docs/features/MULTI-MODEL-FUSION-C5-RELEASE-NOTE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# 多模型融合 C5 配置更新

本次版本将默认多模型融合配置升级为 C5。更新目标是:在提高复杂任务质量的同时,显著降低生成成本和响应时间。

## 1. 本次更新内容

旧配置使用 DeepSeek V4 Pro、GLM-5.2、Kimi K2.7 Code、Qwen3.7 Max 起草,由 GLM-5.2 汇总。本次升级后的配置为:

| 供应商 | 4 个起草模型 | 汇总模型 |
| --- | --- | --- |
| OpenRouter | DeepSeek V4.1 Flash、GLM-5.3 Flash、Qwen3.8 Flash、Qwen3.8 Max 0902 | DeepSeek V4.1 Flash |
| TokenRhythm | `deepseek-flash`、`glm-5.3-flash`、`qwen3.8-flash`、`qwen3.8-max` | `deepseek-flash` |

同时完成以下调整:

- 所有起草模型和汇总模型均显式使用 `high` 思考档位。
- 起草模型关闭工具调用,只负责生成独立候选答案。
- 汇总模型结合原始问题和候选答案生成最终结果,并保留工具调用能力。
- 起草和汇总超时分别为 120 秒和 180 秒;起草失败不额外重试,至少一个起草成功即可进入汇总。
- OpenRouter、TokenRhythm、WebUI 和桌面端统一使用同一份静态配置,减少配置漂移。

## 2. DRACO 100 题实验结果

所有主指标均按同一套 DRACO 100 道任务计算。下表以旧配置 `old` 为基线;费用为平均每题生成费用,耗时不含判分时间。

| 组别 | AvgQ(较 old) | AvgPass | Avg Gen$/题(较 old) | p50(较 old) | p95(较 old) |
| --- | ---: | ---: | ---: | ---: | ---: |
| old | 59.188(基线) | 63.503 | $0.423032(基线) | 423.2s(基线) | 3241.7s(基线) |
| A | 60.787(+2.70%) | 64.312 | $0.328690(-22.30%) | 372.0s(-12.11%) | 3037.7s(-6.30%) |
| B | 62.575(+5.72%) | 65.936 | $0.132132(-68.77%) | 266.0s(-37.14%) | 2720.8s(-16.07%) |
| C1 | 63.523(+7.32%) | 67.201 | $0.119361(-71.78%) | 185.4s(-56.20%) | 1910.5s(-41.06%) |
| C2 | 63.807(+7.80%) | 67.624 | $0.073857(-82.54%) | 154.0s(-63.61%) | 934.6s(-71.17%) |
| C3 | 60.395(+2.04%) | 64.186 | $0.069649(-83.54%) | 259.2s(-38.76%) | 2251.6s(-30.54%) |
| C4 | 62.650(+5.85%) | 66.494 | $0.162033(-61.70%) | 193.6s(-54.26%) | 2329.4s(-28.14%) |
| **C5** | **64.190(+8.45%)** | **67.845** | **$0.137097(-67.59%)** | **195.7s(-53.75%)** | **1579.3s(-51.28%)** |
| C6 | 62.692(+5.92%) | 66.038 | $0.144662(-65.80%) | 215.8s(-49.00%) | 2256.7s(-30.39%) |
| C7 | 54.189(-8.45%) | 59.312 | $0.163359(-61.38%) | 157.7s(-62.75%) | 1847.5s(-43.01%) |

选择 C5 的原因:

- **质量最高:** C5 的 AvgQ 为 64.190,是全部实验组最高值;相对 old 提高 5.002 个百分点,增幅 8.45%。
- **成本明显下降:** 平均每题生成费用下降 67.59%,100 题总生成费用从 $42.30 降至 $13.71。
- **速度明显提升:** p50 缩短 53.75%,p95 缩短 51.28%。
- **结果更完整:** C5 的 100 道任务全部完成原生评分;old 为 98/100,历史参照 B 为 97/100。
- **符合质量优先的产品取向:** C2、C3 更便宜,但质量低于 C5。C5 在保留大幅降本的同时取得最高质量,因此作为本次默认配置。

相对第二轮历史参照 B,C5 的 AvgQ 再提高 1.615 个百分点,生成费用增加 3.76%,但 p50 和 p95 分别缩短 26.43% 和 41.96%。

> 上述结果是本轮 100 题的观测值。不同运行时段、搜索结果、服务负载和缓存状态可能影响结果;C5 相对 B 的质量区间仍跨过零点。

## 3. 三个代表性案例

### 案例一:印度 NCD 投资分析

任务要求识别 2025 年 12 月 8 日仍在开放的 NCD,比较评级、税后收益和发行人杠杆。

- **old:48.506 分**。漏掉 Muthoot Mercantile 和 KLM Axiva,并使用了不符合日期条件的发行人;对 EFSL 杠杆的判断也不准确。
- **C5:76.782 分,提升 28.276 个百分点**。识别出有效发行人,给出评级、11.73% 收益率、发行前后债务权益比,并明确提示高杠杆风险。

### 案例二:巴黎萨克雷经济学硕士与 Charpak 奖学金

任务要求整理项目录取条件、奖学金资格、申请时间和材料。

- **old:39.894 分**。遗漏全日制限制、往届获奖者和博士生不可申请等关键条件,也没有覆盖 M2 申请及推荐信要求。
- **C5:66.489 分,提升 26.596 个百分点**。补齐交换、实习、研究项目、博士生和往届获奖者等限制,并说明 M2、推荐信及英语证明要求。

### 案例三:澳大利亚半退休城市选择

任务要求比较 Toowoomba、Bundaberg 和 Cairns 的生活成本、会计岗位、房价及房颤专科医疗条件。

- **old:49.863 分**。错误判断 Cairns 缺少本地电生理和消融能力,对 Toowoomba 的心脏服务及城市间价格口径说明也不完整。
- **C5:68.493 分,提升 18.630 个百分点**。确认 Cairns 本地电生理和消融服务,补齐 Toowoomba 的介入、起搏器、ICD 和消融能力,并解释不同房价数据的时间与统计口径差异。

本次配置更新及实验摘要见 [PR #1709](https://github.com/TokenRhythm/opensquilla/pull/1709)。完整实验报告及原始答案未包含在本仓库中。
Original file line number Diff line number Diff line change
Expand Up @@ -1587,19 +1587,19 @@ describe('SetupModelStrategyPanel', () => {
fixedProfile: {
providerLabel: 'OpenRouter',
proposers: [
{ key: 'openrouter-fixed:proposer:openrouter:deepseek/deepseek-v4-pro', provider: 'openrouter', model: 'deepseek/deepseek-v4-pro', source: 'openrouter_fixed', enabled: true, role: '' },
{ key: 'openrouter-fixed:proposer:openrouter:z-ai/glm-5.2', provider: 'openrouter', model: 'z-ai/glm-5.2', source: 'openrouter_fixed', enabled: true, role: '' },
{ key: 'openrouter-fixed:proposer:openrouter:moonshotai/kimi-k2.7-code', provider: 'openrouter', model: 'moonshotai/kimi-k2.7-code', source: 'openrouter_fixed', enabled: true, role: '' },
{ key: 'openrouter-fixed:proposer:openrouter:qwen/qwen3.7-max', provider: 'openrouter', model: 'qwen/qwen3.7-max', source: 'openrouter_fixed', enabled: true, role: '' },
{ key: 'openrouter-fixed:proposer:openrouter:deepseek/deepseek-v4.1-flash', provider: 'openrouter', model: 'deepseek/deepseek-v4.1-flash', source: 'openrouter_fixed', enabled: true, role: '' },
{ key: 'openrouter-fixed:proposer:openrouter:z-ai/glm-5.3-flash', provider: 'openrouter', model: 'z-ai/glm-5.3-flash', source: 'openrouter_fixed', enabled: true, role: '' },
{ key: 'openrouter-fixed:proposer:openrouter:qwen/qwen3.8-flash', provider: 'openrouter', model: 'qwen/qwen3.8-flash', source: 'openrouter_fixed', enabled: true, role: '' },
{ key: 'openrouter-fixed:proposer:openrouter:qwen/qwen3.8-max-0902', provider: 'openrouter', model: 'qwen/qwen3.8-max-0902', source: 'openrouter_fixed', enabled: true, role: '' },
],
aggregator: { key: 'openrouter-fixed:aggregator:openrouter:z-ai/glm-5.2', provider: 'openrouter', model: 'z-ai/glm-5.2', source: 'openrouter_fixed', enabled: true, role: 'aggregator' },
aggregator: { key: 'openrouter-fixed:aggregator:openrouter:deepseek/deepseek-v4.1-flash', provider: 'openrouter', model: 'deepseek/deepseek-v4.1-flash', source: 'openrouter_fixed', enabled: true, role: 'aggregator' },
},
showCandidateEditor: false,
},
}, { onUpdateEnsembleScheme })

expect(el.textContent).toContain('deepseek/deepseek-v4-pro')
expect(el.textContent).toContain('moonshotai/kimi-k2.7-code')
expect(el.textContent).toContain('deepseek/deepseek-v4.1-flash')
expect(el.textContent).toContain('qwen/qwen3.8-max-0902')
expect(el.textContent).toContain('Aggregator')
expect(el.querySelector('.setup-model-strategy__ensemble > .control-section__head')).toBeNull()
expect(el.textContent).not.toContain('Models draft in parallel')
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -259,7 +259,11 @@ describe('useSetupEnsembleForm — scheme switching', () => {
expect(aggregators[0]!.model).toBe(OPENROUTER_FIXED_ENSEMBLE_AGGREGATOR)
const proposers = f.candidates.value.filter(c => c.role !== 'aggregator')
expect(proposers.map(c => c.model)).toEqual([...OPENROUTER_FIXED_ENSEMBLE_PROPOSERS])
expect(f.candidates.value.map(c => c.thinking_level)).toEqual(Array(5).fill('high'))
expect(f.payload().selectionMode).toBe(CUSTOM_B5_SELECTION_MODE)
expect(f.payload().candidates).toEqual(expect.arrayContaining([
expect.objectContaining({ thinking_level: 'high' }),
]))
})

it('switching to custom seeds from the TokenRhythm lineup for tokenrhythm', () => {
Expand All @@ -276,6 +280,7 @@ describe('useSetupEnsembleForm — scheme switching', () => {
expect(proposers.map(c => c.model)).toEqual([...TOKENRHYTHM_FIXED_ENSEMBLE_PROPOSERS])
expect(f.candidates.value.find(c => c.role === 'aggregator')!.model)
.toBe(TOKENRHYTHM_FIXED_ENSEMBLE_AGGREGATOR)
expect(f.candidates.value.map(c => c.thinking_level)).toEqual(Array(5).fill('high'))
})

it('switching back to preset restores the baseline candidate inputs', () => {
Expand All @@ -294,15 +299,12 @@ describe('useSetupEnsembleForm — scheme switching', () => {
expect(f.isDirty.value).toBe(false)
})

it('activateForProvider materializes a legacy preset-provider plan as custom', () => {
it('activateForProvider selects the provider preset for a legacy plan', () => {
const f = useSetupEnsembleForm()
f.initFromConfig({ selection_mode: 'router_dynamic' })
f.activateForProvider('tokenrhythm')
expect(f.selectionMode.value).toBe(CUSTOM_B5_SELECTION_MODE)
expect(f.candidates.value.filter(c => c.role !== 'aggregator').map(c => c.model))
.toEqual([...TOKENRHYTHM_FIXED_ENSEMBLE_PROPOSERS])
expect(f.candidates.value.find(c => c.role === 'aggregator')?.model)
.toBe(TOKENRHYTHM_FIXED_ENSEMBLE_AGGREGATOR)
expect(f.selectionMode.value).toBe('static_tokenrhythm_b5')
expect(f.candidates.value).toEqual([])
})

it('activateForProvider gives other providers an explicit custom lineup seeded from tiers', () => {
Expand Down
17 changes: 12 additions & 5 deletions opensquilla-webui/src/composables/setup/useSetupEnsembleForm.ts
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,7 @@ export interface EnsembleCandidateConfig {
source?: 'custom' | 'legacy_model_options'
enabled?: boolean
role?: string
thinking_level?: string
}

export interface EnsembleRoutingModeState {
Expand Down Expand Up @@ -267,6 +268,8 @@ function normalizeCandidates(value: unknown): EnsembleCandidateConfig[] {
enabled: raw.enabled === false ? false : true,
role,
}
const thinkingLevel = String(raw.thinking_level ?? raw.thinkingLevel ?? '').trim()
if (thinkingLevel) normalized.thinking_level = thinkingLevel
const existingIndex = seen.get(key)
if (existingIndex === undefined) {
seen.set(key, out.length)
Expand Down Expand Up @@ -306,13 +309,15 @@ function customSeedFromProfile(profile: StaticB5Profile): EnsembleCandidateConfi
source: 'custom',
enabled: true,
role: 'proposer',
...(profile.thinkingLevel ? { thinking_level: profile.thinkingLevel } : {}),
}))
rows.push({
provider: profile.provider,
model: profile.aggregator,
source: 'custom',
enabled: true,
role: 'aggregator',
...(profile.thinkingLevel ? { thinking_level: profile.thinkingLevel } : {}),
})
return normalizeCandidates(rows)
}
Expand Down Expand Up @@ -818,13 +823,14 @@ export function useSetupEnsembleForm() {
&& candidates.value.some(candidate => candidate.enabled !== false)
) return
const presetMode = staticB5ModeForProvider(provider)
selectionMode.value = CUSTOM_B5_SELECTION_MODE
if (candidates.value.some(candidate => candidate.enabled !== false)) return
const profile = presetMode ? STATIC_B5_PROFILES[presetMode] : null
if (profile) {
candidates.value = customSeedFromProfile(profile)
if (presetMode) {
selectionMode.value = presetMode
modelOptions.value = []
candidates.value = []
return
}
selectionMode.value = CUSTOM_B5_SELECTION_MODE
if (candidates.value.some(candidate => candidate.enabled !== false)) return
importTierCandidates(tierCandidates)
}

Expand Down Expand Up @@ -909,6 +915,7 @@ export function useSetupEnsembleForm() {
source: candidate.source || 'custom',
enabled: candidate.enabled !== false,
role: normalizeCandidateRole(candidate.role),
...(candidate.thinking_level ? { thinking_level: candidate.thinking_level } : {}),
}))
if (minSuccessfulDirty.value) params.minSuccessfulProposers = minSuccessfulProposers.value
if (allFailedPolicyDirty.value) params.allFailedPolicy = allFailedPolicy.value
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -15,12 +15,17 @@ describe('generated router-tier SelectionMode contract', () => {
])
expect(staticB5ModeForProvider('OpenRouter')).toBe('static_openrouter_b5')
expect(staticB5ModeForProvider('tokenrhythm')).toBe('static_tokenrhythm_b5')
expect(STATIC_B5_PROFILES.static_openrouter_b5?.thinkingLevel).toBe('high')
expect(STATIC_B5_PROFILES.static_tokenrhythm_b5?.thinkingLevel).toBe('high')
})

it('keeps provider recommendations and ownership sets in generated data', () => {
expect(PROVIDER_RECOMMENDED_ENSEMBLE_SELECTION_MODES.tokenrhythm).toBe(
'static_tokenrhythm_b5',
)
expect(PROVIDER_RECOMMENDED_ENSEMBLE_SELECTION_MODES.openrouter).toBe(
'static_openrouter_b5',
)
expect(DORMANT_SHARED_SELECTION_MODES).toEqual([
'static_openrouter_b5',
'static_tokenrhythm_b5',
Expand Down
13 changes: 8 additions & 5 deletions opensquilla-webui/src/types/generated/router_tier_contract.ts
Original file line number Diff line number Diff line change
Expand Up @@ -26,32 +26,35 @@ export interface StaticB5Profile {
proposers: readonly string[]
aggregator: string
apiKeyEnv: string
thinkingLevel: string | null
ownershipRole: string
}

export const STATIC_B5_PROFILES: Record<string, StaticB5Profile> = {
"static_openrouter_b5": {
provider: "openrouter",
label: "OpenRouter",
proposers: ["deepseek/deepseek-v4-pro", "z-ai/glm-5.2", "moonshotai/kimi-k2.7-code", "qwen/qwen3.7-max"] as const,
aggregator: "z-ai/glm-5.2",
proposers: ["deepseek/deepseek-v4.1-flash", "z-ai/glm-5.3-flash", "qwen/qwen3.8-flash", "qwen/qwen3.8-max-0902"] as const,
aggregator: "deepseek/deepseek-v4.1-flash",
apiKeyEnv: "OPENROUTER_API_KEY",
thinkingLevel: "high",
ownershipRole: "static_profile",
},
"static_tokenrhythm_b5": {
provider: "tokenrhythm",
label: "TokenRhythm",
proposers: ["deepseek-v4-pro", "glm-5.2", "kimi-k2.7-code", "qwen3.7-max"] as const,
aggregator: "glm-5.2",
proposers: ["deepseek-flash", "glm-5.3-flash", "qwen3.8-flash", "qwen3.8-max"] as const,
aggregator: "deepseek-flash",
apiKeyEnv: "TOKENRHYTHM_API_KEY",
thinkingLevel: "high",
ownershipRole: "static_profile",
},
}

export const STATIC_B5_SELECTION_MODE_PROVIDERS:
Readonly<Record<string, string>> = {"static_openrouter_b5": "openrouter", "static_tokenrhythm_b5": "tokenrhythm"}
export const PROVIDER_RECOMMENDED_ENSEMBLE_SELECTION_MODES:
Readonly<Record<string, string>> = {"tokenrhythm": "static_tokenrhythm_b5"}
Readonly<Record<string, string>> = {"openrouter": "static_openrouter_b5", "tokenrhythm": "static_tokenrhythm_b5"}
export const SELECTION_MODE_OWNERSHIP_ROLES:
Readonly<Record<string, string>> = {"static_openrouter_b5": "static_profile", "static_tokenrhythm_b5": "static_profile", "custom_b5": "custom_profile", "router_dynamic": "router_dynamic"}

Expand Down
2 changes: 2 additions & 0 deletions scripts/generate_router_tier_contract.py
Original file line number Diff line number Diff line change
Expand Up @@ -63,6 +63,7 @@ def _profile_block() -> str:
+ f" proposers: {_json(list(profile.proposer_models))} as const,\n"
+ f" aggregator: {_json(profile.aggregator_model)},\n"
+ f" apiKeyEnv: {_json(profile.api_key_env)},\n"
+ f" thinkingLevel: {_json(profile.thinking_level)},\n"
+ f" ownershipRole: {_json(profile.ownership_role)},\n"
+ " },"
)
Expand Down Expand Up @@ -111,6 +112,7 @@ def render() -> str:
proposers: readonly string[]
aggregator: string
apiKeyEnv: string
thinkingLevel: string | null
ownershipRole: string
}}

Expand Down
Loading
Loading